Skip to content

The conversation bought a discount on checking.

This is one result from the free test Is Your Trust in AI Backed by Evidence? It describes the most common way AI-assisted work goes wrong: not a dramatic hallucination, but an output that crossed into your work with less resistance than its consequence deserved. Something in the exchange made the answer feel settled, and the settling happened before anything outside the conversation had a say.

The discount, described precisely

Here is the mechanism in one sentence. When a session feels like it understands you, you skim output you would have reviewed line by line if it had arrived from a form.

That is the whole trade. You did not decide to check less. Nobody decides that. A social channel carries tags that we never audit, and one of them is understanding. Fluent, on-topic, in-your-own-vocabulary replies get read below awareness as comprehension, and comprehension is something you extend credit to. The credit is real. Nothing on the other end earned it.

The result is that fluency becomes the practical acceptance test. Not because you believe fluency implies correctness, but because the artifact never met anything else.

Four versions of the same gap

The test weighs what the output was compared against before you used it, against how hard the mistake would have been to undo. This profile appears when the gap between those two is the largest thing in your session.

Four versions show up, and they need different fixes.

  • Nothing separate. The answer's only test was whether it read well. This is the pure form of the discount.
  • The conversation as its own witness. You checked later replies against earlier replies. Consistency inside one thread is a property of the thread, not evidence about the world.
  • A spot check that was too narrow. You checked the part that looked suspicious. The consequence ran through a part that looked fine.
  • Real evidence, wrong target. You did check something outside, and it did not settle the claim that would reverse the decision.

What it looks like in practice

A marketer asks for competitor pricing, gets a clean comparison table, and checks one row against the vendor's site. The row was right. The row that decided the positioning was invented, and it was not the one that looked odd.

A developer reviews an AI-written patch after a long session of explaining the codebase. The explanation went well, the model clearly understood the constraints, and the diff gets a scan instead of a read. Review fatigue research describes this exact drift: sustained review shifts from evaluation to surface scanning, and it is not caused by people caring less.

An analyst asks a follow-up question, gets a confident answer that matches the earlier one, and treats the match as confirmation. Two replies from the same system agreeing is not two sources agreeing. It is one source, twice.

Not a claim that the output was wrong

It is not a claim that the output was wrong. Most outputs in this profile are fine, which is exactly why the pattern persists. Automation complacency needs a reliable system to form around: high reliability, competing demands on your attention, and a track record of being right are the three ingredients, and all three are usually present.

It is also not the same as having no owner. If your session ended with a clear answer and no one positioned to catch a mistake, the more precise map is Responsibility drifted into the chat. The discount is about evidence. That profile is about who says no.

The move: find the decisive claim, then leave the room

The principle: one check outside the conversation beats five inside it. Independence is the property that does the work, not thoroughness.

Do it in three steps, and keep it small enough to survive a busy day.

  1. Name the one claim that would change what you do. Not the whole output. The single fact, number, behavior, or assumption where being wrong reverses the decision. Write it down, because writing it exposes how often the decisive claim was never explicitly stated.
  2. Check that claim against something the model did not produce. A primary source rather than another summary. A run rather than a description of what would happen. A recomputation rather than a restated total. A person rather than a paraphrase.
  3. Record the verdict. If the source contradicts the claim, cannot support it, or leaves it uncertain, revise or reject. An unfinished check is not a passed check.

If you only remember one rule from this profile, make it the second step. Do not let the conversation witness itself.

Then raise the floor before the next session, not during it

Checking harder in the moment is a willpower fix, and willpower is exactly what a fluent session spends. The durable version is a floor set in advance: for each recurring task, the minimum evidence required at each level of consequence, written down before you are looking at an answer you like.

The floor should track how hard the mistake is to undo, not how confident the output sounds. Confidence in the text is the one input that carries no information about the world.

A worked example. A team that ships AI-assisted customer emails decides: internal drafts need nothing, anything sent externally needs every factual claim traced to a source document, anything touching billing needs a named approver. That took ten minutes to write and removes the decision from the moment it is most compromised.

Where the fix would overshoot

Do not respond by banning conversational tools or moving all work into pipelines. The channel is not what made the answer wrong, and structured interfaces do not make output more correct. They only make unsupported warmth less available as a reason to trust it, which matters when consequences are high and is overhead when they are not.

If the reason nothing got checked was that the session ended in an apology that felt like closure, go to The apology closed more than it repaired. instead. That is a different failure with a different repair.