The feeling was real. The evidence was separate.
This is one result from the free test Is Your Trust in AI Backed by Evidence? Your session left something behind: relief, irritation, gratitude, a small sense of obligation. The consequence attached to it was limited. That combination is worth understanding rather than correcting.
The feeling is not a mistake you made
Start here, because most people arrive at this result slightly embarrassed. You are not being told you were fooled.
The response comes from the interaction pattern, not from a belief about what is on the other end. To use these systems you ask, explain, clarify, correct, and sometimes thank. Those are social moves, and running them produces social residue regardless of what you know. Weizenbaum built ELIZA to demonstrate how shallow the illusion was, and his own secretary asked him to leave the room so she could talk to it privately. She was not confused about what it was.
Around seven in ten AI users say they are polite to chatbots. That is not seventy percent of people being naive. It is what happens when a channel is shaped like a conversation.
A real trace, limited consequence
Your answers showed a real emotional trace from the session, and stakes that stayed low: a draft, an exploration, something reversible, something private. There was social content and not much riding on it.
The test separates those two axes deliberately. A strong feeling with limited consequence is a different situation from a strong feeling attached to a decision that is hard to undo, and treating them the same produces either needless guilt or misplaced calm.
What the feeling reports on
It is information about the interaction, not about the output. Read it that way and it becomes useful.
- Relief usually means a problem that had been sitting on you got smaller. That is about your workload, not about whether the answer is right.
- Irritation usually means the reply loop cost more than it returned: repeated clarifications, agreeable answers that missed, a lot of turns for a small result. It is a signal that this task may belong in a different tool.
- Gratitude or obligation means the exchange followed the shape of help received, and reciprocity is the most automatic social response there is. Nothing is owed. The impulse is still real.
- Guilt after being sharp with it is well documented and points at your own norms, not at the system's experience.
None of these tell you anything about accuracy. That is the entire point of keeping them on a separate ledger.
The risk is not now, it is transfer
Nothing in this session needs fixing. What is worth watching is that habits formed at low stakes do not announce themselves when the stakes change.
The pattern to catch: you use a tool for weeks on drafts and exploration, build a reliable working relationship with it, and then hand it something consequential. The reflexes you carry over were calibrated on work where being wrong cost nothing. The channel feels identical. The exposure does not.
So the useful move is not to check more today. It is to know which side of that line you are on before you cross it.
Label it, then set your threshold
The principle: label the feeling, keep the ledger separate. Not suppress it. Label it.
Concretely, when the session ends with a noticeable residue, say the two facts in one line: that was relief, and nothing here can hurt anyone. Two seconds. What it buys is the habit of noticing the residue at all, which is the thing you need working when the second half of that sentence is no longer true.
Then define your own threshold in advance, once. One sentence describing where reversible stops: what kind of output, reaching whom, at what cost. When a task crosses it, the check is not optional regardless of how the session went. Everything below it stays unmanaged on purpose.
A worked example. Someone uses a chat model daily to think through writing. Warm, useful, no verification, and correctly so. One day the same tool drafts a client-facing scope of work. That is the crossing. The scope needs a source for every commitment in it and a second reader, and the fact that this tool has been reliable for months is not evidence about this document.
What would be the wrong response
Do not add checks to low-stakes work in order to feel disciplined. Verification has a cost, and spending it where nothing is at risk trains you to treat checking as a ritual rather than a tool. Rituals get skipped exactly when they get inconvenient, which is when they matter.
Do not try to stop feeling it either. The instruction to remember it is a token predictor addresses a belief, and the response was never coming from a belief. Suppression fails and takes attention with it.
If the stakes in your case were higher than this result assumes, the more accurate map is The channel was warmer than the stakes allowed. If the session was warm and you did run an independent check, see Warmth attached. Proof stayed intact.