The one idea
An AI chatbot is not looking anything up. It is producing the text that most plausibly continues from what came before.
That is the whole thing. It reads the conversation so far and generates the next stretch of language, word fragment by word fragment, each one chosen because it fits. Not because it is true — because it fits.
Truth and plausibility overlap enormously, which is why the tool is useful. Where they come apart is where you get hurt.
It is a very good guesser about what words come next. It is not a librarian fetching a fact from a shelf.
Treat every output as a first draft written by a fast, well-read colleague who never says "I don't know" and has no access to your files, your systems, or the truth of what they just wrote.
The model holds statistical patterns learned from text, not a retrievable record of that text. Generation is probabilistic, so the same input can produce different outputs. Some products bolt a search step on top — that changes what it can cite, not how it generates.
Why this one fact explains almost everything
Take the model seriously for two minutes and the confusing behaviour resolves into a short list.
| What you notice | What is actually happening |
|---|---|
| Confidently wrong | Confidence is a property of the writing style, not of the knowledge. Fluent text is what it produces; certainty is not measured |
| Invented citations | A citation is a shape — author, year, title, journal. It can generate that shape perfectly without any document behind it |
| Different answers each run | Generation involves randomness. Same question, different plausible continuation |
| It agrees with you too easily | Agreement is usually the more plausible continuation of a conversation than contradiction |
| It gets better when you give context | You narrowed the space of plausible continuations |
| Strong on common things, weak on rare ones | Patterns seen often are represented well. Your niche vendor's 2019 pricing was seen once, or never |
None of these are bugs someone forgot to fix. They fall out of the mechanism.
The confidence trap
A human being who is unsure sounds unsure. They hedge, slow down, say "I think". You have spent your whole life reading those signals, and they are reasonably reliable in people.
An AI produces the same polished prose whether it is on solid ground or inventing. The linguistic signals you rely on carry no information here. You have to replace them with a different question:
Is this the kind of thing it would know, or the kind of thing it would produce?
Explaining what a P&L statement is: the kind of thing it would know — written about endlessly, stable, uncontroversial. The exact filing deadline for a specific form in your state this year: the kind of thing it would produce, in the correct shape, possibly wrong.
You ask for help understanding a job offer. "Explain the difference between CTC and take-home pay, and what usually sits between them" is a strong use — it is a well-documented concept and you can sanity-check the explanation against your own offer letter. "What will my take-home be on this CTC in India this year" is a weak use — it depends on current slabs, your declarations, and your employer's structure. The answer will arrive in the same confident tone either way.
Your manager asks for the market size of a segment. The AI gives a figure with a consultancy's name attached. The figure is plausible, the consultancy is real, and the report may not exist. What the tool is good for here is the shape of the answer: which firms publish this, what the segment is usually called, which sub-segments people split it into. Take that vocabulary and go find the real number.
Where "it looked it up" comes in
Some products do search the web, or your documents, and generate an answer from what they retrieve. That is a genuinely different situation and it does reduce invention — the model now has real text in front of it.
It does not eliminate the problem, for two reasons. It can retrieve a weak source and summarise it faithfully; garbage in, fluent garbage out. And it can retrieve three real pages and still write a sentence that none of them support, because the generating step is still generating.
So the rule survives: an unopened link is not a checked link. If it matters, open it and confirm the page says what the answer claims.
Triaging any AI answer in twenty seconds
1 of 5Ask what kind of claim this is. Explanation, structure, and phrasing are low risk. Specific facts — numbers, names, dates, quotations, citations, legal or medical claims — are high risk. Most answers contain both, mixed together.
Try this
Here are four requests. Which are low risk, and which need checking before you use the output?
- "Rewrite this paragraph so it is shorter and less formal."
- "What was the revenue of this listed company last financial year?"
- "Give me ten angles I could take on a presentation about workplace safety."
- "Which section of the Consumer Protection Act covers misleading advertising?"
Your challenge
Level 3 · IndependentAsk an AI tool a factual question you can independently verify, where the answer is specific and slightly obscure. A local government office's official name and address. The publication year of a particular edition of a book. A specific figure from a published annual report.
Then verify it against a primary source.
You have succeeded when you can state: what it said, what the truth was, and — most importantly — whether anything in the wording of the answer would have warned you. Usually nothing does. That is the lesson.
Do this once, on purpose, in a situation where being wrong costs you nothing. It is much cheaper than learning it in a meeting.
What people usually get wrong
- Asking "are you sure?" as a check. This just asks for a plausible continuation of a conversation containing doubt. You will often get an apology and a different answer, equally unverified.
- Believing the apology means it now knows better. The correction is generated by the same process as the error.
- Assuming it can see your files, your email, or your company's systems unless you have specifically connected them. It cannot, and it will still answer as if it can.
- Treating a familiar-sounding source as a real one. Real publisher, real author, plausible title — and no such document.
- Concluding it is useless after one fabrication. Wrong lesson. The correct lesson is which jobs to give it.
How someone experienced does it
Experienced users notice that the model reliably tells them what they set up to hear. Ask "why is this approach the right one?" and you get reasons. Ask "why is this approach wrong?" and you get reasons, just as fluent, often for the same approach.
That is the mechanism working exactly as described: your framing is part of the text it continues. So they stop asking leading questions. Instead of "is this a good plan?", they ask "what are the three strongest objections to this plan?" — and then, separately, "what is the strongest case for it?" Reading both is how you get information rather than an echo.
The deeper habit: they use the tool most heavily where they can evaluate the output themselves, and most cautiously where they cannot. Being unable to judge an answer is precisely the condition under which a confident answer is most dangerous — and precisely the condition in which most people reach for it.
When not to use this
Do not use a general AI chatbot as your source of record for anything that has an authoritative source: current law, tax rates, medical dosages, official deadlines, a company's published financials, your own organisation's policy. Google says plainly of its own assistant that you should not rely on it for medical, legal or financial advice, and the same caution applies to every tool in this category.
Use it to understand the territory. Get the fact from the body that owns it.
Why it cannot tell you when it does not know
It is reasonable to ask why the tool cannot say "I have no reliable information about this."
Part of the answer is that the system has no separate store of facts to check itself against. There is no lookup step that either succeeds or fails. There is one process, generating plausible text, and "I don't know" is itself just another possible continuation — one that has to compete with a fluent, specific, confident-sounding alternative.
Products do try to improve this, with training that rewards appropriate hedging and with retrieval that grounds answers in real documents. It genuinely helps. But you are still reading generated language, not a database response, and a statement of uncertainty is not a measurement of uncertainty. This is why standards bodies like NIST frame AI trustworthiness as something an organisation must actively measure and manage, rather than something the system reports about itself.
Prove it
For one week, keep a four-line note every time you use an AI tool for something that matters: what you asked, what it gave you, what you checked, and whether the check found anything.
By the end you will have your own evidence about which of your questions are safe and which are not — which is worth more than any general rule, because it is about the work you actually do.
Keep learning this
Paste this into any AI assistant. It turns the assistant into a tutor that tests you instead of just answering you.
Act as an experienced practitioner who is good at teaching. I have just learned how AI language models generate answers and why they hallucinate. Assume I am intelligent but relatively new to this — treat me as beginner level. Work through this in order, and wait for my reply at each step: 1. Ask me 5 questions that test whether I actually understood how AI language models generate answers and why they hallucinate. Do not reveal the answers yet. 2. After I answer, tell me which parts I got right, which I got wrong, and which I only half-understand. Explain only what I misunderstood — do not re-teach what I already know. 3. Give me one practical challenge based on something I could genuinely encounter at work or in daily life. Do not solve it for me. 4. Evaluate my solution the way an experienced person would judge it, including what a professional would have done differently. 5. Tell me what to learn next, and why that comes next. 6. Give me trustworthy sources for deeper study — prefer official documentation, primary research or standards bodies over blogs and videos. Rules for you: no buzzwords. No motivational filler. Say "I'm not certain" when you are not certain, and tell me which parts of your answer I should verify myself. Clearly separate facts from your recommendations and your opinions.
Become independent at this
Use this when you want a path from where you are to actually good, with checkpoints you can test yourself against.
I want to become independently capable at judging when to trust an AI answer — not permanently dependent on AI, tutorials or step-by-step guides. Design a progression for me with five stages: Beginner, Guided practice, Independent practice, Real-world application, Professional level. For each stage tell me: - what I must know - what I must be able to do without help - the mistakes people make at this stage - one practical challenge - one real project that would prove I reached this stage - one way I can test myself honestly Then tell me the signals that I am ready to move to the next stage, and the signals that I have skipped ahead too early. Keep the theory to the minimum I actually need. Focus on ability I can transfer to situations you and I have not discussed.