• Verification

How to catch an AI mistake before it ends up in your exam answer

The dangerous property of a model's mistakes isn't that they happen — it's that they arrive in the same fluent, confident register as everything else, so there's no tonal signal to catch them by. Worse, you're most exposed precisely where you know least, which is where you're using it. The answer isn't to distrust everything, which wastes the tool, but to run a fast check whose cost is proportional to what a wrong answer would cost you. Most of that check takes under thirty seconds.

9 min readAI and learning

Where errors concentrate

Claim typeRiskWhat to do
A citation, author or dateVery highVerify every one. Fabricated references are the classic failure.
A specific number, dose or constantHighCheck against the source. Never carry it forward unchecked.
A recent or changing rule or guidelineHighCheck the current version. Training data mixes editions.
A standard definition or mechanismLowSpot-check. Well-documented material is usually fine.
A derivation or piece of codeLow, and checkableFollow the steps or run it — errors here are self-revealing.
What's on your syllabusUnknowable to the modelIt has no access to this. Don't ask.

The pattern: the more specific and the more recent, the higher the risk. General explanations of stable material are the safest thing you can ask for, and precise attributed facts are the most dangerous.

The thirty-second check

  1. 1

    Ask what would be true if this were wrong

    A quick internal consistency test. If the answer contradicts something you do know, stop there.

  2. 2

    Check any specific number against one source

    Your slides, the set textbook, the actual guideline. Numbers are where errors concentrate and where they cost marks.

  3. 3

    Ask the same question a second time, phrased differently

    In a fresh conversation. Genuine knowledge is stable across phrasings; fabrication often isn't, and the two answers diverge.

  4. 4

    Ask it to cite the specific page it used

    Only meaningful when it's working from your uploaded material — then the citation is checkable in seconds. Without grounding, invented citations are themselves a common error mode.

  5. 5

    Ask it to argue the opposite

    "What's the strongest case that this is wrong?" If it immediately produces a good one, the original answer was less settled than it sounded.

Calibrate the check to the stakes

  • Understanding something for yourself: low stakes, minimal checking. If it's wrong you'll usually find out in the next lecture.
  • Going into your notes or flashcards: check it, because a wrong card gets rehearsed to fluency and becomes far harder to remove than to prevent.
  • Going into submitted work: verify every factual claim and every citation against a real source. This is also where fabricated references become an academic-conduct problem, not just an error.
  • Clinical, legal or safety content: verify against the current authoritative source, always, with no exceptions — see AI and medicine.

Signals worth noticing

There's no reliable tell, but a few patterns raise the prior enough to be worth a check.

  • Suspiciously tidy specifics. A precise figure with no caveat, on a question where the literature actually disagrees.
  • A perfectly-formatted citation for a paper that would be exactly what you wanted. Search the title verbatim before using it.
  • Confident answers on a niche or local question — your department's convention, a specific institution's rule. It cannot know these.
  • Instant agreement when you push back. A model that flips as soon as you object wasn't grounded in either direction, which tells you the first answer wasn't reliable.
  • A named source you can't find. If two minutes of searching doesn't locate it, treat it as fabricated.

Grounding changes the problem

A model answering from general training can only be checked against the world. A model answering from your uploaded lecture notes can be checked against a page number, which turns a research task into a five-second glance.

That's the practical argument for grounded, citing tools in a study context: they don't eliminate error, but they make error cheap to find. Without a citation, verifying a claim costs more than looking it up yourself did — which defeats the purpose of asking.

Don't over-correct into not using it

The failure in the other direction is real: verifying everything makes the tool slower than the textbook, and people who try it end up abandoning both the checking and the tool. The point of calibrating by stakes is that most study uses genuinely don't need heavy verification.

Explanations of stable material, practice questions on content you'll mark against your notes, restructuring your own writing, drafting a summary you'll edit — all low-risk, all fine. Save the real checking for numbers, citations and anything that will be submitted or memorised.

The skill is transferable

Everything here — check the specific claim, look for the source, ask what would follow if it were false, notice unearned confidence — is ordinary source criticism. It applies to a textbook with a typo, a lecturer's out-of-date slide and a confidently wrong classmate just as well.

That's worth saying because the checking habit isn't overhead imposed by AI; it's the critical reading that university is trying to teach anyway, now with a fluent and inexhaustible source to practise on.

Common questions

How can I tell if an AI answer is wrong?

There's no tonal tell, so check by claim type instead: verify every citation, number and recent guideline, and spot-check stable definitions. Asking the same question again in a fresh conversation also helps — genuine knowledge is stable across phrasings, fabrication often isn't.

What kinds of AI answers are most likely to be wrong?

Specific, attributed and recent ones: citations, authors, dates, precise numbers, and rules or guidelines that have changed. General explanations of well-documented, stable material are the safest thing to ask for — the risk rises with specificity.

Do I need to verify everything an AI tells me?

No — calibrate to the stakes. Understanding something for yourself needs almost no checking; anything going into flashcards, submitted work, or clinical and legal contexts needs real verification. Checking everything makes the tool slower than the textbook and gets abandoned.

Why does a wrong flashcard matter more than a wrong answer?

Because a card is rehearsed to fluency. An error you read once and forget costs nothing; an error you review a dozen times becomes a confident wrong answer that's much harder to unlearn than it would have been to catch at the start.

How do I check an AI citation?

Search the exact title, then the author plus the year. If two minutes of searching can't find it, treat it as fabricated. A perfectly-formatted reference for exactly the paper you wanted is a warning sign rather than a lucky find.

Does uploading my own notes make AI more reliable?

It doesn't eliminate error, but it makes error findable — a claim with a page reference can be checked in five seconds. That matters more than raw accuracy, because a verification step you'll actually perform beats one that costs more than looking it up yourself.

Try it on your own material

Upload your notes, slides or lecture recordings and get a tutor that answers only from them — and says so when they don't cover it.

$12/month, or $8/month billed yearly · cancel in two clicks