You asked a chatbot about a regulation, a statistic, or a quote, and it replied with a clean paragraph and a link. You're about to paste it into a report. How do you know it's right?
You check it, but you don't check everything. The routine below takes about ten minutes for a typical answer: pull out the claims, rank them by what breaks if they're wrong, check the top few against a source you opened yourself, and treat the rest as unverified. I also wrote a script for the mechanical part, checking that cited pages load and that quoted text appears on them, and I'll show you what it got right and where it fails.
This is deliberately general. For paper-level work, the academic research papers post has a DOI checker against Crossref, which I won't repeat. For why models invent things in the first place, the hallucinations deep dive is the place to go. This post is the routine you run afterwards.
Step 1: What are the claims in this answer?
An AI answer feels like one thing. It's actually twelve small statements, and two of them carry all the risk. Split it before you check it.
You can ask a model to do the splitting. This is an untested template, and the output is a to-do list, not a verdict:
Below is an answer. List every factual claim in it as a separate numbered
line: numbers, dates, names, quotes, causal claims, and "X says Y"
attributions. Do not judge whether they are true. For each, say what kind
of evidence would settle it (a document, a database, an official page).
Do not add claims that are not in the text.
[PASTE ANSWER]
Splitting into a list matters because a sentence with three facts in it can be two-thirds right, and a paragraph read at speed tends to get a pass or a fail as a whole. A list makes you look at each fact.
Step 2: Which claims are worth checking first?
You won't check everything, so rank. Score each claim on two questions: how bad is it if this is wrong, and how likely is it to be wrong?
| Likely wrong | Why |
|---|---|
| Exact numbers, percentages, prices | Models produce plausible digits. A number is the easiest thing to get slightly wrong. |
| Direct quotes and who said what | Paraphrases get presented as quotes. |
| Dates and "first/largest/only" claims | Superlatives are rarely checked and often false. |
| Recent events, current prices, current rules | Past the model's training data unless it searched. |
| Niche local detail: a small company, a regional law | Thin training data, so more guessing. |
| Citations: author, year, title, link | The classic fabrication. |
Anything you're going to publish, send to a client, or act on with money or health goes to the top. Background colour, general explanations and widely known facts go to the bottom. Be honest with yourself here: if the claim is load-bearing, it gets checked, however good the answer sounds.
Step 3: What counts as a real check?
A check is real when the evidence comes from somewhere the model can't have shaped. In rough order of strength:
- The primary source. The law text, the company filing, the paper, the vendor's documentation, the statistics office's own table.
- A second independent source. Independent means it didn't copy the first. Two sites repeating one press release is one source.
- A person who knows. Your accountant, your doctor, your colleague who owns the system.
- The same model, asked again. Not a check. It's useful only to see whether the answer is stable, because a claim that changes between runs is a claim to distrust.
Two traps. First, "I Googled it and found the same sentence" often finds content that was itself written or paraphrased by a model. Look for a source with a named author, a date, and something at stake. Second, a source existing is not the source supporting the claim. Open it and find the sentence.
Can a script do any of this?
Part of it. A script can't judge truth, but it can do two boring things quickly: confirm that a cited page loads, and confirm that the quoted wording is on it. I ran this one on 2026-10-08. It uses curl through Python's standard library, strips HTML, and looks for the claimed sentence. If the exact sentence isn't there, it counts how many three-word phrases from the claim do appear. If the page 404s, it asks the Wayback Machine's availability API for a snapshot.
import html
import json
import re
import subprocess
import urllib.parse
# (url the AI cited, sentence the AI says the page contains)
claims = [
("https://en.wikipedia.org/wiki/Python_(programming_language)",
"conceived in the late 1980s by Guido van Rossum"),
("https://en.wikipedia.org/wiki/Python_(programming_language)",
"created in 1995 by Linus Torvalds at Google"),
("https://example.com/",
"This domain is for use in illustrative examples in documents"),
("https://example.com/2024/ai-report-final",
"Seventy percent of analysts now use AI daily"),
]
def curl(url):
out = subprocess.run(
["curl", "-sL", "-m", "20", "-A", "Mozilla/5.0", "-w", "\n%{http_code}", url],
capture_output=True, text=True,
).stdout
body, _, code = out.rpartition("\n")
return int(code or 0), body
def norm(s):
s = html.unescape(re.sub(r"<[^>]+>", " ", s))
return re.sub(r"\s+", " ", s).lower()
def overlap(claim, page):
w = claim.lower().split()
grams = [" ".join(w[i:i + 3]) for i in range(len(w) - 2)]
return sum(g in page for g in grams), len(grams)
def wayback(url):
code, body = curl("https://archive.org/wayback/available?url="
+ urllib.parse.quote(url, safe=""))
try:
return json.loads(body)["archived_snapshots"].get("closest", {}).get("url")
except Exception:
return None
for url, claim in claims:
code, body = curl(url)
if code != 200:
print(f"HTTP {code} UNREACHABLE {url}")
print(f" claim: {claim}")
print(f" Wayback snapshot: {wayback(url)}\n")
continue
page = norm(body)
if norm(claim) in page:
verdict = "EXACT QUOTE FOUND"
else:
hits, total = overlap(claim, page)
verdict = f"NO EXACT MATCH ({hits}/{total} three-word phrases found)"
print(f"HTTP {code} {verdict}")
print(f" url: {url}")
print(f" claim: {claim}\n")
Output from my run:
HTTP 200 NO EXACT MATCH (5/7 three-word phrases found)
url: https://en.wikipedia.org/wiki/Python_(programming_language)
claim: conceived in the late 1980s by Guido van Rossum
HTTP 200 NO EXACT MATCH (0/6 three-word phrases found)
url: https://en.wikipedia.org/wiki/Python_(programming_language)
claim: created in 1995 by Linus Torvalds at Google
HTTP 200 NO EXACT MATCH (4/8 three-word phrases found)
url: https://example.com/
claim: This domain is for use in illustrative examples in documents
HTTP 404 UNREACHABLE https://example.com/2024/ai-report-final
claim: Seventy percent of analysts now use AI daily
Wayback snapshot: None
Now the honest reading, because none of these says what you'd hope.
Row 1 is true, but the script says "no exact match." Wikipedia's text doesn't use my wording. It says Guido van Rossum "began working on Python in the late 1980s", which supports the claim in different words. 5 of 7 phrases matched, and that's what a correct paraphrase looks like. So exact-match tools produce false alarms on honest paraphrases.
Row 2 is false, and 0 of 6 phrases matched. That's the useful signal. A fabricated attribution leaves no trace on the page.
Row 3 is a drifted page. I wrote that quote from memory, and the wording is off. The page, which I fetched, says "This domain is for use in documentation examples without needing permission." I don't know whether the page ever read differently. Either way it shows the pattern: the link works, the quote doesn't match, and you can't tell from the link alone whether the model misremembered or the page changed.
Row 4 is a dead link with no archive copy. A 404 plus no Wayback snapshot means the claim has nothing behind it. That doesn't prove fabrication, since archive coverage is patchy, but it moves the claim to "unverified" until you find the source another way.
So the script sorts, it doesn't decide. Zero overlap and 404 mean look hard. High overlap means read the passage, because overlap can't distinguish "the page says this" from "the page mentions the same words and says the opposite". A number being on the page also doesn't tell you what it refers to. I didn't build the part that checks numbers or negation, and I wouldn't trust a script that claimed to.
Step 4: What do you do when the AI won't give a source?
Don't accept "according to recent studies" or "experts agree". Ask for the specific document, then verify that document exists. If it can't name one, the claim is background at best.
When you have a search-enabled tool, ask it to quote the exact passage and give the URL, then open the URL. Tools that browse can still summarise a page inaccurately. The failure shifts from "invented from memory" to "real page, wrong reading," which is easier to catch if you actually click through.
What are the red flags in an AI answer?
None of these proves an answer wrong. They tell you where to spend your checking time.
- Precision with no source. "Exactly 37.4% of..." and no origin.
- A tidy quote from a famous person. Famous people are misquoted constantly, and models inherit the misquotes.
- Everything agrees with you. If you asked a leading question and got a confirming answer, ask the opposite question in a fresh chat and compare.
- A link that works but goes to the homepage or a generic page. The model got the domain right and guessed the rest.
- Confidence about something recent. Ask the model what its knowledge covers, or whether it searched, before you rely on it.
- Suspiciously smooth coverage of a niche topic. Thin-data topics get filled with plausible guesses.
- Numbers that don't add up. Percentages that sum past 100, a total that doesn't match its parts, a date that falls on the wrong weekday. Mental arithmetic catches more than you'd think.
A cheap consistency test: ask the same question in two or three fresh chats. Details that change between runs (a year, a name, a figure) are the ones the model is making up. Details that stay put could still be wrong, but they're less likely to be pure invention.
What should I never trust without checking?
- Numbers, dates and quotes that matter to your decision
- Legal, medical, tax and regulatory specifics (check the official source, and a professional for anything consequential)
- Anything about events after the model's training data, unless it searched and you read the pages
- Citations: author, title, year, DOI, URL
- Code that touches money, security or user data (run it, test it)
- Summaries of a document you haven't read, when the summary will stand in for the document
You can lean on AI more for brainstorming, structure, rewording, explaining a concept you'll verify later, and first drafts of things you'll fully rewrite. The more the output depends on the model's memory of facts, the less you should lean.
The ten-minute routine
- Split the answer into numbered claims (Step 1 prompt).
- Mark the three claims that cost you most if wrong.
- For each: find the primary source yourself, or run the cited URL through the script and read the passage.
- Mark each claim verified, unverified, or wrong. Keep unverified claims out of anything you publish, or label them as unverified.
- Ask the model to revise the answer using only the claims you verified, and nothing else.
- Keep the list. When you reuse the answer next month, the list tells you what you checked and when.
Step 5 is where the model is useful again: rewriting from vetted facts is a job it does well. The AI research workflows post shows a longer version of this loop, and why your AI results are mediocre covers the prompting side so you get fewer bad claims to check in the first place. If you're pasting sensitive text into these tools while you do this work, read what not to paste into ChatGPT first.



