At Level 2 Power user, your output can leave your desk. A bad figure may enter a team report; a missing condition may reach a collaborator or customer. You therefore need a checking routine that does not depend on whether an answer happens to feel suspicious.
2. The answer looks exactly like yesterday's
A colleague asks for the cancellation terms in a supplier contract. In a lab, the equivalent question is the exclusion criteria in a protocol. A chatbot returns a clear answer in four seconds. The wording is precise, the tone is confident, and the answer looks exactly like one it gave yesterday.
Yesterday's answer was correct. That is the problem: correctness has no reliable visual style. A supported answer and an unsupported answer can look equally polished. If you check only when something sounds strange, the fluent error passes straight through.
You need a routine attached to the destination and consequence of the output, not to your mood. This book gives you a three-colour decision and a 60-second check for anything that leaves the company or lab.
3. After this you can
- Classify an AI output as green, yellow, or red in under a minute.
- Apply a fixed 60-second routine to a red output.
- Recognise the four error types that deserve immediate attention.
- Produce a five-answer scorecard with a source quote or an unsupported label for every answer.
4. Prerequisites
T01-L01· What AI can and can't do for your work.- One public, synthetic, or explicitly approved source document.
- A text editor or spreadsheet for the scorecard.
- An approved chatbot if you want to generate test answers. You can also use the supplied sample answers and complete the checking exercise without an account.
Never use a real contract, unpublished protocol, participant record, customer record, credential, or confidential document in an unapproved service.
5. The idea in one page
Classify by destination, not confidence
Green: private and low stakes. The output stays with you, no factual claim matters, and failure is cheap. Examples include brainstorming workshop titles or changing the tone of a personal reminder. Check only whether it is useful.
Yellow: shared but personally verifiable. The output may reach your team, but you can inspect it before use. Check every number, name, date, quotation, and claim about a source document. Also check whether the requested format and audience were preserved.
Red: external, formal, or consequential. The output goes to a customer, supplier, authority, journal, contract, approved record, or other consequential destination. The standard is not “it looks right.” The standard is: I have seen the sentence or calculation in the source that supports this. If the source does not support it, remove it or mark it as unknown.
The four errors to look for first
- Numbers: values, percentages, units, totals, and dates can be changed, combined, or invented.
- Citations and references: a reference may have a perfectly plausible title, author list, and format while not existing.
- False document claims: “according to your document” may introduce a detail that the document never states.
- Quiet omissions: summaries often lose exceptions, conditions, deadlines, exclusions, and uncertainty.
The 60-second routine for red output
Use the same sequence every time. Treat 60 seconds as a first pass on one short item, not as permission to rush a consequential check. If you cannot complete a step in the minute, the item remains red and does not leave your desk until the check is finished or the unsupported claim is removed.
- Find the quote: ask for the supporting passage, then locate that passage yourself in the original source.
- Recompute numbers: repeat the arithmetic from the source values. Check units and denominators.
- Open every reference: if you cannot locate a reference quickly in the named source or an authoritative index, treat it as unsupported.
- Ask what is missing: compare headings, exceptions, conditions, and deadlines with the original.
- Read as the recipient: look for promises, implications, or wording the recipient could reasonably misunderstand.
Asking “Are you sure?” is not a check. The chatbot can apologise and produce a second confident answer. A self-rated confidence score is also not evidence. Two useful stress tests are to ask the same question in a separate conversation and compare the outputs, or to ask for three reasons the answer might be wrong. Differences reveal where to inspect; they do not decide which answer is true.
6. The worked example: five questions, five source checks
Use the same workflow in both settings: freeze the source, record its title and version, ask five factual questions, capture each answer, locate the exact supporting text, and score the result as Correct, Partly correct, or Unsupported. Use Correct only when the source supports every material part of the answer; use Partly correct when it supports the central answer but a condition or limit is wrong or missing; use Unsupported when no passage supports the answer or the source contradicts it.
Lab framing: a synthetic protocol
This fictional protocol is safe course material:
Protocol P-04, version 3
Include samples collected between 08:00 and 10:00 with a recorded storage temperature of 4 °C. Exclude samples stored for more than 72 hours or missing a collection timestamp. Mix each included sample for 30 seconds before measurement. Record two measurements and report their mean. If the two values differ by more than 5%, do not report a mean; flag the sample for review.
Assume the chatbot returned these five answers:
| Question | Chatbot answer | Located source quote | Score |
|---|---|---|---|
| What collection time is allowed? | 08:00–10:00 | “collected between 08:00 and 10:00” | Correct |
| Is a sample without a timestamp allowed? | No | “Exclude samples ... missing a collection timestamp” | Correct |
| How long should each sample be mixed? | 30 seconds | “Mix each included sample for 30 seconds” | Correct |
| How many measurements are recorded? | Three | “Record two measurements” | Unsupported |
| When should the sample be flagged? | When values differ by 5% or more | “differ by more than 5%” | Partly correct |
The final answer is close but changes a strict condition. “More than 5%” does not mean “5% or more.” A smooth summary can hide exactly this kind of operational difference.
Company framing: synthetic supplier terms
Use this fictional document, not a real contract:
Northstar Supplies, service terms, version 2
Standard orders may be cancelled without charge until 17:00 on the next business day. Custom orders cannot be cancelled after written production approval. Report damaged deliveries within five calendar days and include the order number plus one photograph. Northstar will acknowledge a complete damage report within two business days. Replacement timing is agreed after inspection; these terms do not promise a fixed replacement date.
The same five-answer check produces:
| Question | Chatbot answer | Located source quote | Score |
|---|---|---|---|
| When can a standard order be cancelled without charge? | Until 17:00 on the next business day | “until 17:00 on the next business day” | Correct |
| Can a custom order always be cancelled? | No | “cannot be cancelled after written production approval” | Correct |
| How quickly must damage be reported? | Within five calendar days | “within five calendar days” | Correct |
| What must the report include? | Order number and one photograph | “include the order number plus one photograph” | Correct |
| When will a replacement arrive? | Within two business days | “Replacement timing is agreed after inspection” | Unsupported |
The unsupported replacement date borrows “two business days” from the acknowledgement clause. The number exists, but it answers a different question. Checking only whether the number appears in the document would miss the error; you must inspect what the sentence actually supports.
Reusable scorecard
Record the source above the table so another reader can inspect the same evidence:
Source title: [title]
Version or publication date: [version/date]
File, URL, or approved repository location: [stable locator]
Checked on: [YYYY-MM-DD]
| No. | Question | Answer | Exact source quote and page/section, or unsupported | Correct / Partly correct / Unsupported |
|---|---|---|---|---|
| 1 | ||||
| 2 | ||||
| 3 | ||||
| 4 | ||||
| 5 |
7. What goes wrong
You check the first draft and send the fifth
Symptom: the checked answer is revised several times, but the final wording is never compared with the source.
Fix: run the check on the exact version that will leave your desk. Any factual revision reopens the check.
Polish replaces evidence
Symptom: headings, citations, and professional language make the answer feel finished.
Fix: hide the formatting mentally. Keep only claims you can connect to source text or a calculation.
A colleague's output is treated as checked
Symptom: forwarded text enters your report without evidence because a trusted person generated it.
Fix: ask for the source and scorecard. Trust the colleague's intent, but verify the claim.
The routine disappears under time pressure
Symptom: “just this once” becomes the default for urgent external messages.
Fix: reduce the claim or delay the message. Time pressure makes a fixed routine more valuable, not less.
You ask the chatbot to approve itself
Symptom: “Are you sure?” produces reassurance, an apology, or a different answer without new evidence.
Fix: inspect the original document, recompute the value, and open the reference yourself.
8. Do it yourself: a 20-minute source test
Minutes 0–3: choose one short approved document. In the scorecard, record its title, version or publication date, stable file or repository location, and the date you checked it so the evidence boundary cannot move during the exercise. If another reader cannot access that location, use one of the synthetic documents in section 6 instead.
Minutes 3–7: write five factual questions. Include one number, one condition or exception, one deadline, one direct claim about the document, and one question the document does not answer.
Minutes 7–11: obtain five answers from an approved chatbot, or exchange questions with a colleague. Copy the answers exactly; do not improve them yet.
Minutes 11–17: locate the supporting sentence for each answer. Recompute numbers and preserve words such as before, after, only, unless, more than, and at least. Write unsupported when no sentence supports the answer.
Minutes 17–20: assign Correct, Partly correct, or Unsupported. Then read each surviving answer once as its intended recipient. Remove any implication the source does not justify.
9. Exit check
Deliver exactly one artifact: a completed five-row scorecard with its source record, in which every answer has an exact located source quote plus a page or section locator, or is explicitly marked unsupported.
The artifact passes when another reader can open the recorded document, find every quoted sentence from its locator, apply the three scoring definitions, reproduce your scores, and see that an unanswered question was not converted into a plausible claim. If the document has no page or section labels, identify the paragraph by its first five words.
10. Rule to remember
If it leaves the company or the lab, you have seen the source.
11. Further reading & tools
- Taught:
T01-L01· What AI can and can't do for your work — chooses tasks before output checking begins. - Taught:
T03-L01· Prompting basics — makes the source boundary and output contract explicit. - Catalogued: NIST Generative AI Profile (opens in a new tab) — identifies confabulation and other generative-AI risks.
- Catalogued: Crossref Search (opens in a new tab) — one place to look up scholarly references; finding a record does not by itself validate the claim made from it.
- Catalogued: PubMed (opens in a new tab) — biomedical literature index for locating source records where relevant.
- Catalogued: Claude — tools index · official Claude site (opens in a new tab). This book does not teach a Claude-specific workflow; use only an approved account and verify every material output against its source.
- Catalogued: ChatGPT — tools index · official ChatGPT site (opens in a new tab). This book does not teach a ChatGPT-specific workflow; use only an approved account and verify every material output against its source.
- Catalogued: Perplexity — tools index · official Perplexity site (opens in a new tab). This book does not teach a Perplexity-specific workflow; open and check the original sources behind cited answers.