01One follow-up question split the AI's report into checked and unchecked

In September 2026 the operator noticed that the "latest 10" news list at the top of the site had stopped updating. An AI was asked to find out why. The job that rebuilds the list had never been added to the morning schedule. The AI added it to the schedule and rebuilt the list up to the latest date.

The operator then asked one short question: so does it work now? The AI did not stop at "yes". It split its answer into three things it had actually run and checked, and one thing it had not yet checked.

The three checked items were these. The list now ran up to the latest date. The rebuild job worked correctly under the same conditions as the scheduled run. And the change to the schedule settings was exactly the two intended lines. The one unchecked item was whether, the next morning at the set time, the whole chain would run end to end, from generating the news to rebuilding the list.

This episode takes one pattern from that exchange. When an AI says the work is done, have it separate what it checked from what it did not. Then decide when and how the unchecked part will be checked.

02"Done" mixes checked facts with expectations

A completion report usually arrives as a single word. Its contents are not single. Some of it was run and observed. Some of it was built but never run. Some of it should work, given how the system is set up. All three sit inside the same "done".

Sorted by how each item was checked, the report looks like this.

Item in the reportHow it was checkedCategory
The list runs to the latest dateLooked at the rebuilt listChecked
The rebuild job worksRan it under the same conditions as the scheduleChecked
Only the two intended lines changedCompared the settings before and afterChecked
Tomorrow's run goes end to endNot yet run; it should work as configuredUnchecked

Laid out this way, the remaining risk shrinks to one row. That row is also the only thing the operator needs to look at the next morning.

Figure 1 Splitting one completion report three ways
"Done"a one-word reportCheckedthree items run and seenUncheckeddoes tomorrow's run gothroughReason notchecked"Done"a one-word reportCheckedthree items run and seenUncheckeddoes tomorrow's run go throughReason not checked
Once split, the only thing a person must verify personally is the single unchecked item.

03Taking an expectation as finished work means the gap shows up in production

If you accept "done" without splitting it, the unchecked item lands on the finished side. Close the task there, and if the run fails the next morning, nobody goes to look. The list freezes again, and the first people to notice are readers looking at stale news.

The original fault had exactly this shape. Run by hand, the rebuild job worked. It simply was not in the automatic sequence. "It runs" and "it runs every day" are separate claims, and each needs its own check.

A split report has another advantage: it saves effort. For the three checked items, reading the evidence the AI provides is enough. The person only needs to verify the one unchecked item personally.

04Checked means run and observed; built is not enough

The pattern only works if the split is consistent, so it helps to fix the boundary. Here, "checked" means the item was actually run or opened, and the result was seen.

1

Run and observed

Goes under checked

The job was run and its output or diff was seen. The result itself can be attached.

2

Built or written

Goes under unchecked

Settings or steps were changed but not run. Writing something is not evidence that it works.

3

Should work

Goes under unchecked

The design says it will run. No result exists until the time or outside conditions arrive.

4

Deliberately not run

Unchecked, with a reason

Running it would affect something outside, so the check was skipped. The reason goes in the report.

The fourth case matters. Here the AI judged that news generation was a daily job calling an outside service, and chose not to run it just to test it. Sometimes not checking is the right call. Even then, a report that says what is still unchecked, and why, lets the person decide what to do next.

By contrast, the AI's remark that it considered the risk low does not count as checked. An assessment is material for a decision. It is not a result.

05In material review, AI reports on cross-checks and summaries need the same split

In pharmaceutical material review and medical affairs, reports on work handed to AI also arrive as "done". The figures in a promotional piece were compared with the package insert. The key points of a cited paper were summarised. Reviewer comments were applied. In each case, the single word does not say how far anything was checked.

A report that "all figures match" may mix items where both documents were opened and compared with items judged to match from context. A summary of a cited paper may mix parts written from the full text with parts written from the abstract alone. Splitting the report shows exactly where a person needs to go back to the source.

The idea is familiar from how drug records are handled. The U.S. Food and Drug Administration's 2018 guidance on data integrity expects data to be attributable to the person who recorded it, recorded at the time of the activity, and accurate. The same guidance treats review by a second person as an action that must be attributable to a specific individual. Splitting an AI's report into checked and unchecked points the same way: a record of who checked what.

06AI stops when the work looks done, and leans toward stating rather than hedging

There are mechanical reasons an AI may not split its report on its own. Anthropic's guidance for Claude Code says Claude stops when the work looks done. Without a check it can run, looking done is the only signal it has.

The second reason is a lean toward confident statements. A 2025 paper argues that one reason language models state false things is that training and evaluation have rewarded guessing over admitting uncertainty. A 2024 document from the U.S. National Institute of Standards and Technology lists confidently stated but false content as one of the risks of generative AI.

At the same time, models can estimate how sure they are. A 2022 study reported that large language models can predict, fairly well, the probability that their own answers are correct. So the ability to tell certain from uncertain is there. Unless the report is asked to show it, it simply does not surface.

Figure 2 How an expectation slips in as done
WorkfinishesIs therea check…It looks donethe onlysignalStated withconfidencelean towardguessingUncheckedcounted as…Work finishesIs there a check torun?It looks donethe only signalStated with confidencelean toward guessingUnchecked counted as done
With no check to run and no request to split the report, expectations arrive on the finished side.

07From tomorrow: set three sections for the report, and give each unchecked item a date

This time one follow-up question was enough to get a split report. Rather than relying on that, set the format from the start. There are four steps.

  1. Set the report format first. When you hand over the task, say that the report must have three parts: checked, unchecked, and reasons for not checking.
  2. Ask for evidence under checked. The output of the run, the diff, the passage in the document that was opened: something a person can see.
  3. Ask for a method and a time under unchecked. One line each: what would count as confirmation, and when it can happen.
  4. Close only when unchecked is empty. Look at the result on the due date and move the item to checked. If it cannot move, ask for a fix on the spot.
SectionWhat the AI writesWhat the person does
CheckedWhat was run, what was seen, the result itselfReads the attached result
UncheckedWhat has not been seen; how and when to check itLooks at the result on the due date
Reasons for not checkingWhat running it would affectDecides whether to run it

A follow-up question also works. Ask "what did you actually check, and what is still unchecked?" and an AI will usually answer in the same shape as here. What matters is clearing the unchecked rows before the task is closed.

Figure 3 Four steps before closing a task
Set thereport…checked,unchecked,…Evidence forcheckedMethod andtime for…Look onthe due…Close whenunchecked is…Set the report formatchecked, unchecked, reasonsEvidence for checkedMethod and time for uncheckedLook on the due dateClose when unchecked is empty
The condition for closing is an empty unchecked section, not the AI saying done.
Key Points ── 3 to take away
  1. An AI's "done" mixes facts it checked by running things with expectations that things should work. Have completion reports split into the two.
  2. Checked means run or opened, with the result seen. Items that were only built, or should work by design, go under unchecked, along with any reason a check was skipped.
  3. Give every unchecked item a method and a time. Close the task only when that section is empty.
Closing

There is no need to distrust an AI's report. The point is not to take a one-word "done" as a single fact. Split into checked and unchecked, the part a person needs to look at shrinks to a line or two. The gap between the ideal and the current state is sitting in those unchecked lines.

Sources & references
  1. Anthropic. Best practices for Claude Code. Claude Code Docs. https://code.claude.com/docs/en/best-practices
  2. Kalai, A. T., Nachum, O., Vempala, S. S., Zhang, E. Why Language Models Hallucinate. arXiv:2509.04664, 2025. https://arxiv.org/abs/2509.04664
  3. Kadavath, S., et al. Language Models (Mostly) Know What They Know. arXiv:2207.05221, 2022. https://arxiv.org/abs/2207.05221
  4. National Institute of Standards and Technology. AI Risk Management Framework: Generative AI Profile (NIST AI 600-1). 2024. https://doi.org/10.6028/NIST.AI.600-1
  5. U.S. Food and Drug Administration. Data Integrity and Compliance With Drug CGMP: Questions and Answers. Guidance for Industry, 2018. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/data-integrity-and-compliance-drug-cgmp-questions-and-answers
What this episode is based on The operator's own record of requests to and decisions with Claude since February 2026 (the operator has used generative AI since March 2023), anonymised and generalised into a pattern. No messages are quoted. The sources listed are public material used to check the background of the pattern.