01One follow-up question split the AI's report into checked and unchecked
In September 2026 the operator noticed that the "latest 10" news list at the top of the site had stopped updating. An AI was asked to find out why. The job that rebuilds the list had never been added to the morning schedule. The AI added it to the schedule and rebuilt the list up to the latest date.
The operator then asked one short question: so does it work now? The AI did not stop at "yes". It split its answer into three things it had actually run and checked, and one thing it had not yet checked.
The three checked items were these. The list now ran up to the latest date. The rebuild job worked correctly under the same conditions as the scheduled run. And the change to the schedule settings was exactly the two intended lines. The one unchecked item was whether, the next morning at the set time, the whole chain would run end to end, from generating the news to rebuilding the list.
This episode takes one pattern from that exchange. When an AI says the work is done, have it separate what it checked from what it did not. Then decide when and how the unchecked part will be checked.
02"Done" mixes checked facts with expectations
A completion report usually arrives as a single word. Its contents are not single. Some of it was run and observed. Some of it was built but never run. Some of it should work, given how the system is set up. All three sit inside the same "done".
Sorted by how each item was checked, the report looks like this.
| Item in the report | How it was checked | Category |
|---|---|---|
| The list runs to the latest date | Looked at the rebuilt list | Checked |
| The rebuild job works | Ran it under the same conditions as the schedule | Checked |
| Only the two intended lines changed | Compared the settings before and after | Checked |
| Tomorrow's run goes end to end | Not yet run; it should work as configured | Unchecked |
Laid out this way, the remaining risk shrinks to one row. That row is also the only thing the operator needs to look at the next morning.
03Taking an expectation as finished work means the gap shows up in production
If you accept "done" without splitting it, the unchecked item lands on the finished side. Close the task there, and if the run fails the next morning, nobody goes to look. The list freezes again, and the first people to notice are readers looking at stale news.
The original fault had exactly this shape. Run by hand, the rebuild job worked. It simply was not in the automatic sequence. "It runs" and "it runs every day" are separate claims, and each needs its own check.
A split report has another advantage: it saves effort. For the three checked items, reading the evidence the AI provides is enough. The person only needs to verify the one unchecked item personally.
04Checked means run and observed; built is not enough
The pattern only works if the split is consistent, so it helps to fix the boundary. Here, "checked" means the item was actually run or opened, and the result was seen.
Run and observed
The job was run and its output or diff was seen. The result itself can be attached.
Built or written
Settings or steps were changed but not run. Writing something is not evidence that it works.
Should work
The design says it will run. No result exists until the time or outside conditions arrive.
Deliberately not run
Running it would affect something outside, so the check was skipped. The reason goes in the report.
The fourth case matters. Here the AI judged that news generation was a daily job calling an outside service, and chose not to run it just to test it. Sometimes not checking is the right call. Even then, a report that says what is still unchecked, and why, lets the person decide what to do next.
By contrast, the AI's remark that it considered the risk low does not count as checked. An assessment is material for a decision. It is not a result.
05In material review, AI reports on cross-checks and summaries need the same split
In pharmaceutical material review and medical affairs, reports on work handed to AI also arrive as "done". The figures in a promotional piece were compared with the package insert. The key points of a cited paper were summarised. Reviewer comments were applied. In each case, the single word does not say how far anything was checked.
A report that "all figures match" may mix items where both documents were opened and compared with items judged to match from context. A summary of a cited paper may mix parts written from the full text with parts written from the abstract alone. Splitting the report shows exactly where a person needs to go back to the source.
The idea is familiar from how drug records are handled. The U.S. Food and Drug Administration's 2018 guidance on data integrity expects data to be attributable to the person who recorded it, recorded at the time of the activity, and accurate. The same guidance treats review by a second person as an action that must be attributable to a specific individual. Splitting an AI's report into checked and unchecked points the same way: a record of who checked what.
06AI stops when the work looks done, and leans toward stating rather than hedging
There are mechanical reasons an AI may not split its report on its own. Anthropic's guidance for Claude Code says Claude stops when the work looks done. Without a check it can run, looking done is the only signal it has.
The second reason is a lean toward confident statements. A 2025 paper argues that one reason language models state false things is that training and evaluation have rewarded guessing over admitting uncertainty. A 2024 document from the U.S. National Institute of Standards and Technology lists confidently stated but false content as one of the risks of generative AI.
At the same time, models can estimate how sure they are. A 2022 study reported that large language models can predict, fairly well, the probability that their own answers are correct. So the ability to tell certain from uncertain is there. Unless the report is asked to show it, it simply does not surface.
07From tomorrow: set three sections for the report, and give each unchecked item a date
This time one follow-up question was enough to get a split report. Rather than relying on that, set the format from the start. There are four steps.
- Set the report format first. When you hand over the task, say that the report must have three parts: checked, unchecked, and reasons for not checking.
- Ask for evidence under checked. The output of the run, the diff, the passage in the document that was opened: something a person can see.
- Ask for a method and a time under unchecked. One line each: what would count as confirmation, and when it can happen.
- Close only when unchecked is empty. Look at the result on the due date and move the item to checked. If it cannot move, ask for a fix on the spot.
| Section | What the AI writes | What the person does |
|---|---|---|
| Checked | What was run, what was seen, the result itself | Reads the attached result |
| Unchecked | What has not been seen; how and when to check it | Looks at the result on the due date |
| Reasons for not checking | What running it would affect | Decides whether to run it |
A follow-up question also works. Ask "what did you actually check, and what is still unchecked?" and an AI will usually answer in the same shape as here. What matters is clearing the unchecked rows before the task is closed.
- An AI's "done" mixes facts it checked by running things with expectations that things should work. Have completion reports split into the two.
- Checked means run or opened, with the result seen. Items that were only built, or should work by design, go under unchecked, along with any reason a check was skipped.
- Give every unchecked item a method and a time. Close the task only when that section is empty.
There is no need to distrust an AI's report. The point is not to take a one-word "done" as a single fact. Split into checked and unchecked, the part a person needs to look at shrinks to a line or two. The gap between the ideal and the current state is sitting in those unchecked lines.
- Anthropic. Best practices for Claude Code. Claude Code Docs. https://code.claude.com/docs/en/best-practices
- Kalai, A. T., Nachum, O., Vempala, S. S., Zhang, E. Why Language Models Hallucinate. arXiv:2509.04664, 2025. https://arxiv.org/abs/2509.04664
- Kadavath, S., et al. Language Models (Mostly) Know What They Know. arXiv:2207.05221, 2022. https://arxiv.org/abs/2207.05221
- National Institute of Standards and Technology. AI Risk Management Framework: Generative AI Profile (NIST AI 600-1). 2024. https://doi.org/10.6028/NIST.AI.600-1
- U.S. Food and Drug Administration. Data Integrity and Compliance With Drug CGMP: Questions and Answers. Guidance for Industry, 2018. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/data-integrity-and-compliance-drug-cgmp-questions-and-answers