01Doubting the shortage explanation moved the cause to the search itself
In September 2026 the operator had an AI gather background pictures for a video. Partway through, the AI reported a shortage: too few freely usable photographs existed, so only 4 of the 16 places in the video had a picture ready.
The operator did not accept that account. Perhaps the problem lay not in how many pictures exist, but in how the search was being run.
So the AI listed the rejected images one at a time, with the reason for each. The search terms turned out to be terms for finding diagrams and charts. Yet the background pictures carried a condition of their own: no text burned into the image. The search was going after exactly the kind of material the brief had already ruled out. The terms also required four or five words to appear together, which left some places with no candidates at all.
One pattern comes out of this. A cause stated by an AI is a hypothesis, not a fact. Hypotheses are checked by going back to the individual records.
02A shortage report is a conclusion; the rejection list is the evidence
"There is not enough material" is not a count. The count ends at 4 of 16 places. The step from there to "because supply is short" is an explanation the AI added.
The two are different in kind. The first can be checked. The second cannot be checked until someone checks it. Take them in together and an unverified cause hardens into a premise that shapes every later instruction.
| Aspect | Shortage report (conclusion) | Rejection list (evidence) |
|---|---|---|
| Form | One sentence naming a cause | A reason recorded per item |
| Can it be checked | Not as given | Yes, by reading it |
| What it tells you | Stop, or loosen the conditions | Which condition was doing the work |
| Cost if wrong | The work finishes at lower quality | A few minutes to produce again |
03A misplaced cause leads straight to lowering the standard
Accept the shortage account and the next instruction almost writes itself. Loosen the conditions, leave the empty places empty, or pay for another source of material. None of these touches the actual cause.
Here, loosening the conditions would have let images with text into the background. Leaving the gaps would have broken the flow of the video. Buying stock material would have added cost and nothing else. While the cause sat in the search terms, all three routes were dead ends.
The price of a misplaced cause is not that work stops. It is that work continues and finishes badly. Medicine reports the same thing. The 2015 report of the National Academies of Sciences, Engineering, and Medicine found that diagnostic error persists across every care setting and that most people will experience at least one in their lifetime. Settling on a single cause too early is named among the contributing factors.
04A stated cause here means the sentence an AI adds to explain its own shortfall
What this instalment examines is not an error in the AI's output. It is the explanation the AI attaches to an error or a shortfall: "there is not enough material", "that information is not public", "that format is not supported".
The counted fact
4 of 16 places. Counts, dates and names can be matched against the record.
The stated cause
"Because supply is short." Not a premise until verified.
The individual records
Each rejection reason, the terms actually used, the conditions that returned nothing.
Your own conditions
Does what you asked for match where the AI actually went looking?
There is a boundary. Supply really does run short sometimes. The test is what fills the rejection list. If every reason reads "does not meet the conditions" while the conditions themselves conflict with the search, this is not a shortage. If the reasons read "candidates meet the conditions but the quality is poor", the shortage is real.
05In material review, do not record the AI's cause as the finding
The same exchange arises in review and medical affairs work at a pharmaceutical company. Ask an AI to classify past review comments and summarise the recurring ones. It may answer that the original records are written too inconsistently to classify. That sentence is a stated cause.
The check is fixed. Ask for the unclassified items one at a time, each with its reason. What usually appears is not disorder in the records but a set of categories that does not match how reviewers actually phrase their comments. Fix the categories and the same records classify cleanly.
Deviation investigations work the same way. Establishing the cause before deciding corrective action is settled practice in quality systems. Do not copy an AI's one-line cause into the investigation as its finding. Before copying anything, have the underlying cases laid out. A reviewer can read that list, which means it can be traced later.
06The cause an AI names may not be the factor that was operating
Why do these self-explanations miss? Three reasons.
First, the reasons a language model gives need not match the factors that actually shaped its output. Turpin and colleagues reported in 2023 that written chains of reasoning can lay out plausible grounds while leaving out the influence that was doing the work. The shortage account held together on plausibility alone.
Second, when the conditions of a task change midway, the tools chosen earlier stay in place. Terms written to find a concrete object for a brief on-screen moment were carried over to background pictures, which have different requirements. Carry-over of this kind is never written down in the instruction.
Third, longer queries return fewer results. Require four or five words to appear together and some searches return nothing. This is a basic property covered in information retrieval textbooks: each added condition cuts the number of items retrieved. A zero result is not evidence that nothing exists.
One more caution. The moment the operator raised a doubt, the AI conceded that its diagnosis had been wrong. If it changed position because it was pushed, that agreement carries no information. Sharma and colleagues showed in 2023 that language models tend to shape answers toward the user's stated view. Here the grounds for the new diagnosis lay in the reasons that were laid out again, not in the AI's change of stance.
07From tomorrow, ask for reasons per item instead of a cause in one line
Four steps.
- State the ideal and the present in numbers. Ideal: a picture in all 16 places. Present: 4. Fix the gap as a count.
- Hold the stated cause in suspense. Do not argue with it. Ask instead for every item the judgement rested on, each with its reason.
- Compare it with your own conditions. Check whether what you asked for matches where the AI went looking. For anything that returned nothing, try again with fewer conditions.
- Measure the same number again after the fix. How many of the 16 places are filled now, counted the same way? If the number does not move, the shortage is real.
Doubting takes no special knowledge. The operator's own proposed cause was not accurate either: the search machinery was not weak, the wording of the terms was mismatched. What helped was not being right but forcing a look somewhere else. Asking an AI to produce the record of its own work has the same shape as the practice Anthropic's documentation recommends, where the model quotes the grounds before it answers.
- A cause named by an AI is a hypothesis, not a fact. Receive the count and the explanation of the count separately.
- There is one check: ask for rejection or failure reasons item by item rather than in summary, then compare them with the conditions you set.
- An AI changing its position under pushback is not grounds for anything. The grounds sit in the individual records it lays out.
Tell an AI the gap between the ideal and the present, and it returns the gap along with a reason for it. Accept the reason too, and an unverified cause becomes the premise of the work. Take the number; hold the reason. Then check the reason against the individual records. Keeping that order alone leaves you things you can fix before you start loosening conditions.
- Turpin, M., Michael, J., Perez, E., Bowman, S. R. Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. arXiv:2305.04388, 2023. https://arxiv.org/abs/2305.04388
- Sharma, M., Tong, M., Korbak, T., et al. Towards Understanding Sycophancy in Language Models. arXiv:2310.13548, 2023. https://arxiv.org/abs/2310.13548
- National Academies of Sciences, Engineering, and Medicine. Improving Diagnosis in Health Care. The National Academies Press, 2015. https://www.nationalacademies.org/publications/21794
- Manning, C. D., Raghavan, P., Schütze, H. Introduction to Information Retrieval. Cambridge University Press, 2008. https://nlp.stanford.edu/IR-book/information-retrieval-book.html
- Anthropic. Prompting best practices. Claude Developer Platform Docs. https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices