01With the original request pasted back in, the AI checked its work item by item

Take an example from September 2026. A little earlier, the operator had asked an AI to build a system that publishes articles automatically. The request was specific. Each day, write two articles based on the morning and evening AI news, plus one article on a question about life. Write all three in Japanese and in English, and publish them at 6 a.m.

Once the system was running, the operator pasted that same request back into the conversation, word for word. Then came one narrow question: measured against this request, do the article topics actually include AI?

The AI first gave a breakdown. Two of the three articles start from the AI news. The third is chosen from a list of life questions. Each is written in both languages, so the system produces six texts a day.

Then the AI reported something it had not been asked about. Drafting was scheduled for 3 a.m., but the day's news was not ready until 6:30 a.m. At 3 a.m. there was no news for that day yet. The request had asked for that day's morning and evening news; in practice, the system was leaning on the previous day's. The AI changed the timing.

The operator had asked about one thing only. Because the original request was in front of the AI, it compared every item in the request with the system it had built, and found a gap in a different item. The lesson for anyone using AI is simple. When the AI reports that the work is done, paste in the original request and have it answer, item by item, whether each part is met.

02The yardstick moves from the AI's summary back to the request itself

Ask "is it done?" after a task, and the AI answers against its own understanding of the request. That understanding is a summary the AI built while it worked. It is not the request, word for word.

Paste the original request back in and the yardstick changes. The AI now answers against what was written, not against its summary. Conditions such as "morning and evening", "6 a.m." and "both languages" each become something to check.

Figure 1 Asking again with the original request
Paste theoriginal requestSplit it intoitemscount, topics,languages, timeIs each itemmet?Fix the unmetitemuse that day's newsPaste the original requestSplit it into itemscount, topics, languages, timeIs each item met?Fix the unmet itemuse that day's news
Even a one-point question leads the AI to compare every item in the request with the output, once the request is in front of it.

What comes back is not a single yes or no. It is a list that links each item in the request to the part of the system that meets it. Any item with no link is, by definition, a place that needs fixing.

03Working and meeting the request are two separate checks

The timing gap would never show up in a test run. Drafting starts at 3 a.m. Articles go live at 6. No error appears. New articles arrive every morning. From the outside, the system works.

The only thing wrong is that the source material is a day old. The request wanted that day's news. The system wrote from the day before. That difference stays hidden unless someone holds the output up against the request.

Software quality practice gives these two checks different names. Verification asks whether the product was built to its specification. Validation asks whether it meets the needs of the people who use it. US medical device regulations define the two terms separately.

The previous episode had the AI split its "done" into what it had actually checked and what it had not. This episode goes one step further. It asks whether the part that was checked matches what the request asked for in the first place.

04Checking against the request means linking each item to the output, one by one

In this episode, asking whether the work meets the request means this: break the original request into items, and have the AI show where in the output each item is met. In software engineering, the ability to trace a requirement backwards to its source and forwards to where it was built and tested is called requirements traceability.

CheckYardstickWhat it findsWhat it misses
Run itDoes an error appear?Crashes and failuresParts that run but differ from the request
Split the "done" report (previous episode)Did the AI actually check it?Parts nobody checkedGaps from the request inside the checked parts
Check against the original request (this episode)Each item in the requestUnmet items, items read a different wayNeeds the request never stated

The check has limits. It finds gaps between the request and the output. It cannot find a condition the request never mentioned. That is found by a different check: the person looking at the output against their own purpose.

The second limit is that a single original request has to exist. If the request changed many times, in conversation or in fragments, there is no one text to check against. The method works when there is one written request to hold the output up to.

05In material creation and review, the brief and the review request are ready-made yardsticks

Pharmaceutical promotional work already has documents that play the role of the original request. Marketing writes a brief for each piece. Medical affairs receives a request to prepare content. A review request goes in with each submission. Each states the audience, the key message, the evidence to cite and the limits on what may be claimed.

When an AI drafts a piece of material, reading only the draft draws attention to how well it is written. Paste the brief back in and have the AI check each item, and different gaps appear. A piece meant for healthcare professionals may be phrased for the general public. A figure may come from a source other than the trial the brief named. The more readable the draft, the easier such gaps are to miss.

Document that serves as the requestItems to checkGaps the check can catch
Material briefAudience, key message, evidence to citeWording wrong for the audience, sources not named in the brief
Review requestScope of review, points to look at closelyRequested points the review never addressed
Internal procedureRequired steps, record formatSkipped steps, wrong format

For a reviewer, the gain is that every comment points back to the request. Instead of "this wording bothers me", the comment becomes "this does not match item 3 of the brief". The person who receives the comment knows exactly what to change to meet the request.

Having the AI run this check is a first pass. The reviewer reads the list and makes the final judgment.

06Over a long task, the original request drifts out of the AI's reach

Why would the AI that built the system drift from the request? The first reason is length. While building, the AI reads files, runs commands and receives their output. All of it piles into the AI's input, and the original request ends up buried near the start of that pile.

Anthropic's guide to Claude Code says performance degrades as the context window fills, and that Claude may start forgetting earlier instructions or making more mistakes. A 2023 study found that language models perform much worse when the information they need sits in the middle of a long input than when it sits at the beginning or the end.

The second reason is the number of conditions. This request packed five of them into one sentence: how many articles, which topics, which languages, what time to publish, and which day's news. A study that measures how well language models follow instructions built 25 types of conditions that a machine can check, such as word counts or required keywords, and tested about 500 prompts. Measuring whether an instruction was followed starts with making each condition checkable on its own.

The third reason is that "it runs" easily passes for "it is done". The same guide notes that without a way to check its work, the AI stops when the work looks done. A gap that raises no error slips straight past that signal.

Figure 2 Three reasons an AI drifts from the request
Output drifts from therequestLong taskrequest buried in the inputMany conditionsfive in one sentenceRunning looks like donegaps that raise no errorOutput drifts from the requestLong taskrequest buried in the inputMany conditionsfive in one sentenceRunning looks like donegaps that raise no error
All three are addressed by putting the original request back at the end of the input and checking it item by item.

Pasting the original request back in puts it at the end of the input, the position a language model uses best. Each condition becomes a separate check, and there is now a yardstick other than whether the system runs. One step addresses all three reasons.

07When the work is reported done, paste the request and ask for an item-by-item table

There are four steps.

  1. Keep the request. Save the first request in its original words. Add any later conditions to the same place.
  2. When the AI reports the work done, paste the request back in. Do not summarise it. Give the text exactly as written.
  3. Ask for an item-by-item table. Write: "Split this request into items. For each, show where the output meets it, and list any item it does not meet."
  4. Fix only the unmet items. After the fix, have the AI produce the same table again.

Adding one specific concern makes the check easier to start. The operator's question was about a single point, whether AI was among the topics. Even so, with the original request in front of it, the AI reviewed the other items as well.

The check can also go to a different reviewer from the AI that did the work. Anthropic's guide shows how to have a subagent with a fresh context review the changes against a plan, checking that every requirement is implemented and that nothing outside the task's scope changed. That keeps the builder's assumptions out of the check.

Figure 3 Four steps from tomorrow
Keep the requestPaste it backwhen doneno summaryGet anitem-by-item…Fix only unmetitemsthen rebuild thetableKeep the requestPaste it back when doneno summaryGet an item-by-itemtableFix only unmet itemsthen rebuild the table
The check can go to a reviewer other than the AI that built the work. The final judgment stays with the person.
Key Points ── 3 to take away
  1. When the AI reports the work done, paste the original request back in, word for word, and have it answer item by item whether each part is met. The yardstick moves from the AI's summary to what was written.
  2. Running and meeting the request are different things. A gap that raises no error only shows when the output is held up against the request.
  3. Over a long task, the original request gets buried deep in the AI's input. Pasting it back puts it at the end and turns each condition into a separate check.
Closing

When you hand work to an AI, the last check is not only whether the system runs. It is whether the system does what you first asked, in the words you first used. That yardstick is the request in your hands, not the AI's memory of it. Keep the request, and paste it back each time the work is reported done. That alone brings out gaps that cannot be seen from the outside.

Sources & references
  1. Anthropic. Best practices for Claude Code. Claude Code Docs. https://code.claude.com/docs/en/best-practices
  2. Liu, N. F., Lin, K., Hewitt, J., et al. Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172, 2023. https://arxiv.org/abs/2307.03172
  3. Zhou, J., Lu, T., Mishra, S., et al. Instruction-Following Evaluation for Large Language Models. arXiv:2311.07911, 2023. https://arxiv.org/abs/2311.07911
  4. Wikipedia. Verification and validation. https://en.wikipedia.org/wiki/Verification_and_validation
  5. Wikipedia. Requirements traceability. https://en.wikipedia.org/wiki/Requirements_traceability
What this episode is based on The operator's own record of requests to and decisions with Claude since February 2026 (the operator has used generative AI since March 2023), anonymised and generalised into a pattern. No messages are quoted. The sources listed are public material used to check the background of the pattern.