01First, what this system does
Drug companies make booklets and advertisements to explain their drugs. We call these materials. A material must not say anything beyond what the government has approved. The approved scope covers which disease the drug treats, at what dose, and with what warnings. We call this the approved facts. The approved facts are written in the package insert. The package insert is the official leaflet the government has checked for each drug.
A reviewer reads dozens of pages and looks for wording that goes beyond the approved facts. Most of the reviewer's time goes into this searching. So we let the AI help with the searching. The AI finds suspicious spots, attaches the grounds for suspicion, and hands them to a human. The human decides.
The division of labor can be said in one sentence. "What the rules state, a program checks directly. What they do not state, the AI judges by reading the context. Every flag can be traced back to the rule's original text." That division is the foundation.
02Three problems with reviewing a material one element at a time
Reviewing a material one sentence or chart at a time has three problems. This system was designed to answer those three.
First, flags overlap. Our policy is to flag generously to avoid misses, so we accept a large number. The trouble is the overlap. For example, "the efficacy chart is large," "the side-effect warning is small," and "the two sit side by side" come out as three separate flags. But what stays in the reader's head is one wrong picture: "this drug works, and I need not worry about side effects." The human ends up reading the same thing three times.
Second, the ways of reading get mixed together. There are four ways of reading a material. Read one sentence alone. Read two neighboring items in a row. Read the booklet from start to finish. Look at one chart cut out on its own. If these four are bundled into one question for the AI, what should be kept apart gets mixed inside the instructions to the AI. For the same reason, the cast of characters in a check called "the public eye" will drift from document to document unless it is fixed in advance.
Third, a way to count misses is needed. Declaring that "a miss is the gravest failure" does not tell you whether anything was actually missed. Neither does a plan to "run alongside humans before going live and compare." A priority that is not counted is not a priority.
03Updating the thinking ── review the picture in the head, not the material
The foundation is one idea. The side that makes the material knows more than the side that reads it. If the reader misunderstands, the reader is the one harmed. So a material must not leave a wrong picture in the reader's mind.
What matters is what gets reviewed. We do not check whether the material's sentences and charts go beyond the approved facts. We check whether the picture left in the head of someone who read the material goes beyond the approved facts. From here on we call this picture in the head the image.
What is an image? It is the small bundle of statements that stays in your head after reading. Two kinds of statement go in. One is a claim written plainly in the material. The other is a conclusion the reader naturally draws from how things are arranged and what is left out, even though it is not written. "The efficacy chart is large" is an observation about the material. "This drug works well" is an image. What the law forbids is excess on the side of the image. In the law's own words, that is "exaggerated" or "misleading."
Reviewing the image solves the second problem. The four ways of reading are no longer four separate checks. They are four conditions for building the same image. One person who read a single sentence, one who read two neighbors in a row, one who read the whole booklet, one who saw only the chart. Four people read the same material, four images are built, and each is compared with the approved facts. Only one flag is attached to each difference between an image and the approved facts. Which sentences or charts produced the difference is attached as a path. The human now reads one wrong image once.
Here we make the division of roles between the AI and the program clear. Building the image is the AI's job. But taking the difference between the image and the approved facts, and labeling that difference "heavy" or "light," is done by the program under fixed rules. Suppose the difference is "it says the drug works for a disease that is not approved." If that is written plainly in the material, it gets the heaviest label. If it is only hinted, it gets the second label. Rules decide this. The AI describes the image; the program judges it. The division "the program decides, the AI reads the meaning" works at the level of the image.
04The second foundation ── a miss that is not counted is not a priority
There are three devices to prevent misses. Keep low-confidence suspicions instead of discarding them. Ask the AI the same question three times, and adopt "problem" if it comes up even once. Always record the reason when a flag is removed. All three are right.
But these are efforts not to miss. They are not confirmation that nothing was missed. Effort and confirmation are different things. Checking the locks every morning is effort. Counting how many times you were burgled is confirmation. Without counting, you cannot tell whether the habit works.
So we add a second foundation. Misses are counted per material as the difference from the human's final verdict, totaled every week, and if they cross a set line, live operation stops. For this, we build a path that always returns the human reviewer's verdict, "fix" or "no fix," and the reason, to the AI side. Among the verdicts that come back, the ones the AI stayed silent on are the misses. Misses are recorded in a ledger. A ledger is a record book you can look back through later. In the ledger we add which check item should have fired and which way of reading should have built the image.
With this, the go-live rule, "zero misses for four weeks running," becomes a measured number rather than a wish. The AI not only searches before human review. It also learns after it.
05Seven principles ── six carried over, one added
There are seven principles. Each is built to answer one common objection. The table below shows which objection each principle answers and how it works in this system.
| Principle | Objection it answers | Meaning in this system |
|---|---|---|
| ① The standard is the approved facts | By what standard do you call something correct? | The standard is three documents: the official leaflet (package insert), the government's review report, and the safety-measures plan. The image is always compared with these three |
| ② Even facts must not mislead | If what is written is true, what is wrong? | Even a list of true facts is a problem if it forms a wrong image. We build the image under four ways of reading to check |
| ③ Do not pick only the good news | Surely the company may choose what to include? | Big on benefits, small on side effects: that kind of choice shows up in the image. If the safety image is lighter than the approved facts, that is a difference |
| ④ No flag without grounds | Who guarantees that verdict is correct? | Every flag must carry the rule's original text and which difference in which image. The program does not accept a flag without them |
| ⑤ A miss is the gravest failure | AI makes mistakes. Which mistake do you choose? | Add confirmation to effort. Return the human verdict, count the misses, stop if they cross the line (section 04) |
| ⑥ What the rules leave out is not free | Surely what is not written is free? | A material that hits no rule is not marked "no problem." It is reread through the public eye, and we record whose eyes read it |
| ⑦ Same judgment, same verdict | It passed last month. Why does it fail this month? | The same image, the same version of the rules, and the same human judgment give the same verdict, however much time passes. When a verdict changes, we show which changed: the rules or a past judgment |
The seventh, "same judgment, same verdict," is the first principle to be tested once real use begins. Review does not end after one round. Material for the same drug is revised. Material for another drug is built the same way. The reviewer changes. If the verdict wavers each time, the AI cuts the time spent searching but adds no reason to trust it. So we keep the human verdicts in a notebook of past judgments. When a similar image appears, that notebook is attached to the flag.
06Judgment in three tiers. The AI is a part; the words of the verdict are the constitution
Important verdicts go through a step in which 60 virtual reviewers debate and vote. But if calling that step is optional, verdicts that did not call it leave no record of debate. So judgment is fixed at three tiers, and rules decide how far each flag goes.
One more thing to settle: swapping the AI. AI products are replaced by newer ones every few months. When we swap, is the review still the same review? The answer is this. The words used for verdicts, the shape of a flag, and the gates of the checks are the constitution. The AI is a part. The verdict words are fixed at five levels, from "heaviest" to "no problem," and whatever the AI says, no verdict outside them exists. Before a new AI is used, we run it on a bundle of materials whose answers are known (an exam with an answer key). We confirm that the same images and differences come out as before.
07Turning the thinking into a mechanism that cannot be broken ── eight gates
Anyone can write down a way of thinking. What matters is whether there is force to stop it when the thinking is broken at run time. There are eight gates. A gate is a check you cannot pass until its condition is met.
- The entry gate ── Do not start unless the material and the reference documents are all present. Anything missing becomes the first flag, "check with a human." Never proceed in silence.
- The grounds gate ── A flag without the rule's original text, and without which difference in which image, is discarded the moment it is made.
- The removal gate ── A flag cannot be removed without a reason and a record.
- The coverage gate ── Every run confirms that images were built under all four ways of reading and compared with every item of the approved facts.
- The reproduction gate ── Every run confirms that the program's checks give the same answer for the same input.
- The independence gate ── Questions to the AI are asked three separate times, and "problem" is adopted if it appears even once. The cast of the public eye is managed by version.
- The continuity gate ── When a rule is revised, update the check items, confirm with the exam, and re-examine material already passed.
- The follow-up gate ── A review whose human verdict has not come back is not treated as finished. The difference from the returned verdict is counted. If the weekly miss rate crosses the set line, live operation stops.
Thinking, mechanism and check correspond one to one. That is the backbone of this design. The "image" becomes a part that builds images and a part that takes differences. "Counting" becomes the ledger of misses and the eighth gate. "Same verdict" becomes the notebook of past judgments and the exam at swap time. Which principle becomes which part, and when it works within one review, is told in the next two articles. The conclusion is short. This system was built to stop the image that tries to escape the approved facts, and built to count what it has missed.