01The whole picture ── seven columns and the learning loop
The steps fall into seven columns: input, rule checks, picture and difference, the public eye, merging, human, and learning. Steps 1 to 11 are the machine, 12 and 13 are the human and the panel, and 14 and 15 are the learning plane. Arrows show that the output of one step is the input of the next. The dashed loop shows that check items corrected at step 15 return to steps 4 and 5 of the next review.
The order matters. Checks with an unhesitating answer come first; checks with heavy judgment come later. The rule checks run before the picture is made and hand over their results as cues. The public eye runs after the difference is taken and reads only what no written rule covered. The exception desk sits before flags are combined, so that a drop is recorded before merging. The human verdict sits before the panel. The panel makes candidates to help the human's judgment; it is not a step that keeps the human waiting.
02Steps 1 to 3 Fixing the input ── facts in a fixed form (entry gate)
- Receive ── The material's PDF and the three documents to compare against (the official leaflet, the government's review report, and the safety-measures plan) are received. The entry gate works here. If the documents are not in the right form, the review does not start. Anything missing becomes the first flag, "check with a human." Nothing proceeds in silence.
- Intake ── The intake clerk turns the material into a map of sentences, figures and notes: column layout, rejoining cut sentences, tying figures to their notes, measuring size and how much things stand out. Anything unreadable is kept, not discarded.
- Fix the basis ── The approved-facts store holds the three documents in a fixed form and stamps the version used into the record. Every comparison from here on is against that version.
03Step 4 Rule checks ── settle the unhesitating part first (reproduction gate)
The rule-checker runs the clear-cut half or so of the 694 check items, sentence by sentence and figure by figure. Forbidden words. Where the chart's vertical axis starts. The number of cases. Required statements. The same material always gives the same answer, and the reproduction gate confirms this every time with the exam that has an answer key. There are two outputs. Flags (at full certainty; weak hits are kept together at low certainty), and cues for the next step. Facts such as "this chart's vertical axis does not start at zero" or "this page has three figures for effect and none for safety" guide the reading that follows.
04Steps 5 to 6 Make the picture, take the difference ── four readers read (coverage gate, independence gate)
At step 5, the image maker makes the picture in four ways of reading. Someone who reads one sentence alone. Someone who reads two neighbors in a row. Someone who reads the booklet through. Someone who sees only the chart. For each way, the AI makes the picture three separate times, and everything is combined. Any implication that appeared even once is kept. The independence gate works here: the question is phrased differently across the three runs, to confirm that the same implication does not depend on chance. The coverage gate checks, at the end of the review, that all four pictures were made and compared with every item of the approved facts.
At step 6, the difference taker compares the four pictures with the approved facts. Things not in the approved facts. Things stated more lightly than the approved facts. Things in the approved facts but missing from the picture. Each difference gets a label under fixed rules. Whether it is stated plainly or implied, under which way of reading it got worse, and which sentences or figures produced it are attached as the path. There is one flag per difference in the picture, with several paths if several sentences or figures were involved. The grounds gate works at this point, and a flag missing any required item is not made.
A word flag, an arrangement flag and a whole-booklet flag never attach separately to the same wrong picture. The three cues become three paths of one flag. The number of items a human reads drops here, and what was removed stays in the record as paths.
05Steps 7 to 8 Checking the documents, and the public eye ── beyond the approval, beyond the rules
Step 7 compares the differences in the picture with the government's review report and the safety-measures plan. The review report says what the government accepted and did not accept about effect, safety, target disease and how the drug is used. If the picture contains a claim that was not accepted, the difference is "goes beyond the approval." If an important risk from the safety-measures plan is missing from the picture, the difference is "makes safety look lighter." Numbers (response rates, risk ratios, statistical certainty, ranges) are checked by rule to see whether they trace back to the official leaflet.
The public eye at step 8 reads what no written rule covered. The picture from reading the booklet through is read by the fixed cast of characters, who attach a label on strong unease and resemblance to past incidents. Unease from even one of them becomes a flag. The second-highest label goes to a human, the highest to management. There is no power to stop the material. Which version of the cast did the reading stays on the flag.
06Steps 9 to 11 Exceptions, merging, scoring ── a reason to drop, a path to merge
- Exception desk ── Only the three fixed kinds are dropped, by rule. A quotation where only the dose or interval differs. Results carrying the fixed note. Early-stage trials with the stage clearly stated. The removal gate works here: a dropped flag does not vanish but stays in the record tagged "withdrawn," with the reason and the rule version.
- Merge ── Merge and score combines overlapping flags on the same picture into the heaviest label and keeps the paths. Differences that got worse on reading neighbors, or reading the whole, go up one step. Guaranteeing safety outright, writing that something unapproved is approved, and tampering with a chart's axis are fixed at the heaviest label.
- Score and record ── A score of weight × certainty × visibility is given, and the material is rated on four levels. A single heaviest label means the lowest rating. Public-eye labels are not scored; they only raise a flag. The record and precedent notebook adds the material received, the check items, rules, cast and AI version used, and each flag's path. The coverage gate confirms the full list of checks for the review. The list of flags is complete here and goes to the human.
07Steps 12 to 13 The human verdict and the panel ── three tiers of judgment (independence gate)
At step 12, the human reviewer reads the list. Each flag comes with the quotation from the material, the original text of the rule, the picture and its difference, the path, and past verdicts on similar pictures drawn from the precedent notebook. For each flag the reviewer records fix or no fix, and the reason. That record is the condition for the review to end. A review without it stays "incomplete," and the follow-through gate counts it every week.
The panel at step 13 runs alongside step 12 for heavy flags only, making candidates to help the reviewer's judgment. Which flags go to it is set by rule. The 30 debaters give reasons, counterarguments and minority opinions, and a majority of the 30 voters makes a "should fix" candidate. When the vote looks unanimous, when it is close, when it disagrees with the machine's weight, or when the AI gave no answer, a "check with a human" mark is added. The panel neither deletes nor softens a flag, and its record is kept. When the reviewer overturns the panel's candidate, the reason becomes a precedent.
08Steps 14 to 15 The precedent notebook and the miss ledger ── the review ends when the verdict returns (follow-through gate, continuity gate)
At step 14, the precedent notebook stores the human verdict, tied to the type of picture and difference. When a similar picture appears in the next review, the precedent is attached to the flag. If the earlier verdict and this candidate disagree, and it cannot be shown whether the rule version or the precedent changed, the flag goes to a human. This upholds the principle that the same judgment gives the same verdict.
At step 15, the miss ledger matches the human verdict against the machine's candidates. Where the human said fix and the machine stayed silent is a miss; where the machine flagged and the human said no fix is an overreach. Each miss is annotated with the check item that should have fired, the way of reading that should have been made, and the precedent that should have been consulted. The follow-through gate totals the rate weekly and stops live operation if misses at the heaviest label exceed 1% or misses at the second label exceed 5%. Even if they do not, for each type of miss it drafts fixes to the check items, thresholds, or instructions to the AI. A human approves, the exam with an answer key confirms (the continuity gate), and the fixes return to steps 4 and 5 of the next review. The dashed line in Figure 3 is this loop.
The loop also turns when a rule is revised. When the rule-change tracker spots a revision, it drafts the changes to the check items, a human approves, the exam confirms, and material already passed is re-examined. If material that passed before the revision fails after it, the reason is shown as the difference in rule version.
09Five promises every step keeps
Wherever the steps are, these five are never broken.
- No silent misses ── Missing items, an AI failure, and a conditional "pass" are always brought into the open as "check with a human" or an immediate halt.
- The program decides, the AI reads the meaning ── The AI describes the picture; the program judges it. Nothing a program can decide is handed to the AI.
- No flag without grounds ── A flag without a check item, the rule's original text, the picture and the way of reading is discarded the moment it is made.
- A reason and a record for every drop ── Suppressing, allowing and withdrawing are all recorded against a rule and can be checked afterward.
- A review ends when the human verdict returns ── A review whose verdict has not returned is counted as incomplete. The difference from the returned verdict goes into the miss ledger, and live operation stops if the weekly rate crosses the line. The machine searches before human review and learns after it.
That is the whole of one review. A review does not end when it is handed to the human; it ends when it comes back from the human. That turns "a miss is the gravest failure" from a declaration into a measurement, keeps "same judgment, same verdict" across time, and makes the way of thinking a mechanism that cannot be broken.