Three of us read the same piece of promotional material, and three different findings came back. One asked where the numbers had come from. One flagged the gap between what the headline claimed and what the body text actually supported. One imagined the meeting where the page would be shown, and worried about what a representative would add out loud. All three were right. None of them overlapped. Looking at that list, I started wondering what exactly I had spent years trying to smooth away.

01Three findings, pointing in three directions

It was a re-read of a product piece. Instead of splitting the work, we put the same material in front of three reviewers. The point of the exercise was alignment. We did not get it.

The first reviewer went after the patient population behind the graph: the restrictions on the study group were not stated anywhere in the body text. The second wrote that the headline asserted a degree of effect the body text never established. The third started somewhere else entirely, from a question: "This page gets used in a briefing. When someone opens it, what will they say out loud that isn't printed here?"

In the meeting we tried to reconcile the three and decide which was the real finding. We could not. Of course we could not — the three were not competing. We were trying to narrow down a set of things that had no conflict in them. What we were doing was not comparison. It was discarding.

On the way home I kept turning over one thought: the work of removing variation might actually have been the work of throwing information away.

02We have been calling it error

When we talk about review quality, we say "reduce the variation" almost reflexively. Underneath sits an assumption: the same material should produce the same conclusion regardless of who reads it. That assumption is strong, and it rarely gets examined.

I understand why. From the outside, variation looks like unfairness. If the same phrasing passes with one reviewer and comes back with another, the people submitting cannot trust the standard. So we flatten. We thicken the manual, add rows to the check sheet, narrow the room for judgment.

Kahneman, Sibony and Sunstein argued that the scatter among expert judgments on identical cases is itself a problem organizations fail to measure. Their point is not that someone is wrong; it is that judgment has an inherent spread, and if you never measure it, your response to it stays at the level of exhortation. So far I agree.

Past that point I slow down. What follows measurement is not necessarily suppression. First you have to separate which part of the scatter is error and which part is information. Flatten without separating, and you lose the errors along with the range of sight the organization actually possessed.

Back to those three findings. That was not a scatter of mistakes. Three people were looking at three different things. Flatten it, and two thirds disappears.

03Angle, standpoint, and association are not one thing

The single word "variation" swallows three phenomena of quite different kinds. I have started naming them separately.

1

Angle — what gets looked at

Which element of the material draws attention first. The figure, the headline, the axis of the chart, the placement of the footnote. Where the eye lands first differs by person, shaped by training and by the cases someone has handled. Same page, different object of attention.

2

Standpoint — where it is looked at from

Whose position the reader occupies. The regulator. The physician deciding whether to prescribe. The internal briefing room. The moment five years out when this page gets cited back. Move the position and the same sentence reads as safe or as dangerous.

3

Association — how it gets connected

How this material links to past cases or knowledge from elsewhere. "This has the same shape as something that caused trouble on an unrelated product." That leap lives here. It is the hardest to put into words and the most personal of the three.

The three are independent. Same angle, different standpoint, and the conclusions split. Same standpoint, different association, and the depth of the finding changes. Because we folded all three into "variation," our response stayed uniform. Thickening the manual does align angle. It does almost nothing for standpoint. For association it does nothing at all. We kept applying a remedy where it could not work, then complained that nothing improved.

Separate them and the work separates too. Angle can be shared. Standpoint can be declared. Association can only be recorded and collected.

04Putting it into words is not the same as making it reproducible

This is the part I most want to get right.

Putting a judgment into words means moving what is inside you to the outside. Turning "something bothers me here" into "what bothers me is that the headline asserts a degree of effect the body text does not establish." That step alone matters. But it is still closed inside one person. Nothing guarantees another reader of the same page arrives at the same place.

Making it reproducible is one step further: putting it in a form another person can retrace. What was looked at, from where, against which premises, along what chain of reasoning. If the premises and the steps are written, someone else can follow them. And having followed them, that person can push back — "that premise does not hold in this case." Only when a judgment can be argued with does it become shared property.

What to look atStopping at wordsGoing through to reproducible form
What comes outA conclusion and a personal reasonPremises, standpoint, what it was checked against, the chain
What others can doAgree, or notRetrace it, test it, argue with it
What carries to the next caseExperience held by one personA conditional rule the group can hold
Does it aggregateNo. The count just growsYes. There is a surface for judgments to meet on

Polanyi wrote that we know more than we can tell. Much of an experienced reviewer's judgment sits on the untellable side. Nonaka and Takeuchi described how that untellable knowing takes form through dialogue and shared work, and begins to circulate as knowledge the organization holds. In our setting, writing down the reason behind a finding is the first stroke of that circulation.

It usually stops at one stroke. A record whose reason field reads "inappropriate in context" is words, but it is not reproducible form. The next reader learns that this person felt it was inappropriate; they learn nothing about what to do when a similar page arrives. On busy days I have written exactly that.

05Collective judgment is not a vote

Gathering variation does not automatically make a group smarter. Take that lightly and the collection makes it duller.

Surowiecki set out the conditions under which a group beats its individuals: diversity of opinion, independence of the members, distributed knowledge, and a means of aggregating. Drop any one and the group falls below the individual. When independence goes, it falls fast.

Asch showed that people will revise their answer about the plain length of a line when everyone around them says otherwise. In a review meeting, the most experienced person speaks first and a junior reviewer then states the same conclusion. That may not be agreement. It may be conformity. Unanimity with conformity mixed in produces confidence with nothing under it.

Page gave a mathematical account of when diversity outperforms individual ability: if people err in different directions, the errors cancel when the judgments are combined. If everyone errs in the same direction, no number of people reduces the error. A group where every reviewer took the same training, read the same manual and was coached by the same senior colleague drifts toward the second case. Standardization quietly removes the group's capacity to cancel its own errors.

So the moment you decide to treat variation as individuality, something else has to be designed alongside it. Time to judge independently before anyone speaks. A defined procedure for combining. And a record of who looked from which standpoint. Saying "diversity matters" without those three is borrowing the conclusion while dropping the conditions.

06Review AI 1.0 and what 2.0 would be

The tooling around material review has grown considerably these past few years. Nearly all of what is running now, as far as I can see, rests on one design idea: checking against rules. Against the list of prohibited expressions. Against the approved indication text. Against the source figures. Fast, tireless, hard to slip past. Call it 1.0.

1.0 is an instrument for aligning angle. On what gets looked at, it is far steadier than a person. But it holds no standpoint. Asked where it is reading from, it can only read from the side of the rules. Association is further still. "This has the same shape as a case from another therapeutic area" is a leap that falls outside checking altogether.

1

Review AI 1.0 — checking against rules

Holds one set of correct answers and matches material against it. Output is a pass or fail plus the clause relied on. It beats people on speed and coverage. It goes silent wherever the rules are unwritten. In this design, variation is noise to be removed.

2

Review AI 2.0 — several standpoints applied at once

Holds the reasoning chains of multiple reviewers, once those chains have been written down in reproducible form, as separate ways of seeing. One page gets read from the regulator's side, the prescriber's side and the briefing room at the same time. Output is not a single verdict but the concern from each standpoint, its reason, and where the standpoints disagreed.

What matters in 2.0 is that the disagreements are not hidden. 1.0 looks reassuring because nothing splits. 2.0 looks unsettling because things do. But the genuinely dangerous state is the other one: a single answer emerging while nobody notices the split underneath it. A visible split is a marker that judgment is required there. Given the marker, people can spend their time on it.

What to look atReview AI 1.0Review AI 2.0
Treatment of variationNoise to removeInformation to keep
Knowledge taken inRules, prohibited wording, approved textAll of that, plus each reviewer's reasoning chain
Shape of the outputOne verdict and its basisThe view from each standpoint, and the points of disagreement
What people doConfirm the resultPlace judgment at the split, and feed the reason back in

Tetlock and Gardner found that the most accurate forecasters break their judgments into parts, record their reasons, and when they are wrong go back to find which premise failed. If 2.0 ever runs, it runs only on top of that habit. The instrument does not come first. The practice of writing down reasons comes first, and the accumulation becomes the instrument.

07Does the individual disappear

I want to face this rather than step around it. Once individual judgments are collected, once the reasoning becomes formal and anyone can retrace it, does the person who made the judgment become unnecessary? Is offering your way of seeing the same as offering up your place?

For a long time I had no answer. I kept writing the records anyway. Here is what I think now.

What becomes formal is only the past judgment. A chain: this material, from this standpoint, under these premises. That does become something others can retrace. But arriving at a standpoint nobody has taken yet does not become formal. Association is association precisely because it happens outside the existing forms. The more forms there are, the wider the outside grows.

Schön described how practitioners rework their own method while in the middle of the work — not executing a procedure but questioning it as they execute. That activity does not fit inside the procedure once it is written out. It stays with the person who wrote it out.

So the individual does not disappear. What changes is the form of remaining. From "the person who can make this judgment" to "the person who can supply this standpoint." The first is replaceable. The second remains by producing the next standpoint each time the last one is absorbed. What you hand over becomes form, and you move to the outside of the form. Continuing to move is what your individuality looks like.

This is not a comfortable arrangement. It means letting go, repeatedly, of judgments you worked to acquire. But hold on to them and they simply age inside you, reaching no one.

Closing

Those three findings are still in the record. It does not say which one was right. It says what each person looked at, where they looked from, and what they connected it to. Six months later, when a page with a similar structure came through, another reviewer opened that record and added a fourth way of seeing.

The variation was not error waiting to be smoothed. Put into words, made reproducible, and combined with independence intact, it became individuality. The individuality stacked up into footing for the next person to stand on.

One person's review becomes a review carried out by the collective judgment of a Team. At that point the work changes: from raising the accuracy of one pair of eyes to handling several ways of seeing at once. That, I think, is where Review AI 2.0 begins. What the beginning needs is not a new instrument. It is writing the reason next to today's finding.

Key Points ── 3 to take away
  1. "Variation" swallows three different things: angle, standpoint and association. Flatten without separating them and you lose the organization's range of sight along with the errors. A thicker manual works on angle, and barely touches the other two.
  2. Putting a judgment into words moves it outside you; making it reproducible lets someone else retrace and argue with it. Stop at the first and the count of records grows without becoming shared judgment. A reason field holding only a conclusion gives the next reader nothing to stand on.
  3. Collective judgment has conditions: diversity, independence, distributed knowledge, a means of combining. Independence is the one that fails first and hurts most. If you decide to treat variation as individuality, design independent judgment time and a combining procedure at the same moment.
Sources & references
  1. Daniel Kahneman, Olivier Sibony, Cass R. Sunstein. Noise: A Flaw in Human Judgment. Little, Brown Spark, 2021. (On the scatter among expert judgments of identical cases, and organizations' failure to measure it.)
  2. James Surowiecki. The Wisdom of Crowds. Doubleday, 2004. (Sets out diversity, independence, decentralization and aggregation as the conditions under which groups outperform individuals.)
  3. Scott E. Page. The Difference: How the Power of Diversity Creates Better Groups, Firms, Schools, and Societies. Princeton University Press, 2007. (A formal account of when diversity outweighs individual ability, given errors in differing directions.)
  4. Solomon E. Asch. "Opinions and Social Pressure." Scientific American, vol. 193, no. 5, 1955. (Conformity experiments in which unanimous wrong answers shift individual judgment on a plainly visible fact.)
  5. Michael Polanyi. The Tacit Dimension. Routledge & Kegan Paul, 1966. (The argument that we know more than we can tell.)
  6. Ikujiro Nonaka, Hirotaka Takeuchi. The Knowledge-Creating Company. Oxford University Press, 1995. (How tacit and explicit knowledge convert into one another to generate organizational knowledge.)
  7. Cass R. Sunstein, Reid Hastie. Wiser: Getting Beyond Groupthink to Make Groups Smarter. Harvard Business Review Press, 2015. (Conditions under which groups decide worse than their members, and how to design discussion against them.)
  8. Philip E. Tetlock, Dan Gardner. Superforecasting: The Art and Science of Prediction. Crown, 2015. (Breaking judgments into parts, recording reasons, and returning to failed premises as drivers of accuracy.)
  9. Donald A. Schön. The Reflective Practitioner: How Professionals Think in Action. Basic Books, 1983. (On practitioners reworking their own method in the middle of the work.)
  10. Standards for Fair Advertising of Pharmaceuticals. Ministry of Health, Labour and Welfare, Japan. (The practical basis for grounds and record-keeping in material review.)