Under scoring that punishes whoever admits ignorance or error, errors stay hidden for people and AI alike. On wards, fear of responsibility discourages reports; in AI evaluation, "I don't know" scores zero, so errors miss the record and guesses appear. The remedy is to decide first not to penalise the speaker: psychological safety, changed scoring and first-person uncertainty make errors visible and fixable.
Image abstract — the whole article on one page (click to enlarge)

Answer "I don't know" and you score zero; guess, and you might be right. In September 2025, OpenAI researchers argued that because AI models are graded this way, they guess instead of admitting they don't know. What does showing your weakness, admitting you don't know or that you got something wrong, make possible? It makes it possible for errors to be found and fixed. That chance lasts where the person who speaks up is not penalised for it.

01Being able to say "I don't know" creates the chance to fix errors

When someone says "I don't know" or "I got that wrong", the error moves into a place where other people can see it. An error that can be seen can be fixed. One that stays hidden cannot. The management scholar Amy Edmondson observed this sequence on hospital wards, and later in teams at a manufacturing company.

The same sequence is at work in how AI is evaluated. In a 2025 paper, Adam Tauman Kalai and colleagues at OpenAI pointed out that most benchmarks for language models give "I don't know" zero points. As a result, models learn to make confident guesses.

Whether the speaker is a person or a model, errors stay out of sight when the scoring punishes whoever admits them. So anyone who evaluates a colleague's report or an AI's answer has to decide, in advance, not to penalise the one who speaks up.

02Eight hospital units, 1996: the better-led ones recorded more medication errors

The first study to put numbers on this sequence looked at hospital records. In a 1996 paper, Edmondson compared the recorded medication errors of eight units in university hospitals. Units with stronger direction from the nurse manager and better working relationships recorded more errors, not fewer.

Read naively, that says good units make more mistakes. Edmondson read it another way. The number of errors that end up in the record depends on two things: how many errors actually happen and how often they are reported. On units with good relationships, nurses found it easier to report their errors, so more of them reached the record. That was her interpretation.

Figure 1 How an error reaches the record
YesAn error occursA medicationerrorCan it besaid?It enters therecordThe cause canbe fixedYesAn error occursA medication errorCan it be said?It enters the recordThe cause can be fixed
Only the errors that can be spoken reach the record, and only recorded errors can be fixed.

When this essay speaks of showing weakness, it does not mean lacking ability in general. It means saying out loud that you don't know or that you got something wrong. A nurse reporting their own medication error is one form of it.

03From drug administration to AI answers, errors surface only in places that allow them to be spoken

Recorded errors depend on how easy it is to report them. Edmondson found this in the administration of medicines on hospital wards. On the scale of hospital error, the U.S. Institute of Medicine published its report To Err Is Human in 2000. The report estimated that as many as 98,000 people die each year from medical errors that occur in hospitals. It also found that fear of being held responsible discourages people from reporting errors.

Tests that measure AI models have a similar kind of scoring. Many benchmarks mark an answer only as right or wrong. Under that scheme "I don't know" scores zero, the same as a wrong answer, while a guess has some chance of scoring.

AspectNurses on a hospital unitLanguage model evaluation
Admitting an error or uncertaintyFear of being held responsibleScores zero
Staying silentThe error never reaches the recordPlausible wrong answers appear
What makes speaking up possibleGood relationships on the unitScoring that penalises wrong answers more heavily

The two settings differ in the work being done and in who is speaking. What they share is that whether anyone speaks up depends on what comes back to the one who does.

04People and AI alike withhold errors when the scoring punishes whoever admits them

On the ward, admitting an error invites blame; in an AI benchmark, admitting ignorance costs points. Why, then, do people fall silent and models start guessing?

In 1999 Edmondson studied 51 teams at a manufacturing company. She looked at a shared belief among team members: that in this team, asking questions or reporting mistakes would not get you blamed. She defined it as a shared belief that the team is safe for interpersonal risk taking, and called it psychological safety.

Teams with more psychological safety engaged more in learning behaviour, such as discussing mistakes and asking for help. Teams that engaged more in learning behaviour also performed better.

Figure 2 Scoring that punishes the speaker
Nurses on aunitLanguagemodel…Admit an errorHeld responsibleStay silentAnswer "I don'tknow"Score zeroGuess insteadNurses on a unitLanguage modelevaluationAdmit an errorHeld responsibleStay silentAnswer "I don'tknow"Score zeroGuess instead
In both settings, when speaking up costs the speaker, errors and uncertainty stay hidden.

Kalai and his co-authors showed the same sequence for language models in mathematical form. Under right-or-wrong grading, "I don't know" is the worst possible answer. A confident guess earns a higher expected score. The paper itself mentions students who fill in blanks on an exam by guessing. Models, too, adapt to the grading and learn to produce an answer even when they do not know.

05When AI added "I'm not sure, but", fewer users followed its wrong answers

If scoring that punishes the speaker makes errors hide, how can scoring be built that does not, and how much difference does it make? Three studies give examples that can be counted.

1

Change the scoring

Kalai and colleagues proposed telling a model to answer only if its confidence exceeds t, with wrong answers penalised t/(1−t) points. "I don't know" scores zero and is no longer worse than a wrong answer.

2

Voice uncertainty in the first person

In an experiment by Kim and colleagues with 404 participants, when the AI added "I'm not sure, but" in the first person, users' accuracy on medical questions the AI got wrong rose from 43.6% to 52.0%.

3

A shared belief in the team

Google's internal research named psychological safety the most important of five dynamics that set effective teams apart.

According to Kim and colleagues, first-person expressions of uncertainty made users less likely simply to agree with the AI's answer. The gain in accuracy was largest on the questions the AI got wrong. When an AI states its own uncertainty, users are less easily pulled along by its mistakes.

06A unit that reports many errors, or an AI that says "I don't know", is not necessarily a bad sign

When the AI voiced uncertainty in the first person, fewer people followed its wrong answers. Whether a report or an answer that shows weakness turns out to be useful depends on how the receiver treats it. Three things are worth reading differently.

1

Units that report more

Across the eight units, the better-led ones recorded more errors.

2

An AI's "I don't know"

First-person uncertainty reduced over-reliance on the AI's wrong answers.

3

The rules of deduction

Under right-or-wrong grading, guessing pays best.

Anyone who compares departments by their error reports should not assume that the department reporting more is the worse one. A department with few reports may make few errors, or it may simply be unable to speak. The count alone cannot tell the two apart.

Anyone who uses AI answers should not treat "I don't know" or "I'm not certain" as a defect. The same goes for documents drafted with generative AI. Where the AI has marked a passage as needing checking, keep the mark and start the review there.

Anyone who writes the rules of evaluation should first settle on scoring that does not punish the one who speaks up. For staff appraisals, that means not lowering someone's rating because they reported their own error. For AI evaluation, it means penalising a wrong answer more heavily than "I don't know".

07More "I don't know" from AI hinges on the scoring set by builders and users

Will scoring that does not punish the speaker spread? Kalai and colleagues wrote that AI guessing will not decline unless the scoring of widely used benchmarks changes. Adding one more benchmark that measures errors is not enough. If most benchmarks keep grading only right or wrong, models will keep leaning towards guessing. Whether AI says "I don't know" more often depends on how the people who build evaluations, and the people who use AI, treat that answer.

Figure 3 What scoring without penalty sets in motion
Say "I don't know"The error becomesvisibleFix the causeNo penalty forspeakingSay "I don't know"The error becomesvisibleFix the causeNo penaltyfor speaking
With a rule that does not penalise the speaker, saying, seeing and fixing can repeat.

What this essay can say about human workplaces has limits. Edmondson's 1996 study observed eight units. The reading that more records reflect easier reporting was not tested by experiment. The 98,000 deaths in To Err Is Human is the upper end of an estimate.

Still, I am sure of one thing. How to treat the person who admits an error is something the receiver can decide. The time to decide is before anyone speaks up.

Key Points ── 3 to take away
  1. In the eight-unit study, better-led units recorded more medication errors. Do not assume a unit that reports more errors makes more.
  2. When AI voiced uncertainty in the first person, users' accuracy on its wrong answers rose from 43.6% to 52.0%. Do not treat an AI's "I don't know" as a defect.
  3. Under right-or-wrong scoring, "I don't know" scores worst and guessing pays. Set scoring that does not punish whoever speaks up, for people and AI alike.
Closing

Showing your weakness, saying you don't know or that you were wrong, brings errors into view and makes fixing them possible. That chance does not last on courage alone. It lasts where the scoring does not punish the person who speaks.

For human reports and an AI's "I don't know" alike, it is the receiver who sets that scoring first. For my part, I have decided in advance what my first words will be when someone tells me, "I got it wrong."

Sources & references
  1. Amy C. Edmondson. Learning from Mistakes Is Easier Said Than Done: Group and Organizational Influences on the Detection and Correction of Human Error. Journal of Applied Behavioral Science 32(1), 1996.(Medication error records on eight units, and the reporting interpretation)
  2. Amy Edmondson. Psychological Safety and Learning Behavior in Work Teams. Administrative Science Quarterly 44(2), 1999.(Definition of psychological safety; links to learning behaviour and performance in 51 teams)
  3. Institute of Medicine. To Err Is Human: Building a Safer Health System. National Academies Press, 2000.(Estimate of up to 98,000 deaths a year from hospital errors; barriers to reporting)
  4. Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, Edwin Zhang. Why Language Models Hallucinate. arXiv:2509.04664, 2025.(Binary grading rewards guessing; the t/(1−t) scoring proposal; the need to change mainstream benchmarks)
  5. Sunnie S. Y. Kim, Q. Vera Liao, Mihaela Vorvoreanu, Stephanie Ballard, Jennifer Wortman Vaughan. "I'm Not Sure, But...": Examining the Impact of Large Language Models' Uncertainty Expression on User Reliance and Trust. FAccT 2024 / arXiv:2405.00623.(404 participants; 43.6% to 52.0%)
  6. Think with Google. Team dynamics: The five keys to building effective teams. 2023.(Psychological safety as the most important dynamic in Google's research)