On 23 September 2026, Anthropic and OpenEvidence announced they would give physicians in about 100 low- and middle-income countries free access to clinical decision support built on AI. When AI that shows its sources is given away worldwide, who checks whether the answers are right? Listing sources only makes checking easier; the burden of being right stays with the clinician who reads the answer and with the national systems that demand records.
01The wider the free rollout, the more the checking falls to whoever receives it
Four things about this announcement can be confirmed. The recipients are physicians in roughly a hundred low- and middle-income countries. There is no cost to them. Each answer arrives with its sources attached. And the product will be tuned to local medical conditions.
All four reduce the work the receiving clinician does at hand. One task does not shrink at all: deciding whether to use that answer on the patient in front of you.
Who receives it
Physicians in about 100 low- and middle-income countries, at no charge.
What it is
A local edition of a product that answers clinical questions by drawing on peer-reviewed research and treatment guidelines.
Division of labour
Anthropic supplies the underlying technology; OpenEvidence handles the tuning to local medical conditions.
There is one change a reader can make tomorrow morning. When you act on an answer that carries citations, write down which of those citations you opened and which you did not. That single line leaves a path back to any error.
02OpenEvidence builds its answers from peer-reviewed papers and treatment guidelines
You cannot argue about responsibility until you have settled what is being handed out. OpenEvidence answers a physician's question by drawing on peer-reviewed medical research and treatment guidelines, and it prints the papers it drew on alongside the answer.
Clinicians in the United States used the product 42 million times in August 2026 alone. The literature it quotes enters through content agreements with medical journals, so the route differs from scraping whatever a general search turns up.
The free edition puts Anthropic's technology underneath and will be adjusted to local medical conditions. What that adjustment actually changes is not stated in the announcement.
So far as the announcement describes it, there is no stage between the question and the answer where a person has to stop. A person enters at exactly one point: the decision to take the answer or leave it.
03Scarce literature and scarce specialists leave no one to cross-check the answer
In some places the literature behind the answer was never within reach to begin with. The free edition is going to countries including Uganda, Angola, Sudan, Haiti and Mongolia. Journal subscriptions, access to specialists and continuing education after graduation all arrive thinly there.
It is not only information that is thin. There are fewer people to cross-check drug information against. In a Japanese hospital a clinician can set the package insert, the society guideline and a colleague's experience side by side and doubt any one answer. Where only one item can be set down, the act of doubting never gets started.
| What to look at | Literature and specialists within reach | Neither easily within reach |
|---|---|---|
| What the answer is checked against | Colleagues and the literature at hand | Sometimes nothing but the AI's answer |
| Speed of noticing an error | Another source catches it early | It can continue unnoticed |
| What the record keeps | Clinical notes and society reports | No established way of keeping it yet |
The difference is not one of ability. The same physician reads the same answer in the same way. What differs is how many places the hand can reach afterwards. Where it can reach only one, what remains is to use the answer or to use nothing; the option of using it after cross-checking is gone.
The third row of the table is the one that returns later. To measure from the outside whether an error occurred, somebody has to have written it down.
04The more citations appear, the more readers feel they have already checked
Where nothing else is available to check against, the attached citation is the only handhold left. How strong a handhold is it?
Large language models sometimes produce well-formed references to papers that do not exist. Author, journal, year and volume are all in place, and nothing distinguishes them until someone actually looks. Work from medical education pushed the problem one step further: readers treat the presence of references as itself a mark of credibility.
That creates a force pointing the wrong way. With one citation, a reader opens it. With ten, the work of opening them one by one is ten times as large. As the count rises, the motive to check each one falls.
Fabrication also varies by subject. An experiment measuring citation accuracy in mental health research found the rate differed by disorder, running from 59 percent at the lowest to 64 percent at the highest. The rate moves with the subject, and in every disorder measured it stayed around six in ten. No single figure describes how often a language model's citations are right.
None of this says wrong answers always come back. What has been reported is that unsafe answers have come back to medical questions posed by patients. That is the extent of what is established.
05WHO placed the job of setting standards on national governments, in more than 40 recommendations
If the presence of a citation reads as proof, then leaving the work of opening citations to individual conscientiousness is a weak arrangement. A document already exists that puts that work outside the individual.
In January 2024 the World Health Organization issued guidance on large multi-modal models. It carries more than forty recommendations, allocated separately to those who develop, those who provide and those who deploy.
Within it, the primary responsibility for setting standards on development and deployment was placed on national governments. The principles named include transparency, explainability, and making clear where responsibility sits.
Who sets the standard
The primary responsibility for standards on development and deployment rests with national governments.
What must be shown
Transparency, explainability, and clarity about where responsibility lies are named as principles.
How many recommendations
More than forty, distributed across the development, provision and deployment stages.
If governments set the standard, then goodwill on the part of whoever gives the tool away does not stand in for oversight in the country receiving it.
06Free, cited and locally tuned: three conditions that reshape the checking
Even granting that governments will set standards, the product is used during the years before those standards exist. That interval divides into three.
1. It costs nothing
Volume rises, and if the share taken without checking stays the same, the total quantity of unchecked answers entering care rises with it. In the United States alone the product was used 42 million times in August 2026. That figure is not a forecast for the free edition. It does indicate how far volume can run where the price barrier is gone.
2. Answers carry citations
The feeling of having checked replaces the act of checking. Citations are sometimes fabricated, and readers treat their presence as proof. Put those two together and the reference list stops prompting verification and starts excusing it.
3. It is tuned to local conditions
Nobody is assigned to check the tuned portion. WHO gave the standard-setting role to national governments, but the hands doing the tuning belong to the party distributing the product. The party that decides and the party that edits are not the same.
One policy document will not close all three. Volume is answered by the form of the record, citation reading by a stated procedure for opening sources, and tuning by oversight that sits with the receiving country. Each needs a different instrument.
07The range of AI error will become visible first in countries that keep records of checking
There is currently no external way to see which of the three has been settled. Any such way has to be built out of records.
The numbers that will surface are counts of use by country. Counts are easy to take. The thing worth knowing is harder to take: the occasions when a physician read an answer and did not act on it.
If rejected answers are recorded together with the reason, it becomes visible where answers diverged from clinical judgment, and from there the range over which AI errs comes into view. If they are not, whether this giveaway improved care will never be established.
So the first thing to put in place is not a usage dashboard. It is a single line in the record for whether the answer was taken.
- US clinicians used OpenEvidence 42 million times in August 2026 alone; a free rollout to about 100 countries increases the number of answers used where there is nothing to check them against.
- Language models can produce well-formed references that do not exist; when citations themselves read as proof, the motive to verify each one falls.
- WHO placed the primary duty of setting standards on national governments; the goodwill of the party giving the tool away does not substitute for the receiving country's records and oversight.
What carries the rightness of an answer is not the machinery that lists sources. It is the clinician who decides whether to use it, and the national rules that ask for a record of that decision.
I am not doubting the value of giving the product away. Where neither journals nor specialists are within reach, an answer grounded in peer-reviewed research arriving at all is a gain. But at the receiving end, nobody outside the consultation can yet say whether those answers were right. The only instrument that would let anyone say so is a record.
Writing down the answer you did not take costs more effort than writing down the one you did. The extra effort buys the only view anyone will have later.
- Reuters. Exclusive: Anthropic, OpenEvidence partner to bring medical AI worldwide. 23 September 2026.(The announcement, the recipients, the absence of cost, and the split between underlying technology and local tuning)
- The Next Web. Anthropic and OpenEvidence to offer free medical AI to doctors in about 100 countries. 23 September 2026.(The roughly 100-country scope and free access for physicians)
- JAMA Network. OpenEvidence and the JAMA Network sign strategic content agreement. 14 January 2025.(Source literature entering through agreements with peer-reviewed journals)
- World Health Organization. Ethics and governance of artificial intelligence for health: large multi-modal models. WHO guidance, 18 January 2024.(More than forty recommendations; primary responsibility for standards on national governments)
- arXiv. Large language models provide unsafe answers to patient-posed medical questions. 25 July 2025.(Unsafe answers returned to medical questions)
- JMIR Medical Education. Citation Accuracy Challenges Posed by Large Language Models. 24 April 2025.(Fabricated citations, and references read as a sign of credibility)
- Journal of Medical Internet Research. Accuracy of Large Language Models When Answering Clinical Research Questions: Systematic Review and Network Meta-Analysis. 2025.(Accuracy on clinical questions being measured in aggregate)
- PubMed. Influence of Topic Familiarity and Prompt Specificity on Citation Fabrication in Mental Health Research Using Large Language Models. 6 November 2025.(Accuracy from 59 percent to 64 percent depending on disorder)
