A diagram from the 91% headline to a re-test. The 91% counts criteria passed; counted per task, only 16.4% pass every criterion. An AI graded both evaluation and training, with 80% of the reward on criteria passed. Weights and the evaluation record are public, so users can re-test on their own contracts. A lawyer's review surrounds everything, and the output is draft quality.
Image abstract — the whole article on one page (click to enlarge)

On 1 October 2026 the legal AI company Ivo released Ivo Sage, a free open-source model post-trained for contract work, and reported that its criterion pass rate on contract tasks rose from 0.701 to 0.913. Does a legal AI's "91%" mean it can be trusted with contract review? It does not. Only 16.4% of contract tasks met every criterion, and the grading was done by an AI rather than a person. What the score can speak to is the quality of a draft that a lawyer will still review.

01Ivo Sage's 91% averages criteria; it is not a rate for handing over contracts

Ivo is a legal AI company based in San Francisco. Ivo Sage is built on DeepSeek V4 Flash, a model released by DeepSeek, and post-trained for contract work. Ivo published the weights under the MIT licence, so anyone can use them.

The announcement carries two numbers. On contract tasks, Ivo Sage passed 91.3% of rubric criteria. Yet only 16.4% of tasks had every one of their criteria passed. And the grader was not a lawyer but an AI acting as judge.

My reading is that the score tells you how good a draft is before a lawyer looks at it, and no more. Anyone thinking of bringing AI into contract review should not take the number at face value, but use the published record to re-test the model on their own contracts.

0291.3% counts criteria passed; 16.4% counts tasks that passed them all

The 91.3% and the 16.4% come out of the same evaluation. They differ in what they count.

According to the model card on Hugging Face, Ivo tested Ivo Sage on 180 tasks held out from training, 55 of them contract tasks. An AI grader scored each task's deliverables against that task's rubric in three independent passes. The model card calls this grader an LLM judge.

The 91.3% is the share of rubric criteria that passed. The 16.4% is the share of tasks in which every criterion passed in all three passes. If a single criterion fails, the task does not count as an all-criteria pass. In my view the all-criteria count is closer to asking whether the work could go straight to a lawyer as it stands.

Figure 1 One grading, two numbers
countcriteriacount tasks55 contracttasksLLM judge, 3passesper rubriccriterionCriterion passrate0.913All-criteriapass16.4%count criteriacount tasks55 contract tasksLLM judge, 3 passesper rubric criterionCriterion pass rate0.913All-criteria pass16.4%
Counted by criterion, the same grading gives 0.913; requiring every criterion per task gives 16.4%. The headline 91% is the first.

The base model, DeepSeek V4 Flash, scored 0.701 on criteria and 5.5% on all-criteria passes under the same conditions. Ivo Sage raised the criterion rate by 21 points and nearly tripled the all-criteria rate. The gain is real. Even so, on roughly five contract tasks in six, something on the rubric still failed.

The evaluation used LAB, the Legal Agent Benchmark, a set of tasks published by the legal AI company Harvey. LAB was built to test AI agents on tasks close to real legal work.

03Both the model card and the ABA say a lawyer reviews the output

If only 16.4% of tasks met every criterion, then in the rest the AI grader failed at least one criterion, and someone has to check those points. In contract review, that someone is a lawyer.

Ivo Sage's model card states its intended use plainly: it produces legal work product for review by a qualified lawyer. Ivo itself assumes a lawyer will check the output.

On 29 July 2024 the American Bar Association issued Formal Opinion 512 on lawyers' use of generative AI. As summarised by The National Law Review, lawyers must evaluate the tool they use, analyse its output, and not rely solely on the tool's conclusions. Nor may they let it replace their own judgement.

In Japan, the Ministry of Justice published new guidelines on 21 August 2026 on services that support legal work with AI. They address how AI-assisted contract review sits with Article 72 of the Attorney Act, which reserves legal services to licensed lawyers. I have not yet read the guidelines in full, so I do not go into their content.

I keep the scope to contracts and legal practice. Whether the same holds for review work in other fields cannot be settled from the sources here.

04Training rewarded the LLM judge's score, 80% of it on criteria passed

An AI grader produced the evaluation scores. The same kind of grading also served as the reward during training.

According to the model card, Ivo Sage underwent reinforcement learning on 1,040 LAB tasks. In reinforcement learning, the model's answers are scored, and the model is nudged step by step toward answers that score higher. The scores in this training also came from the AI grader.

The reward formula puts a weight of 0.8 on the criterion pass rate and 0.2 on whether every criterion passed. For the model, passing one more criterion is the easier way to earn more reward.

Figure 2 Training and evaluation used the same kind of grader
TrainingEvaluation1,040 trainingtasksLLM judge scores80% weight oncriteria180 held-out tasksLLM judge, 3 passes91% criterion passTrainingEvaluation1,040 trainingtasksLLM judge scores80% weighton criteria180 held-outtasksLLM judge, 3passes91% criterionpass
The tasks were kept apart, but the grading shared one form. Eighty per cent of the reward sat on the criterion pass rate.

The results fit that pattern. The criterion rate rose by 21 points; the all-criteria rate stopped at 16.4%. That shape of improvement is consistent with where the reward weight was placed. But no source has shown that the weighting caused it. Consistency is as far as I can go.

To keep the picture balanced, here is another number. On RedlineBench, an evaluation not used in training, Ivo Sage's score rose from 30.8 to 45.7. That makes it hard to argue the gains came only from pleasing the AI grader used in training.

05The weights and the 180-task record are public, so users can re-read the grading

Users do not have to take the AI grader's numbers on trust. Ivo has published the evaluation record as a whole.

A Hugging Face dataset holds the list of validation tasks and, for each task, the transcript, the deliverables and the grader's verdicts. The weights are MIT-licensed and distributed as an adapter to be added on top of DeepSeek V4 Flash.

1

Weights

MIT-licensed, distributed as an adapter on top of DeepSeek V4 Flash.

2

Evaluation record

Transcripts, deliverables and grader verdicts for every validation task.

3

Task lists

The training and validation task lists, so overlap with your own contracts can be checked.

The model card also compares Ivo Sage with other models. Cost is each model's own token usage priced at its provider's list price.

Point of comparisonIvo SageDeepSeek V4 Flash (base)Fable 5.1
Criterion pass rate0.9130.7010.941
All-criteria pass16.4%5.5%29.1%
Cost per task$1.24$1.33$142.39

Fable 5.1 passes every criterion on 29.1% of tasks, close to twice Ivo Sage's rate. Its cost per task, though, is $142.39, more than a hundred times Ivo Sage's $1.24. In the same table, GPT-6 Astra matches Ivo Sage at 16.4% all-criteria passes, at $74.29 per task.

These costs are list-price conversions, not what real contract work would cost. The claim that Ivo Sage gets the same score for far less money holds for these 55 tasks and nothing wider.

06The count changes the score, the grader is an AI, and the means to re-test are public

The weights, the transcripts and the cost table are all public. From the facts set out in the sections above, I draw three points about reading the score.

When you see "91%", check the counting method first

On the same 55 contract tasks, the criterion pass rate was 0.913 and the all-criteria rate was 16.4%. Which of the two goes into the headline changes how readers take it. Faced with any legal AI score, first find out whether it counts criteria or whole tasks. Until you know, the numbers cannot be compared.

The score measures work against an AI grader's rubric

Both the evaluation scores and the training reward came from the AI grader. The score should therefore be read as something other than the share of work a lawyer would pass. The AI grader's verdict is itself a tool's conclusion, and ABA Formal Opinion 512 tells lawyers not to rely solely on a tool's conclusions.

Adoption can wait for a re-test on your own contracts

Figure 3 Re-testing the score on your own contracts
Read the verdicts180 public tasksRun own contractsEverycriterion…Lawyer reviewsRead the verdicts180 public tasksRun own contractsEvery criterion met?Lawyer reviews
Read where the model fails in the public record, then count all-criteria passes on your own contracts. A lawyer still does the final review.

The weights are MIT-licensed and the transcripts and verdicts are public. A firm weighing adoption does not have to stop at reading the score; it can run the same test on its own contracts. Ivo's chief executive, Min-Kyu Jung, said in the announcement that results will come from giving models the context of the work: the documents, the playbooks and the way a legal team operates. Testing on your own contracts also shows whether that context is reaching the model.

07No published source yet measures accuracy on Japanese contracts or a firm's own templates

The material for re-testing is public. Even so, some things about Ivo Sage are not measured in any published source.

Ivo Sage was trained on English only. The model card gives no figures for Japanese contracts. Nor do we know whether it fits a firm's own forms: no source shows whether it picks up in-house templates or a team's habitual edits.

Nor does the model card say how often the AI grader's verdicts agree with those of human lawyers. I have not read the technical report itself.

1

Japanese contracts

Training was in English only; no scores on Japanese contracts are given.

2

Grader versus lawyer

Agreement between the AI grader and human lawyers is not reported in the model card.

A Japanese company that wants to try Ivo Sage on contract review would have to measure it on Japanese contracts from scratch. None of the sources consulted for this article reports such results.

What I will look for next is a study that sets the AI grader's verdicts beside lawyers' verdicts. Once that number exists, the 91% the AI gave can be compared with the scores lawyers give.

Key Points ── 3 to take away
  1. On contract tasks Ivo Sage scored 0.913 by criteria passed and 16.4% by tasks passing every criterion. Any headline score needs its counting method checked first.
  2. Both the evaluation and the training reward came from an AI grader, with 80% of the reward on criteria passed. The score is not a lawyer's pass or fail.
  3. The weights are MIT-licensed and every validation transcript and verdict is public. Adoption can now rest on re-testing with one's own contracts rather than on the headline number.
Closing

A legal AI's 91% does not show that it can be trusted with contract review. It shows how well the work met an AI grader's rubric, criterion by criterion, and only 16.4% of contract tasks met every criterion.

Adoption should wait until the public record is re-run on one's own contracts, and even after that the reviewing lawyer stays. Before I look at a score, I now check how it was counted and who did the grading.

Sources & references
  1. LawFuel. Legal AI – Ivo Becomes The First Legal AI Company to Publish a Free Open-Source Model Post-Trained for Long-Horizon Contract Work. 2026-10-01.
  2. Ivo AI. Ivo Sage (model card, Hugging Face). 2026-10-01.
  3. Ivo AI. Ivo Sage training artifacts (dataset, Hugging Face). 2026-10-01.
  4. Harvey. harvey-labs: Legal Agent Benchmark (LAB). GitHub, accessed 2026-10-05.
  5. The National Law Review. American Bar Association Issues Formal Opinion on Use of Generative AI Tools. 2024-08-14.
  6. Ministry of Justice, Japan. Attorney Act (other) — AI and other legal-work support services. 2026-08-21.