1. High accuracy does not mean the whole scale is being used
Instead of asking a large language model to write a long answer, you can ask it to read a text and return only a label or a rating. The authors call models used this way direct-decision models. They are fast and their output is structured, which makes them easy to plug into classification pipelines and automatic evaluation.
The authors' concern is that the reliability of such models cannot be judged by accuracy alone. If a user supplies a scale running from "very poor" to "very good", the model has to use those levels faithfully. A model can score reasonably well on accuracy while pushing most of its decisions toward the middle and rarely using the ends. In that case the scale has effectively become coarser than the one the user asked for.[1]
The analysis covers version 1.13 of a model called JEV and three open KEV models. The abstract does not spell out what JEV and KEV stand for or how they work. The source is an arXiv preprint and has not yet been peer reviewed.
2. The core idea: fix the conditions and measure how much of the scale gets used
This is a measurement paper, not a new model. Its core can be put in one sentence: hold the items and source scores fixed, balance the gold labels and the positions of the options, change only the number of levels on the scale, and measure what share of the scale the decisions actually use.
Controlling the conditions this way removes competing explanations one at a time. Decisions could be skewed because the gold labels are skewed, because certain positions are favored, or simply because there are many options. The authors isolate what remains after all of those are balanced and describe it as a compression of the candidate space at the decision stage. They name it ordinal scale-utilization bias and argue it is distinct from accuracy, gold-label imbalance, fixed position effects, and the number of candidates alone.
3. What the abstract reports
The starting point is ANLI, a natural language inference benchmark.[4] JEV reaches 74.95% accuracy there. Yet even with nearly balanced gold labels and balanced candidate positions, it assigns 38.8% of all predictions and 51.3% of its errors to Neutral. When unsure, it retreats to the middle.
The pattern is not specific to ANLI. Across 36 datasets with ordinal scales, final decisions use only 67–76% of the effective gold support, meaning the range of levels the correct answers actually occupy. On four nominal tasks, where the labels have no order, the figure is 87–102%. Randomizing the order of the candidates weakens the compression but does not remove it.
Making the scale finer makes the effect sharper. As the number of levels is raised from K=2 to 14, utilization falls for every model, reaching 26–75% at K=14. At the same time, the probabilities the models assign across candidates remain broad for most models. The uncertainty is still visible in the probabilities, but the final decisions collapse onto fewer levels.
Finally, targeted post-training with a method the authors call BA-LoRA raises gold-relative utilization from roughly 47% to 86% on eight supervised scales, for both KEV sizes. The authors take this as evidence that the compression is learned and can be changed, rather than an immutable architectural limit. The LoRA family that this builds on freezes the pretrained weights and trains only small added matrices.[5] Code and data are public on GitHub.[3]
4. A worked example: a meeting about making the rating scale finer
Picture a team that uses an AI decision model to sort incoming inquiry records by priority. Someone proposes a finer priority scale to improve precision.
Operations lead: The current scale is too coarse. I want to move to a finer one so we can capture differences in urgency.
Quality lead: Before we change it, have we looked at how often each level is used now? Accuracy can look fine while most decisions sit in the middle and the extremes almost never appear.
Operations lead: We have only tracked accuracy. If the top level is rarely used, the most urgent cases could be landing in "moderate".
Quality lead: Exactly. If that is happening, adding levels might just add more levels that never get used. Let us first put the distribution of decisions next to the distribution of correct labels.
The second half of this exchange is the paper's question. A finer scale and finer decisions are not the same thing, and within this paper's results, adding levels lowered the share of the scale that was used.
5. What is not new, and what the abstract does not tell us
Survey research has long discussed the tendency of human respondents to favor middle options and avoid the extremes of a rating scale. Many studies have also shown that language models are sensitive to the position and order of answer options. The contribution here is to isolate the compression that survives after those factors are controlled, treat it as a decision-stage phenomenon, and give it a name and a measure.
There are limits. First, the abstract does not explain what kind of models JEV and KEV are, and it does not show whether other direct-decision models, or models that reason in text before deciding, behave the same way. Second, the abstract does not give the details of how utilization relative to gold support is defined; values above 100% on nominal tasks are hard to interpret without that definition. Third, the abstract does not say what happened to accuracy or decision validity after BA-LoRA raised utilization. Using more of the scale is only useful if the extra levels are used correctly. Fourth, the abstract does not describe BA-LoRA itself in any detail.
The public repository also has only 3 stars, so there is little sign yet of independent replication.
6. A single paper, not a cluster
The paper drew 60 upvotes on Hugging Face Daily Papers.[2] In this site's selection only the reader-vote signal fired; there is no implementation uptake or press coverage yet. An upvote is a vote that a paper is worth reading. It is not evidence that its claims are correct.
No cluster has formed around the question; the cluster size is 1. Using AI models as automatic graders is a growing research area, but the specific question of whether a grader uses the full scale it is given is not yet being taken up by several groups at once.
The paper is worth covering anyway because it adds a concrete check, separate from accuracy, for anyone building an evaluation system. Swapping in a different grading model or refining a scale are routine decisions in practice, and this paper gives a measurable way to see what those decisions actually change.
7. What connects to pharma and regulatory work
Pharmaceutical work is full of ordinal ratings: severity and causality assessments for adverse events, triage of medical inquiries by urgency, risk grading of promotional wording. Teams are beginning to consider AI decision models for first-pass sorting in these tasks.
Two practices follow. One is to check the distribution of decisions across levels against the distribution of correct labels, not just accuracy. If decisions cluster in the middle and the extremes rarely appear, the cases that most need attention may be buried in "moderate". In safety-related ratings, that can matter more than the headline accuracy figure. The other is to avoid confusing a finer scale with finer judgment. Before adding levels, check whether the existing levels are being used at all.
The evaluation in this paper used public datasets, and it does not show whether the same bias appears, or at the same size, in pharmaceutical rating tasks. This review, too, is limited to what the preprint's abstract states.