1. Who scores generated images, and how

Improving a model that generates or edits images requires something that turns each output into a number. That component is a reward model. Its scores are used both to compare models and as the training signal when a model is tuned with reinforcement learning.

The paper's complaint is that most existing reward models take the instruction and a candidate image and map them directly to a single score. If the instruction is "a person holding a red umbrella on a rainy street corner," the things worth checking are specific to that request: the umbrella's color, how the hand holds it, whether the rain and the street read correctly. In the usual approach, what should be evaluated for that case is left implicit inside the model. A low score does not tell anyone what was missing.[1]

The primary source is a preprint on arXiv and has not been peer reviewed.

2. The core idea: decide what matters before scoring

The authors call their principle "Think Before You Score." Its core is a single step: before judging how well a candidate performs, determine explicitly what matters for this particular case.

The Thinking Reward Model (TRM) follows that principle. It first formulates a case-adaptive rubric, then assesses the candidate against that rubric, and finally produces a fine-grained pointwise reward for each candidate. Because the rubric comes out as text before the score, a person can trace what the score was based on.

The second contribution concerns training. The authors observe that conventional pairwise preference optimization, which teaches the model which of two candidates is better, can push scores toward the extremes, a problem they call score polarization. They introduce Pairwise Dual-Group Relative Policy Optimization (PD-GRPO), which uses pairwise supervision to sharpen discrimination while keeping fine-grained pointwise scores. GRPO, the method it builds on, was originally introduced to train mathematical reasoning.[5]

3. What the abstract reports

The authors ran experiments on reward-modeling benchmarks for both image generation and image editing. They report that TRM achieves state-of-the-art performance among open-source reward models while remaining highly competitive with proprietary alternatives.

They also report that using TRM as the reward in reinforcement learning consistently improved a range of visual generation models. They read this as evidence that fine-grained, case-adaptive rewards work as an optimization signal.

The abstract gives no figures, no benchmark names and no names of the proprietary models compared. How large "state-of-the-art" and "highly competitive" are cannot be checked from the abstract alone. All of these are claims under the paper's own experimental conditions.

4. A worked example: checking a generated training illustration

Suppose a team uses an image generator to produce an illustration for internal training material, and a reward model scores the result. The instruction was "show the steps of handwashing with a person standing at a sink."

Staff member: The conventional reward model gave this image a fairly low score. It gives no reason.

Reviewer: Then there is no way to tell the designer what to fix.

Staff member: With a case-by-case rubric, the model first lists criteria such as "the person is standing at a sink," "the hand movement reads as handwashing," and "the fingers do not look unnatural," then scores each one. Here the deduction came from the finger criterion.

Reviewer: That makes the fix clear. The list of criteria itself still has to be checked, though. If something the training purpose cannot do without, such as the tap or the soap, is missing from the rubric, a high score doesn't make the image usable.

The second half of the exchange is the point. Once the basis of a score is visible, the thing that needs checking moves from the score to the rubric. The same model writes the rubric, so gaps or bias in it can only be caught by a person reading it.

5. What is not new, and what the abstract leaves open

Having a language model write its evaluation criteria before scoring, or produce written reasons alongside a score, has been widely tried for evaluating text. GRPO is an existing method. The contribution here is carrying that approach into reward models for image generation and editing, combining case-specific rubrics with fine-grained pointwise scores in a single model, and addressing score polarization through the training method.

Several limits remain. First, the abstract does not say whether the quality of the rubrics TRM writes was itself evaluated. If a rubric has gaps, a fine-grained score simply measures those gaps in fine detail. Second, any reward model used in reinforcement learning invites reward hacking, where the generator learns to exploit the reward model's quirks and raise the score without real improvement; the abstract does not address this. Third, the abstract does not show how general the score-polarization observation is. Fourth, writing a rubric for every case should add compute per evaluation, and that cost is not reported.

6. Why this topic is clustering now

This paper was selected as the representative of a cluster of 2. On the paper-sharing page it has 93 upvotes and 2 comments, and the public repository has 32 stars.[2][3] Upvotes and stars indicate the interest of readers and implementers; they do not show that the claims are correct.

The other paper in the cluster, OmniTaskonomy, asks when and how training a model to generate images improves its ability to understand them.[4] That is a different question from reward modeling, and the link between the two is loose. What they share is an interest in measuring image generation not only by the quality of the output but by how it connects to evaluation and understanding. The cluster is thin, and it is better read as an early stage of a topic than as an established one.

7. Where this connects to pharmaceutical and regulatory practice

Review of promotional materials is, at bottom, a process of deciding what to look at for each case before judging it. Even with the same checklist, a piece built around efficacy claims and a piece built around safety tables call for different emphasis. TRM's approach mirrors that review discipline on the image-evaluation side.

What the paper offers, however, is a tool that supports evaluation, not one that can be handed the evaluation. Because the model writes its own rubric, gaps in the rubric do not show up in the score. If generated images are used in materials, or a reward model is used to check them, the workflow needs a step where a person reads the model's criteria and confirms that the points regulation requires are actually there.

A score with a visible basis is easier to work with than one without. But a visible basis is not the same as a correct one. Keeping that distinction, and remembering that this is a preprint, decides whether such a tool can be brought into regulated work.