How to create grounds for believing in AI

The explanations, scores, and diagnoses provided by AI appear plausible to the reader. However, its accuracy cannot be determined just by looking at it. The explanations provided by research agents about ``why this result should occur'' are usually not measured to see whether they improve predictions. One study reported that it could not confirm the contribution of natural explanations. On the other hand, a randomized trial of a respiratory diagnostic chatbot with 2,400 people showed that it improved the accuracy of judgments made by ordinary people. What we will look at in this half-day is the difference between claims that AI output is correct and results that are backed up by human confirmation and verification procedures. It is becoming clear that the decision to introduce AI is determined not by its performance itself, but by whether the support is available.
When an AI research agent plans an experiment, it includes an explanation of why the result should be this way. Readers use these as clues to estimate the outcome, but whether the explanation improves the prediction is usually not measured. One preprint (on its own website) calls the amount that an explanation adds to a prediction the ``predictive contribution,'' and sets up a procedure to measure it by pairing predictions with only the explanation swapped. They reported that they could not confirm the contribution of natural explanations, as measured by toxicity tests and machine learning tasks. Whether an explanation can be read plausibly and whether it is useful for prediction are two different questions.
There is the same type of review on the scoring side. When performing LLM reinforcement learning on tasks where there is no single correct answer, it is becoming increasingly common to award partial points based on scoring criteria (rubrics) that list the requirements for the answer. Another preprint (on their own site) points out that the scorer may give points for information not included in the answer, calling this Vacuous Credit. As a countermeasure, we are proposing MetaRubric, which alternates between a learning process in which points are awarded only when there is evidence, and a process in which standards are revised while looking at the answers. A high score does not prove that the written content is good. This is because the habits of the grader directly determine the direction of learning.
The handling of success records is also being questioned. To help students learn difficult tasks, records of actual successful tasks are useful. However, if success relies on the assistance of a special harness, that assistance cannot be used in a real environment. The third preprint (own site) proposed Recursive Self-Rewrite (RSR), which extracts successes found under a variety of harnesses into a procedure manual, checks for leaks, and redoes them with a common harness. This is a procedure to rewrite the successful trajectory of terminal work into learning data with the same conditions as the actual one. A record of success can only be used if it is confirmed that it can be reproduced without assistance.




