Who decides the evaluation criteria?
Numbers that convey the effectiveness of AI implementation are often disseminated without mentioning in which language, in which context, and under which rules they were measured. For example, even if a student gets a high score on a question set written in English, it does not necessarily mean that the same standard will be maintained in other languages. Some are beginning to document what they consider safe, such as China's national efforts and the New York City Council's efforts. As long as standards, procedures, and rules for each country are not finalized, the person holding the yardstick that produces the numbers will have more control over the distribution of results and responsibilities than the numbers themselves. In this article, we'll take a look at four things you should check before accepting numbers: language, rules, human involvement, and verification.
DeepMind called for a review of AI evaluations that are biased towards English and called for consideration of cultural differences (Chosunbiz). If the assessment question sets and grading standards are written in English, high scores may only indicate performance in English-speaking countries. The score alone does not tell us whether the same standards will be maintained in other languages and cultures. An article (Vocal) about conversational bots for mental health points in the same direction. Even with bots designed to be culturally savvy, challenges remain when it comes to crisis response. Being easy to use in daily consultations and being able to move appropriately in serious situations are two different performances, and evaluation of the former does not guarantee the latter. Those considering the introduction of the system need to check in which language and culture, and in which situation, the scores proposed were measured.
It is not only in terms of language that the standards for what is considered an achievement are wavering. Regarding AI awareness, Anthropic and Microsoft each expressed their views (Yahoo Finance). Development companies do not agree on whether they are conscious or not, and how to handle them. An article about Anthropic's biological experiments (IBM) discusses what the experiments showed. What constitutes success or significance of an experiment can vary depending on the reader's perspective. The two cases mentioned here are both topics that require us to decide from which point of view we should evaluate them before coming to a conclusion.
These questions are not limited to discussions among experts, but are also reaching local decision-making forums. West Lafayette school board candidate spoke about AI and the budget during a student-sponsored panel (Purdue Exponent). How much a school spends on AI is inseparable from how it measures the effectiveness of its implementation. The numbers and expectations presented by candidates also have different weight depending on which criteria they are measured against.





