测试小模型在严肃游戏中的评分可靠性,发现其评估结果差异大。
Meta-Evaluating Local LLMs: Rethinking Performance Metrics for Serious Games
- 用五个小型LLM评估玩家在能源社区游戏中的回答
- 部分模型准确识别正确答案,但存在误判和结果不一致问题
- 适合关注AI评估可信度的研究者与开发者
严肃游戏中的开放式回答评估面临独特挑战,因正确性常具主观性。大型语言模型(LLMs)正被探索用于此类场景的评估,但其准确性和一致性仍不确定,尤其针对本地部署的小型模型。本研究系统评估了五种小型LLM在模拟能源社区决策的《En-join》游戏中的表现,采用二分类指标(包括准确率、真正例率、真负例率),对比不同评估情境下的模型表现。结果揭示各模型在敏感性、特异性和整体性能间的权衡,表明部分模型虽能有效识别正确回答,但存在误报或评估不一致问题。研究强调需构建上下文感知的评估框架,并谨慎选择模型以部署为评估工具。该工作推动了对AI驱动评估工具可信度的讨论,揭示了不同LLM架构在处理主观评估任务时的表现差异。
原文摘要 · Abstract (English)
The evaluation of open-ended responses in serious games presents a unique challenge, as correctness is often subjective. Large Language Models (LLMs) are increasingly being explored as evaluators in such contexts, yet their accuracy and consistency remain uncertain, particularly for smaller models intended for local execution. This study investigates the reliability of five small-scale LLMs when assessing player responses in \textit{En-join}, a game that simulates decision-making within energy communities. By leveraging traditional binary classification metrics (including accuracy, true positive rate, and true negative rate), we systematically compare these models across different evaluation scenarios. Our results highlight the strengths and limitations of each model, revealing trade-offs between sensitivity, specificity, and overall performance. We demonstrate that while some models excel at identifying correct responses, others struggle with false positives or inconsistent evaluations. The findings highlight the need for context-aware evaluation frameworks and careful model selection when deploying LLMs as evaluators. This work contributes to the broader discourse on the trustworthiness of AI-driven assessment tools, offering insights into how different LLM architectures handle subjective evaluation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。