用游戏化对抗机制动态评估大模型在垂直领域的知识与推理能力。
GuessArena: Guess Who I Am? A Self-Adaptive Framework for Evaluating LLMs in Domain-Specific Knowledge and Reasoning
- 设计对抗性游戏框架,动态生成领域知识与推理任务。
- 在5个垂直领域中显著区分不同模型的知识覆盖与推理完整性。
- 适合需要高适应性评估的行业应用开发者与研究者。
大语言模型的评估长期依赖静态基准,存在两大缺陷:(1)预设测试集难以适配多样应用场景;(2)标准化流程难以捕捉领域知识与上下文推理的细粒度差异。为此,我们提出GuessArena,一种基于对抗游戏交互的自适应评估框架。受《猜我是谁?》游戏结构启发,该框架融合动态领域知识建模与渐进式推理评估,提升评估真实性。在金融、医疗、制造、信息科技和教育五个垂直领域上的实证研究表明,GuessArena能有效区分模型在领域知识覆盖率与推理链完整性的表现。相比传统基准,本方法在可解释性、可扩展性和场景适应性上具有显著优势。
原文摘要 · Abstract (English)
The evaluation of large language models (LLMs) has traditionally relied on static benchmarks, a paradigm that poses two major limitations: (1) predefined test sets lack adaptability to diverse application domains, and (2) standardized evaluation protocols often fail to capture fine-grained assessments of domain-specific knowledge and contextual reasoning abilities. To overcome these challenges, we propose GuessArena, an adaptive evaluation framework grounded in adversarial game-based interactions. Inspired by the interactive structure of the Guess Who I Am? game, our framework seamlessly integrates dynamic domain knowledge modeling with progressive reasoning assessment to improve evaluation fidelity. Empirical studies across five vertical domains-finance, healthcare, manufacturing, information technology, and education-demonstrate that GuessArena effectively distinguishes LLMs in terms of domain knowledge coverage and reasoning chain completeness. Compared to conventional benchmarks, our method provides substantial advantages in interpretability, scalability, and scenario adaptability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。