arXiv:2504.07114cs.CLcs.AI2025-04ACL被引 27

将静态评测转向人机协作评估,揭示模型表现与真实交互的差距

ChatBench: From Static Benchmarks to Human-AI Evaluation

  • 把MMLU题目转为用户与AI对话形式,构建人机协作评测数据集
  • 发现AI单独答题准确率无法预测人机协同表现,各学科差异显著
  • 用少量数据微调用户模拟器,可大幅提升对协作效果的预测能力

随着大语言模型聊天机器人快速普及,亟需评估人类与模型协同所能达成的效果。然而,传统基准如MMLU仅衡量模型独立表现(即‘仅AI’)。本文通过用户研究,将MMLU问题转化为用户与模型的对话任务:用户携带问题,与大模型互动以求解。我们发布ChatBench,包含396个问题、两个大模型的‘仅AI’、‘仅用户’及‘人机协作’三类数据,共144,000条回答和7,336段人机对话。结果表明,‘仅AI’准确率无法预测‘人机协作’准确率,在数学、物理和道德推理等科目中差异显著。通过对对话的分析,揭示了二者在推理路径与错误模式上的根本分歧。最后,我们在部分ChatBench数据上微调用户模拟器,使其对未见问题的预测准确率提升超过20个百分点,为规模化交互式评估提供新可能。

原文摘要 · Abstract (English)

With the rapid adoption of LLM-based chatbots, there is a pressing need to evaluate what humans and LLMs can achieve together. However, standard benchmarks, such as MMLU, measure LLM capabilities in isolation (i.e., "AI-alone"). Here, we design and conduct a user study to convert MMLU questions into user-AI conversations, by seeding the user with the question and having them carry out a conversation with the LLM to answer their question. We release ChatBench, a new dataset with AI-alone, user-alone, and user-AI data for 396 questions and two LLMs, including 144K answers and 7,336 user-AI conversations. We find that AI-alone accuracy fails to predict user-AI accuracy, with significant differences across multiple subjects (math, physics, and moral reasoning), and we analyze the user-AI conversations to provide insight into how they diverge from AI-alone benchmarks. Finally, we show that fine-tuning a user simulator on a subset of ChatBench improves its ability to estimate user-AI accuracies, increasing correlation on held-out questions by more than 20 points, creating possibilities for scaling interactive evaluation.

人机协作评估基准大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。