现有大模型评估对模糊任务失效,揭示了评测设计的根本缺陷。
Evaluating Ill-Defined Tasks in Large Language Models
- 分析模糊任务评估中评测基准与指标的失效原因
- 发现指令遵循与自然语言转流程图任务存在评分不稳、标准不一问题
- 强调需更可解释、可诊断的评估设计,适合研究者优化模型
大量大语言模型(LLMs)评估针对的是本质模糊的任务,其输入输出空间不明确,成功标准模糊。我们分析了现有评测基准和指标为何无法为这类任务提供可靠或诊断性信号。通过两个案例研究:复杂指令遵循(CIF),识别出覆盖不足、对指令措辞敏感、度量不一致且不可比、以及基于LLM的评判者引入不稳定性等问题;自然语言转Mermaid序列图(NL2Mermaid),展示多维度评价标准可提供超越总体分数的切实洞察。两项研究共同表明,当前评估常混淆不同失败模式,导致评分不稳定、非诊断性且难以采取行动。研究揭示了现有模糊任务评估实践的根本局限,并推动更稳健、可解释的评估设计。
原文摘要 · Abstract (English)
Many evaluations of Large Language Models (LLMs) target tasks that are inherently ill-defined, with unclear input and output spaces and ambiguous success criteria. We analyze why existing evaluation benchmarks and metrics fail to provide reliable or diagnostic signals of model capability for such tasks. We examine two case studies: Complex Instruction Following (CIF), where we identify recurring issues including limited coverage of real-world instruction complexity, sensitivity to instruction phrasing, inconsistent and non-comparable metrics, and instability introduced by LLM-based judges; and Natural Language to Mermaid Sequence Diagrams (NL2Mermaid), where we show how multi-faceted evaluation criteria can yield actionable insights beyond aggregate scores. Together, these case studies show that current evaluations frequently conflate distinct failure modes, yielding scores that are unstable, non-diagnostic, and difficult to act upon. Our findings expose fundamental limitations in existing evaluation practices for ill-defined tasks and motivate more robust, interpretable evaluation designs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。