arXiv:2503.01986cs.CLcs.AI2025-03EMNLP被引 4

自动挖掘模型缺陷任务,揭示前沿大模型在多个领域的系统性错误。

Adaptively profiling models with task elicitation

  • 通过自然语言生成新评估任务,自动发现模型行为偏差
  • 发现数百个任务中模型存在系统性失败,如量子与AGI过度关联
  • 适合模型安全评测与鲁棒性研究者使用

语言模型评估常无法捕捉关键的失效模式,迫使专家手动检查输出并构建新基准。本文提出任务诱导方法,可自动构建新评估以刻画模型行为。该方法发现了数百个自然语言任务——数量级远超以往工作——在从预测到网络欺凌等多个领域,前沿模型表现出系统性失败。例如,Sonnet 3.5过度关联量子计算与人工智能;o3-mini在上下文重复虚构内容时易产生幻觉。

原文摘要 · Abstract (English)

Language model evaluations often fail to characterize consequential failure modes, forcing experts to inspect outputs and build new benchmarks. We introduce task elicitation, a method that automatically builds new evaluations to profile model behavior. Task elicitation finds hundreds of natural-language tasks -- an order of magnitude more than prior work -- where frontier models exhibit systematic failures, in domains ranging from forecasting to online harassment. For example, we find that Sonnet 3.5 over-associates quantum computing and AGI and that o3-mini is prone to hallucination when fabrications are repeated in-context.

模型评估系统性错误任务挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。