arXiv:2502.07577cs.LGcs.AI2025-02被引 5

用大模型自动发现另一大模型的能力与缺陷,省去人工设计测试题。

Automated Capability Discovery via Foundation Model Self-Exploration

  • 让一个大模型充当科学家,自动生成开放任务来探测目标模型能力。
  • 在多个大模型上生成数千个任务,发现数十种新能力与失败模式。
  • 可自动评分且结果与人工评估高度一致,适合大规模模型评测。

基础模型已成为跨领域的通用助手,但其能力谱系仍难以精确刻画。现有评估方法依赖大量人工,且随着模型变强,设计挑战题的难度持续上升。我们提出自动化能力发现(ACD)框架,让一个基础模型担任‘科学家’,系统性地生成开放任务以探测目标模型(可能包括自身)的能力。结合前沿开放性研究思想,ACD能自动、系统地揭示目标模型中多样且意外的能力与失败模式。我们在GPT、Claude、Llama系列等多款基础模型上验证该方法,成功生成数千个不同任务,并通过聚类识别出数十个能力范畴与失效模式,这些成果单个团队极难独立发现。我们还通过大规模人工调查验证了模型自评的可靠性,发现模型生成评分与人工评估高度一致。借助大模型的任务生成与自我评估能力,ACD为可扩展、自动化的新型AI系统评估迈出关键一步。所有代码与评估日志已开源:https://github.com/conglu1997/ACD。

原文摘要 · Abstract (English)

Foundation models have become general-purpose assistants, exhibiting diverse capabilities across numerous domains through training on web-scale data. It remains challenging to precisely characterize even a fraction of the full spectrum of these abilities and potential risks in any new model. Existing evaluation approaches often require significant human effort, and it is taking increasing effort to design ever harder challenges for more capable models. We introduce Automated Capability Discovery (ACD), a framework that designates one foundation model as a scientist to systematically propose open-ended tasks probing the abilities of a subject model (potentially itself). By combining frontier models with ideas from the field of open-endedness, ACD automatically and systematically uncovers a diverse spectrum of surprising capabilities and failures in the subject model. We demonstrate ACD across a range of foundation models (including the GPT, Claude, and Llama series), showing that it automatically generates thousands of distinct tasks, which are then clustered to reveal dozens of broader capability areas and failure modes, that would be challenging for any single team to uncover. We further validate our method's automated scoring with extensive human surveys, observing high agreement between model-generated and human evaluations. By leveraging foundation models' ability to both create tasks and self-evaluate, ACD is a significant step toward scalable, automated evaluation of novel AI systems. All code and evaluation logs are open-sourced at https://github.com/conglu1997/ACD.

自动化评测大模型能力自探索能力发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。