提出测量AI行为倾向的新框架,揭示倾向比能力更能预测模型表现与安全。
Capabilities Ain't All You Need: Measuring Propensities in AI
- 用双逻辑函数建模模型在理想倾向区间内表现最优
- 六类大模型测试显示倾向偏移可被量化并影响任务表现
- 结合倾向与能力预测效果优于单独使用任一指标
AI评估长期聚焦于能力测量,现有方法多基于项目反应理论(IRT),但忽略了模型行为倾向——即表现出特定行为的倾向性。传统IRT将成功视为能力与任务难度的单调函数,不适用于倾向性分析,因过度或不足均可能引发问题。本文首次提出基于双逻辑函数的正式框架来度量AI倾向:当模型倾向处于“理想区间”时,成功概率最高。我们利用新开发的无任务特异性评分标准,通过大语言模型估算该理想区间的边界。对六类大模型在双向倾向诱导下的测试表明,可准确测量倾向偏移程度及其对任务的影响。关键发现是:仅用基准测试估计的倾向能有效预测未见任务上的行为;结合倾向与能力的预测性能显著优于单独使用任一指标。本框架展示了如何严谨测量倾向,并证明其在预测AI行为上的优势。
原文摘要 · Abstract (English)
AI evaluation has primarily focused on measuring capabilities, with formal approaches inspired from Item Response Theory (IRT) being increasingly applied. Yet propensities - the tendencies of models to exhibit particular behaviours - play a central role in determining both performance and safety outcomes. However, traditional IRT describes a model's success on a task as a monotonic function of model capabilities and task demands, an approach unsuited to propensities, where both excess and deficiency can be problematic. Here, we introduce the first formal framework for measuring AI propensities by using a bilogistic formulation for model success, which attributes high success probability when the model's propensity is within an "ideal band". Further, we estimate the limits of the ideal band using LLMs equipped with newly developed task-agnostic rubrics. Applying our framework to six families of LLM models whose propensities are incited in either direction, we find that we can measure how much the propensity is shifted and what effect this has on the tasks. Critically, propensities estimated using one benchmark successfully predict behaviour on held-out tasks. Moreover, we obtain stronger predictive power when combining propensities and capabilities than either separately. More broadly, our framework showcases how rigorous propensity measurements can be conducted and how it yields gains over solely using capability evaluations to predict AI behaviour.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。