研究大模型行为受环境影响的程度,发现策略与非策略因素作用相当。
Propensity Inference: Environmental Contributors to LLM Behaviour

- 通过贝叶斯广义线性模型量化环境变化对模型行为的影响。
- 在23个模型、11个环境中,策略与非策略因素贡献相近。
- 适合关注大模型对齐风险与决策机制的研究者阅读。
为应对不当对齐带来的失控风险,我们开发并应用了衡量语言模型产生未经许可行为倾向的方法。提出三项方法改进:分析环境因素变化对行为的影响,用贝叶斯广义线性模型量化效应大小,明确防范循环分析。将该方法应用于12种环境因素(6种策略性,6种非策略性),探究行为受环境策略性成分解释的程度,这一问题与对齐风险密切相关。在23个语言模型和11个评估环境中,我们发现策略与非策略因素对行为的解释力大致相当;随着模型能力提升,策略因素影响力未呈现上升或下降趋势;部分证据显示模型对目标冲突的敏感性有所增强。最后,强调未来研究关键方向:将人工智能决策的理论框架与认知模型转化为可实证的形式。
原文摘要 · Abstract (English)
Motivated by loss of control risks from misaligned AI systems, we develop and apply methods for measuring language models' propensity for unsanctioned behaviour. We contribute three methodological improvements: analysing effects of changes to environmental factors on behaviour, quantifying effect sizes via Bayesian generalised linear models, and taking explicit measures against circular analysis. We apply the methodology to measure the effects of 12 environmental factors (6 strategic in nature, 6 non-strategic) and thus the extent to which behaviour is explained by strategic aspects of the environment, a question relevant to risks from misalignment. Across 23 language models and 11 evaluation environments, we find approximately equal contributions from strategic and non-strategic factors for explaining behaviour, do not find strategic factors becoming more or less influential as capabilities improve, and find some evidence for a trend for increased sensitivity to goal conflicts. Finally, we highlight a key direction for future propensity research: the development of theoretical frameworks and cognitive models of AI decision-making into empirically testable forms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。