arXiv:2606.22797cs.AIcs.CY2026-06

测试大模型在相似决策环境间的行为迁移能力,发现效果不稳定。

Measuring Behavior Portability in Large Language Models

论文配图:Measuring Behavior Portability in Large Language Models
图 1 · 摘自论文原文
  • 用多个源环境数据训练可解释模型,再测其在目标环境表现。
  • 实验显示七类经济决策中行为迁移性能显著下降。
  • 适合关注大模型泛化能力与评估可靠性的研究者。

大型语言模型越来越多地被用作自主决策者,但其行为表现可能在表面不同但收益结构相同的决策环境中存在显著差异。这种敏感性使基于套件的评估变得脆弱,并引发一个核心问题:在某一环境中学习到的行为映射,在保持相同激励结构的另一环境中是否仍具参考价值?本文提出一种正式框架来度量该特性。协议将来自一组源环境的数据合并训练一个可解释的行为模型,并在保留的目标环境中评估其外样本预测性能,以直接在目标数据上训练的最优模型为基准。通过一种不依赖损失函数的度量方式,给出目标环境中诱导出的预测-行动映射性能的最坏情况边界。在涵盖七个经典经济决策问题的受控实验中,我们记录到显著且系统性的行为迁移损失,表明从单一环境中获得的大模型行为表征无法可靠地推广至结构等价的其他环境。

原文摘要 · Abstract (English)

Large language models are increasingly deployed as autonomous decision makers, yet the behavioral mapping they exhibit can vary substantially across decision environments that are payoff-equivalent by construction-environments that share identical payoff-relevant structure but differ in surface presentation. This sensitivity renders suite-based evaluation fragile and raises a fundamental question of behavioral portability: how well does a behavioral mapping learned in one decision environment informative on another that preserves the same underlying incentive structure? We introduce a formal framework to measure this property. Our protocol fits an interpretable behavioral model on data pooled from a set of source environments and evaluates its out-of-sample predictive performance in a held-out target environment, benchmarking against an oracle trained directly on target data. Portability is quantified via a loss-agnostic measure that delivers worst-case bounds on the performance of the induced prediction-action mapping in the target environment. In controlled experiments spanning seven canonical economic decision problems, we document substantial and systematic portability losses, suggesting that behavioral characterizations of LLMs obtained in one decision environment cannot be assumed to transfer reliably to structurally equivalent alternatives.

行为迁移大模型评估决策模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。