arXiv:2509.05624cs.MAcs.LG2025-09被引 1

大模型行为中,动机可精准推断,信念体系却难捕捉。

Behavioral Inference at Scale: The Fundamental Asymmetry Between Motivations and Belief Systems

论文配图:Behavioral Inference at Scale: The Fundamental Asymmetry Between Motivations and Belief Systems
图 1 · 摘自论文原文
  • 用36种行为配置生成百万级轨迹,构建可验证的推理基准。
  • 动机识别准确率98%-100%,信念系统最高仅34.0%。
  • 信念推断是瓶颈,中立类行为易混淆,需关注信号差异。

从可观测行为中能恢复多少关于智能体潜在价值观的信息?这一问题对基于动作序列推断智能体属性的方法至关重要,但在大规模场景下仍缺乏实证答案。我们通过受控实验解决该问题:使用Llama 3.1-8B的LLM代理,在网格世界中分配36种行为配置(9种信念系统 × 4种动机),生成超过150万条行为序列,提供人类研究无法获得的真实标签。经筛选后,分类器在包含10,338个回合和1,200,834条序列的共享标准数据集上训练与评估。结果揭示出显著的不对称性:动机识别准确率达98%-100%,可恢复97%的可用互信息;而信念系统在所有架构下表现受限——即使使用变压器模型,准确率也仅达34.0%,仅恢复16.3%的可用信息,提取效率比动机低6.1倍。各对齐类别准确率介于23.2%(守序中立)至59.4%(混乱邪恶)之间。混淆分析显示,中立区域的行为模糊性集中在“真实中立”,其利他或平衡行为缺乏明确信号,导致相邻类别样本误归。联合推断使36类完整配置分类性能较随机基线提升12.2倍,瓶颈完全来自信念系统推断。信号增强与解释性查询仅带来微弱提升(+3.8%),证实循环结构存在架构上限,非数据不足所致。变压器的34.0%上限是架构限制还是更根本的理论边界,仍有待探究。这些结果刻画了行为观察所能揭示与无法揭示的大型语言模型代理价值的边界。

原文摘要 · Abstract (English)

How much information about an agent's underlying values can be recovered from its observable behavior? This question matters for any approach that infers agent properties from action sequences, yet remains empirically open at scale. We address it through controlled experiments: LLM-based agents (Llama 3.1-8B) assigned one of 36 behavioral profiles (9 belief systems x 4 motivations) generate over 1.5 million behavioral sequences in grid-world environments, providing ground truth unavailable in human behavioral studies. After filtering, classifiers train and evaluate on a shared canonical dataset of 10,338 episodes and 1,200,834 sequences. A fundamental asymmetry emerges in both magnitude and structure. Motivations achieve 98-100% accuracy and recover 97% of available mutual information across all architectures. Belief systems plateau at 24% for LSTMs regardless of capacity, and even transformers reach only 34.0%, recovering 16.3% of available information, a 6.1x asymmetry in extraction efficiency. Per-alignment accuracy ranges from 23.2% (Lawful Neutral) to 59.4% (Chaotic Evil). Confusion analysis maps the failure structure: a neutral zone of behavioral ambiguity centers on True Neutral, absorbing misclassified samples from adjacent Neutral and Good alignments whose prosocial or balance-keeping behavior lacks distinctive signal. Combined inference yields 12.2x improvement over random baseline for full 36-class profile classification, with the bottleneck located entirely in belief system inference. Signal enhancement and explanatory queries yield only marginal LSTM gains (+3.8%), confirming the recurrent ceiling is architectural rather than data-limited. Whether the transformer's 34.0% ceiling reflects a similar architectural-class limit or a more fundamental bound remains open. These results characterize what behavioral observation can and cannot reveal about LLM agent values.

行为推断价值观识别大模型对称性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。