arXiv:2604.05939cs.AIcs.HC2026-04ACL

提出CVA架构,让大模型智能体行为更真实多样。

Context-Value-Action Architecture for Value-Driven Large Language Model Agents

  • 用价值验证器分离思考与行动,模拟人类动态价值观。
  • 在110万真实交互数据上,行为多样性提升且偏差减少。
  • 适合追求真实、可解释智能体的研究者和开发者。

大语言模型虽能模拟人类行为,但现有智能体常显僵化,这一问题在依赖自我评判的评估中被掩盖。通过真实世界数据验证,我们发现强化提示驱动的推理反而加剧价值极化,导致群体多样性坍塌。为此,我们提出基于S-O-R模型和舒瓦茨基本价值观理论的上下文-价值-行动(CVA)架构。CVA通过在真实人类数据上训练的价值验证器,将行动生成与认知推理解耦,显式建模动态价值激活。在包含超过110万条真实交互轨迹的CVABench数据集上的实验表明,该方法显著优于基线,在降低极化的同时提升了行为保真度与可解释性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown promise in simulating human behavior, yet existing agents often exhibit behavioral rigidity, a flaw frequently masked by the self-referential bias of current "LLM-as-a-judge" evaluations. By evaluating against empirical ground truth, we reveal a counter-intuitive phenomenon: increasing the intensity of prompt-driven reasoning does not enhance fidelity but rather exacerbates value polarization, collapsing population diversity. To address this, we propose the Context-Value-Action (CVA) architecture, grounded in the Stimulus-Organism-Response (S-O-R) model and Schwartz's Theory of Basic Human Values. Unlike methods relying on self-verification, CVA decouples action generation from cognitive reasoning via a novel Value Verifier trained on authentic human data to explicitly model dynamic value activation. Experiments on CVABench, which comprises over 1.1 million real-world interaction traces, demonstrate that CVA significantly outperforms baselines. Our approach effectively mitigates polarization while offering superior behavioral fidelity and interpretability.

智能体价值对齐行为建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。