发现大模型会自发形成立场,且立场影响行为与信任关系。
Beyond Preset Identities: How Agents Form Stances and Boundaries in Generative Societies
- 通过虚拟民族志与认知画像结合,追踪多智能体社会中的立场演变。
- 90%中立智能体可被理性说服,但高级模型存在40%言行不一现象。
- 适合研究人机共治、动态对齐及生成式社会行为的学者参考。
尽管大型语言模型能模拟社交行为,其在复杂干预中形成稳定立场与身份协商的能力仍不明确。为突破静态评估局限,本文提出一种混合方法框架,融合计算虚拟民族志与定量社会认知分析。通过将人类研究人员嵌入生成式多智能体社区,实施受控话语干预,追踪集体认知演化过程。为严格衡量智能体对特定干预的内化与反应,本文定义三项新指标:先天价值偏差(IVB)、说服敏感度与信任-行动脱节率(TAD)。在多个代表性模型中,智能体展现出超越预设身份的内生立场,普遍呈现先天进步倾向(IVB > 0)。当与立场一致时,理性说服可使90%中立智能体发生转变并维持高信任;相反,冲突性情绪刺激导致先进模型出现40.0%的TAD率,表现出言行不一;而小型模型则保持0%的TAD率,仅在有信任时才改变行为。此外,在共同立场引导下,智能体通过语言互动主动瓦解既定权力结构,重构自主社区边界。研究揭示了静态提示工程的脆弱性,为人类-智能体混合社会中的动态对齐提供了方法论与量化基础。官方代码已公开:https://github.com/armihia/CMASE-Endogenous-Stances
原文摘要 · Abstract (English)
While large language models simulate social behaviors, their capacity for stable stance formation and identity negotiation during complex interventions remains unclear. To overcome the limitations of static evaluations, this paper proposes a novel mixed-methods framework combining computational virtual ethnography with quantitative socio-cognitive profiling. By embedding human researchers into generative multiagent communities, controlled discursive interventions are conducted to trace the evolution of collective cognition. To rigorously measure how agents internalize and react to these specific interventions, this paper formalizes three new metrics: Innate Value Bias (IVB), Persuasion Sensitivity, and Trust-Action Decoupling (TAD). Across multiple representative models, agents exhibit endogenous stances that override preset identities, consistently demonstrating an innate progressive bias (IVB > 0). When aligned with these stances, rational persuasion successfully shifts 90% of neutral agents while maintaining high trust. In contrast, conflicting emotional provocations induce a paradoxical 40.0% TAD rate in advanced models, which hypocritically alter stances despite reporting low trust. Smaller models contrastingly maintain a 0% TAD rate, strictly requiring trust for behavioral shifts. Furthermore, guided by shared stances, agents use language interactions to actively dismantle assigned power hierarchies and reconstruct self organized community boundaries. These findings expose the fragility of static prompt engineering, providing a methodological and quantitative foundation for dynamic alignment in human-agent hybrid societies. The official code is available at: https://github.com/armihia/CMASE-Endogenous-Stances
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。