大模型在政治分析中常无法坚持分配角色,导致观点失真。
When Roles Fail: Epistemic Constraints on Advocate Role Fidelity in LLM-Based Political Statement Analysis
- 通过推理文本识别角色立场,不依赖表面词汇
- Mistral Large角色保持率67%,远超Claude Sonnet的39%
- 语言和事实核查工具会影响角色稳定性,需纳入评估
民主话语分析系统越来越多地采用多代理大模型流水线,为不同评估模型分配对立角色以生成结构化、多视角的政治声明评估。核心假设是模型能可靠维持其角色。本文首次基于TRUST管道对这一假设进行系统性实证检验。我们开发了一种无需依赖表面词汇的先验立场分类器,通过四个指标(角色漂移指数、期望漂移距离、方向漂移指数、熵基角色稳定性)评估60条政治声明(30条英文,30条德文)的角色忠实度。发现两种失效模式:认知下限效应(事实核查结果形成不可逾越的下限)与角色优先冲突(训练知识覆盖角色指令),二者本质均为认知角色覆盖机制。模型选择显著影响角色忠实度:Mistral Large表现优于Claude Sonnet 28个百分点(67% vs. 39%),且呈现角色放弃但不反转极性的质变行为,而Claude则主动转向对立立场。角色忠实度具有语言鲁棒性。事实核查提供者并非中立:困惑度显著降低Claude在德语陈述上的角色忠实度(Δ = -15pp, p = 0.007),而对Mistral无影响。这些发现直接指向多代理大模型验证的盲区:若未测量角色忠实度,系统可能系统性扭曲其本应提供的认知多样性。
原文摘要 · Abstract (English)
Democratic discourse analysis systems increasingly rely on multi-agent LLM pipelines in which distinct evaluator models are assigned adversarial roles to generate structured, multi-perspective assessments of political statements. A core assumption is that models will reliably maintain their assigned roles. This paper provides the first systematic empirical test of that assumption using the TRUST pipeline. We develop an epistemic stance classifier that identifies advocate roles from reasoning text without relying on surface vocabulary, and measure role fidelity across 60 political statements (30 English, 30 German) using four metrics: Role Drift Index (RDI), Expected Drift Distance (EDD), Directional Drift Index (DDI), and Entropy-based Role Stability (ERS). We identify two failure modes - the Epistemic Floor Effect (fact-check results create an absolute lower bound below which the legitimizing role cannot be maintained) and Role-Prior Conflict (training-time knowledge overrides role instructions for factually unambiguous statements) - as manifestations of a single mechanism: Epistemic Role Override (ERO). Model choice significantly affects role fidelity: Mistral Large outperforms Claude Sonnet by 28pp (67% vs. 39%) and exhibits a qualitatively different failure mode - role abandonment without polarity reversal - compared to Claude's active switch to the opposing stance. Role fidelity is language-robust. Fact-check provider choice is not universally neutral: Perplexity significantly reduces Claude's role fidelity on German statements (Delta = -15pp, p = 0.007) while leaving Mistral unaffected. These findings have direct implications for multi-agent LLM validation: a system validated without role fidelity measurement may systematically misrepresent the epistemic diversity it was designed to provide.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。