用推理分歧识别真实价值观不确定性,提升多智能体决策质量
Consensus is Strategically Insufficient: Reasoning-Trace Disagreement as a Knowledge-Representation Signal
- 将推理过程与结论差异抽象为四种符号状态,实现分歧可解释化
- 在内容审核中证明分歧感知路由比盲目共识更有效
- 适合需要价值判断的复杂决策场景,如伦理审核、安全治理
多智能体系统通常通过投票、共识协议或辩论减少分歧。我们指出,在涉及价值判断的任务中,这种追求一致的目标是不足的,因为分歧可能反映真实的规范性不确定性,而非错误。基于人类-人工智能协作审核中推理轨迹分歧的研究,我们提出一种知识表征层,将推理轨迹和决策抽象为符号化的分歧状态。给定具有明确推理过程和二元决策的智能体,我们区分四类状态:收敛一致、分歧一致、收敛分歧与发散分歧。这些状态支持可撤销的战略路由规则。我们在内容审核任务中实现了该框架,并论证了分歧感知路由能连接子符号的大语言模型推理与符号化的知识表示,促进多智能体的战略性推理。
原文摘要 · Abstract (English)
Multi-agent systems are commonly designed to reduce disagreement through voting, consensus protocols, debate, or fault-tolerant aggregation. We argue that this objective is insufficient for value-laden tasks, where disagreement may reflect genuine normative uncertainty rather than agent error. Building on prior work on reasoning-trace disagreement in human-AI collaborative moderation, we propose a knowledge-representation layer in which reasoning traces and agent decisions are abstracted into symbolic disagreement states. Given agents producing explicit reasoning traces and binary decisions, we distinguish four states according to reasoning similarity and conclusion agreement: convergent agreement, divergent agreement, convergent disagreement and divergent disagreement. These states support defeasible strategic routing rules. We instantiate the framework in content moderation and argue that disagreement-aware routing provides a bridge between sub-symbolic LLM deliberation and symbolic knowledge representation for multi-agent strategic reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。