提出检测协作对话中隐藏分歧的新方法,让未说出口的分歧可见。
Illusion of Alignment: Detecting Hidden Disagreement in Collaborative Dialogue

- 设计诊断选择题,通过回答差异暴露隐藏分歧
- 真实会议中每场平均发现2.89个未言明分歧
- 适合研究人机协作、团队沟通效率的学者
协作对话中常出现表面一致实则目标、假设或执行计划不一致的现象,称为‘对齐错觉’(IoA)。一项涵盖18场真实会议的研究证实,这种现象在人类协作中频繁发生。然而,若参与者意识到分歧,就不会保持沉默;若未意识到,又无法主动表达,导致其对所有人隐形。本文提出通过生成诊断性多选题,使不同参与者的分歧答案成为可观察的行为证据,构建了覆盖五类任务与六个领域的IoA-Suite数据集与评估协议。最佳模型仅达49.5% F1,瓶颈在于对话未呈现的私有上下文。基于该数据集训练的IoA-Prober-8B模型达到51.8% F1,能在18场真实会议中每场识别出2.89个参与者未言明的分歧,并成功应用于多人智能体协作,在BigCodeBench-Hard与HiddenBench上提升下游任务表现。
原文摘要 · Abstract (English)
Collaborative dialogue can end with apparent agreement while participants still differ on goals, assumptions, or execution plans, creating an \textbf{illusion of alignment (IoA)}. A real-user study across 18 meetings confirms that IoA arises routinely in human collaboration. Yet IoA poses a paradox: if participants were aware of such disagreements, they would already be explicit; if not, they cannot articulate them when asked, leaving IoA invisible to both participants and observers. In this work, we make IoA detectable by generating diagnostic multiple-choice questions whose divergent answers across participants provide direct behavioral evidence of hidden disagreement. We construct \textbf{IoA-Suite}, a dataset and evaluation protocol for detecting hidden disagreement, spanning five task types and six domains. We find that even the best model attains only 49.5\% F1, with the bottleneck traced to private context that the dialogue does not surface. We then train \textbf{IoA-Prober-8B} based on IoA-Suite, reaching 51.8\% F1 on IoA-Suite. Across the aforementioned 18 real meetings, it surfaces 2.89 hidden disagreements per meeting that participants confirm they had not voiced, transferring to live human dialogue. Further, in multi-agent collaboration, pairing IoA-Prober-8B with LLM agents improves downstream task performance on BigCodeBench-Hard and HiddenBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。