arXiv:2608.30719cs.CL2026-08

用心智理论建模对话中隐性误解,提升对话对齐的精准度。

Mind the Gap: Theory-of-Mind-Grounded Friction for Epistemic Alignment

  • 基于心智理论构建四重信念结构,量化对话双方的认知分歧。
  • 消融实验显示,该机制使误解识别率从65%降至26%。
  • 适合需要精准干预的复杂对话系统开发与评估。

高效对话对齐需区分表面协作(如回应流畅)与认知对齐(信念状态趋同);现有基于偏好方法通常仅优化回复层级,未显式建模后者。本文在摩擦型策略优化中引入心智理论(ToM)推理,于每个指代表达处提取四部分信念结构:说话者意图指代对象、听者理解、双方对对方信念的建模。由此可机械计算摩擦信号,捕捉双方自信推进但指向不同指代的‘静默分歧’。在表征层面,去除二阶信念通道后,误解召回率由65%降至26%。在策略层面,奖励塑形(FAR)与信任区域(FTR)变体在干预F1和可信上下文校准上优于DPO,Brier分数独立验证了校准优势。三次训练中,FAR与FTR保持稳定,而DPO波动剧烈,甚至削弱基线策略已有的干预能力。因此,基于心智理论的摩擦信号能有效驱动指代分歧下的上下文敏感干预。

原文摘要 · Abstract (English)

Productive dialogue alignment requires distinguishing \emph{surface coordination} (acknowledgments and smooth task progression) from \emph{epistemic alignment} (convergence of belief states); standard preference-based methods typically optimize response-level preferences without explicitly modeling the latter. We operationalize Theory-of-Mind (ToM) inference as a control signal within Frictive Policy Optimization by extracting, at each referring expression, a four-part belief structure: the speaker's intended referent, the addressee's interpretation, and each participant's model of the other's belief. This makes friction mechanically computable from epistemic-state comparisons, capturing \emph{silent divergence}, where both participants proceed confidently while grounding to different referents. We evaluate the signal at two levels. At the representation level, ablating the second-order channel reduces misunderstanding recall from $65\%$ to $26\%$. At the policy level, reward-shaping (FAR) and trust-region (FTR) variants improve intervention F1 and warranted-context calibration over DPO, with Brier scores independently supporting the calibration gains. Across three training runs, FAR and FTR remain substantially more stable, whereas DPO varies widely and can degrade intervention competence already present in the base policy. Thus, ToM-grounded friction provides a trainable signal for context-sensitive intervention under referential belief divergence.

对话对齐心智理论干预策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。