arXiv:2604.21144cs.CLcs.AI2026-04

让对话智能体用视觉化记忆保持共同认知,避免信息混淆。

Using Machine Mental Imagery for Representing Common Ground in Situated Dialogue

论文配图:Using Machine Mental Imagery for Representing Common Ground in Situated Dialogue
图 1 · 摘自论文原文
  • 用动态视觉记录对话状态,构建可追溯的共享场景
  • 在IndiRef上比纯文本推理提升12.3%,减少语义模糊
  • 适合需要长期上下文理解的对话系统研究者

情境对话要求说话者维持可靠的共享上下文表征,而非仅依赖孤立话语。当前对话智能体常在此方面表现不佳,尤其当共同认知需超出即时上下文窗口时。此时细粒度差异常被压缩为纯文本表示,导致我们称为“表征模糊”的关键失败模式:相似但不同的实体被合并为可互换描述。这种语义扁平化制造了虚假的对话一致感,使智能体看似局部连贯,却无法持久追踪共享上下文。受人类思维中心理意象作用的启发,并基于多模态模型的可用性,我们探索对话智能体是否可通过生成某种具象中间表示来解决此问题。为此,我们提出一种主动视觉支架框架,将对话状态逐步转化为持久的视觉历史,以供后续生成有依据的回应。在IndiRef基准上的评估表明,增量外部化本身已优于全对话推理,而视觉支架通过减少表征模糊并强制具体场景承诺带来额外收益。同时,文本表示在非可视信息上仍具优势,混合多模态设置取得最佳整体性能。这些发现表明,融合具象与命题信息的显式多模态共同认知表征,能显著提升对话智能体表现。

原文摘要 · Abstract (English)

Situated dialogue requires speakers to maintain a reliable representation of shared context rather than reasoning only over isolated utterances. Current conversational agents often struggle with this requirement, especially when the common ground must be preserved beyond the immediate context window. In such settings, fine-grained distinctions are frequently compressed into purely textual representations, leading to a critical failure mode we call \emph{representational blur}, in which similar but distinct entities collapse into interchangeable descriptions. This semantic flattening creates an illusion of grounding, where agents appear locally coherent but fail to track shared context persistently over time. Inspired by the role of mental imagery in human reasoning, and based on the increased availability of multimodal models, we explore whether conversational agents can be given an analogous ability to construct some depictive intermediate representations during dialogue to address these limitations. Thus, we introduce an active visual scaffolding framework that incrementally converts dialogue state into a persistent visual history that can later be retrieved for grounded response generation. Evaluation on the IndiRef benchmark shows that incremental externalization itself improves over full-dialog reasoning, while visual scaffolding provides additional gains by reducing representational blur and enforcing concrete scene commitments. At the same time, textual representations remain advantageous for non-depictable information, and a hybrid multimodal setting yields the best overall performance. Together, these findings suggest that conversational agents benefit from an explicitly multimodal representation of common ground that integrates depictive and propositional information.

对话系统多模态共同认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。