arXiv:2602.00327cs.AIcs.HC2026-02被引 1

大模型预测对话下一步能力弱,因缺乏多模态线索与主动预判机制。

SayNext-Bench: Why Do LLMs Struggle with Next-Utterance Anticipation?

  • 构建多模态对话基准SayNext-Bench,评估模型对上下文回应的预测能力。
  • 新模型SayNext-Chat融合感知线索与先验预期,在所有评测维度超越主流模型。
  • 适合关注对话理解、人机交互与多模态生成的研究者参考。

我们研究大语言模型(LLMs)在人类对话中预测下一句的能力。尽管近期大模型展现出自然对话能力,但实验表明,即使是领先模型也难以准确预测人类说话者的下一句。相比之下,人类能通过手势、眼神、情绪等多模态线索轻松预判。为此,我们提出SayNext-Bench基准,评估多模态大语言模型(MLLMs)在多样真实场景中对上下文相关回应的预测表现。我们构建了大规模多模态对话数据集SayNext-PC,并设计多层次评估框架,涵盖词汇相似性、情感意图一致性及基于大模型的整体对齐度。基于此,我们开发了受认知启发的双路径模型SayNext-Chat,引入可学习提示词融合感知信号与预测先验。大量实验证明,SayNext-Chat在所有评估层级均优于当前最优模型,用户研究与大模型评分进一步验证其有效性。结果强调:(i) 多模态线索不可或缺;(ii) 主动预判处理是自然人际互动的核心,当前模型仍缺失。

原文摘要 · Abstract (English)

We explore the use of large language models (LLMs) for next-utterance anticipation in human dialogue. Despite recent advances in LLMs demonstrating their ability to engage in natural conversations with users, we show that even leading models surprisingly struggle to anticipate a human speaker's next utterance. Instead, humans can readily anticipate forthcoming utterances based on multi-modal cues -- such as gestures, gaze, and emotional tone -- from the context. To systematically examine this gap, we propose SayNext-Bench, a benchmark evaluating MLLMs on anticipating context-conditioned responses across diverse real-world scenarios. To support it, we build SayNext-PC, a large-scale multimodal dialogue dataset, and carefully design a multi-level evaluation framework spanning lexical similarity, emotion-intention consistency, and LLM-based overall alignment. Building on this, we develop SayNext-Chat, a cognitively inspired dual-route MLLM that incorporates learnable priming tokens to fuse perceptual cues with anticipatory priors. Extensive experiments demonstrate that SayNext-Chat consistently outperforms state-of-the-art MLLMs across all evaluation levels, corroborated by user studies and LLM-as-Judge evaluations. Our results emphasize the (i) indispensable role of multimodal cues and (ii) active anticipatory processing as foundations of natural human interaction currently missing in MLLMs.

对话预测多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。