arXiv:2502.05887cs.CLcs.AI2025-02NAACL被引 4

构建多模态时间感知对话数据集,提升聊天机器人理解时序信息能力

MTPChat: A Multimodal Time-Aware Persona Dataset for Conversational Agents

  • 构建融合语言、视觉与时间维度的多模态对话数据集
  • 提出两个新任务,评估模型对隐含时间线索的理解能力
  • 设计自适应时间模块,有效捕捉多模态时序依赖

理解时间动态对对话代理至关重要,有助于内容分析和决策。然而,特别是面向人格化对话的时间感知数据集仍十分有限,限制了其应用范围和复杂性。为此,我们提出MTPChat,一个包含语言、视觉和时间元素的多模态时间感知人格对话数据集,集成于对话与人格记忆中。基于MTPChat,我们设计两个时间敏感任务:时间下一回复预测(TNRP)和时间定位记忆预测(TGMP),用于评估模型对隐含时间线索和动态交互的理解能力。此外,我们提出一种创新框架,包含自适应时间模块,可有效整合多模态流并捕捉时间依赖关系。实验结果验证了MTPChat带来的挑战,并证明该框架在多模态时间敏感场景中的有效性。

原文摘要 · Abstract (English)

Understanding temporal dynamics is critical for conversational agents, enabling effective content analysis and informed decision-making. However, time-aware datasets, particularly for persona-grounded conversations, are still limited, which narrows their scope and diminishes their complexity. To address this gap, we introduce MTPChat, a multimodal, time-aware persona dialogue dataset that integrates linguistic, visual, and temporal elements within dialogue and persona memory. Leveraging MTPChat, we propose two time-sensitive tasks: Temporal Next Response Prediction (TNRP) and Temporal Grounding Memory Prediction (TGMP), both designed to assess a model's ability to understand implicit temporal cues and dynamic interactions. Additionally, we present an innovative framework featuring an adaptive temporal module to effectively integrate multimodal streams and capture temporal dependencies. Experimental results validate the challenges posed by MTPChat and demonstrate the effectiveness of our framework in multimodal time-sensitive scenarios.

对话系统多模态时间感知数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。