arXiv:2412.08529cs.CL2024-12中稿 · PACLIC 2024被引 3

用常识知识增强文本,提升多模态意图识别准确率

TECO: Improving Multimodal Intent Recognition with Text Enhancement through Commonsense Knowledge Extraction

  • 从生成与检索知识中提取关系,丰富文本语境信息
  • 融合增强后的文本与视觉音频特征,形成统一表征
  • 适合需要精准理解用户意图的对话系统研究者

多模态意图识别(MIR)旨在结合文本、视频、音频等多源信息以识别用户意图,对对话系统理解语言与上下文至关重要。尽管该领域已有进展,但仍面临两大挑战:(1) 有效提取并利用文本中的语义信息;(2) 非语言模态与语言模态的有效对齐与融合。本文提出文本增强框架TECO,通过从生成与检索的知识中提取关系,增强文本模态的上下文信息,并将强化后的文本与视觉、听觉表示对齐融合,构建统一的多模态表征。实验表明,该方法显著优于现有基线模型。

原文摘要 · Abstract (English)

The objective of multimodal intent recognition (MIR) is to leverage various modalities-such as text, video, and audio-to detect user intentions, which is crucial for understanding human language and context in dialogue systems. Despite advances in this field, two main challenges persist: (1) effectively extracting and utilizing semantic information from robust textual features; (2) aligning and fusing non-verbal modalities with verbal ones effectively. This paper proposes a Text Enhancement with CommOnsense Knowledge Extractor (TECO) to address these challenges. We begin by extracting relations from both generated and retrieved knowledge to enrich the contextual information in the text modality. Subsequently, we align and integrate visual and acoustic representations with these enhanced text features to form a cohesive multimodal representation. Our experimental results show substantial improvements over existing baseline methods.

多模态意图识别文本增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。