综述多模态意图识别的深度学习方法与进展
Deep Learning Approaches for Multimodal Intent Recognition: A Survey
- 从单模态到多模态的演进,融合文本、语音、视觉等数据
- 基于Transformer的模型在多模态意图识别中表现突出
- 适合关注人机交互与多模态学习的研究者参考
意图识别旨在识别用户的潜在意图,传统上聚焦于自然语言处理中的文本。随着对自然人机交互需求的增长,该领域通过深度学习和多模态方法不断发展,融合了音频、视觉及生理信号等数据。近期,基于Transformer的模型在此领域取得显著突破。本文综述了深度学习在意图识别中的应用,涵盖从单模态到多模态技术的演变、相关数据集、方法论、应用场景及当前挑战,为研究者提供多模态意图识别(MIR)最新进展与未来方向的洞察。
原文摘要 · Abstract (English)
Intent recognition aims to identify users' underlying intentions, traditionally focusing on text in natural language processing. With growing demands for natural human-computer interaction, the field has evolved through deep learning and multimodal approaches, incorporating data from audio, vision, and physiological signals. Recently, the introduction of Transformer-based models has led to notable breakthroughs in this domain. This article surveys deep learning methods for intent recognition, covering the shift from unimodal to multimodal techniques, relevant datasets, methodologies, applications, and current challenges. It provides researchers with insights into the latest developments in multimodal intent recognition (MIR) and directions for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。