arXiv:2509.01337cs.MMcs.AI2025-09EMNLP被引 10

用大模型引导小模型做多模态意图理解,提升细粒度语义推理能力。

LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition

  • 借助大模型自动生成语义线索,无需人工设定先验知识。
  • 在多个任务上超越现有方法,提升显著且稳定。
  • 适合需要复杂语义推理的对话与行为分析场景。

从多模态信号中理解人类意图对分析人类行为和增强人机交互至关重要。然而,现有方法在模态层面依赖过重,限制了细粒度语义关系推理,难以实现复杂意图理解。本文提出一种新型的LLM-Guided Semantic Relational Reasoning(LGSRR)方法,利用大语言模型(LLMs)的广泛知识构建语义基础,提升小型模型的关系推理性能。具体而言,提出基于浅到深思维链(Chain-of-Thought, CoT)的策略,自主发现、描述并按重要性排序语义线索,无需依赖人工定义的先验信息。此外,基于逻辑原则形式化建模三种基本语义关系,并分析其细微交互机制,以支持更有效的关系推理。在多模态意图识别与对话行为标注任务上的大量实验表明,LGSRR在多种语义理解场景下均显著优于当前最优方法,表现一致提升。完整数据与代码已公开于https://github.com/thuiar/LGSRR。

原文摘要 · Abstract (English)

Understanding human intents from multimodal signals is critical for analyzing human behaviors and enhancing human-machine interactions in real-world scenarios. However, existing methods exhibit limitations in their modality-level reliance, constraining relational reasoning over fine-grained semantics for complex intent understanding. This paper proposes a novel LLM-Guided Semantic Relational Reasoning (LGSRR) method, which harnesses the expansive knowledge of large language models (LLMs) to establish semantic foundations that boost smaller models' relational reasoning performance. Specifically, an LLM-based strategy is proposed to extract fine-grained semantics as guidance for subsequent reasoning, driven by a shallow-to-deep Chain-of-Thought (CoT) that autonomously uncovers, describes, and ranks semantic cues by their importance without relying on manually defined priors. Besides, we formally model three fundamental types of semantic relations grounded in logical principles and analyze their nuanced interplay to enable more effective relational reasoning. Extensive experiments on multimodal intent and dialogue act recognition tasks demonstrate LGSRR's superiority over state-of-the-art methods, with consistent performance gains across diverse semantic understanding scenarios. The complete data and code are available at https://github.com/thuiar/LGSRR.

多模态意图识别大模型关系推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。