arXiv:2608.04054cs.MMcs.AI2026-08

区分模态一致与冲突,提升多模态意图理解准确率

Modality Agreement- and Conflict-Aware Prototype Hypergraph Learning for Multimodal Intent Understanding

论文配图:Modality Agreement- and Conflict-Aware Prototype Hypergraph Learning for Multimodal Intent Understanding
图 1 · 摘自论文原文
  • 构建分层原型超图,分别捕捉模态间一致与冲突关系
  • 在多个基准数据集上实现显著性能提升,优于主流融合方法
  • 适合需要精准识别讽刺、反语等复杂语义场景的研究者

多模态意图识别不仅需理解文本、语音和视觉信号的共性,还需解析其不一致之处。这种不一致常具类别判别性——例如语言积极但声调或表情不符,可能表示讽刺或嘲讽,而现有融合方法通常强制对齐或视不一致为噪声予以抑制。本文提出MACH(模态一致与冲突感知原型超图)框架,通过分层原型超图结构,将模态一致性与冲突分别建模为可复用的关联模式。模型逐级整合单模态表征为双模态与三模态抽象,在每一层级,模态组合锚点激活稀疏的一致性原型超图以捕获通用共识,同时独立的冲突路径将跨模态差异映射至专用冲突原型超图。两条路径通过样本自适应的特征级仲裁机制融合,既能保留有信息量的不一致,又可抑制偶然性噪声。渐进式优化策略稳定层级间的相互依赖关系,实现一致与冲突的联合学习。在多个基准数据集上的实验验证了该方法的有效性,组件分析与鲁棒性测试进一步证实了分层组合、原型驱动语义精炼及仲裁机制的关键作用。

原文摘要 · Abstract (English)

Multimodal intent recognition requires understanding not only what textual, acoustic, and visual signals share, but also how they disagree. Such disagreement is frequently class-informative; for example, lexical positivity accompanied by incongruent vocal or facial behavior may indicate sarcasm or taunting, yet most fusion methods either encourage modality alignment or treat inconsistency as uncertainty to be suppressed. We propose MACH (Modality Agreement- and Conflict-aware prototype Hypergraph), a hierarchical prototype-hypergraph framework that represents multimodal agreement and conflict as distinct, recurring relational structures. MACH progressively composes unimodal representations into bimodal and trimodal abstractions. At each applicable level, modality-composition anchors activate sparse agreement prototype hypergraphs that capture reusable consensus patterns, while a separate conflict pathway maps cross-modal discrepancies to dedicated conflict prototype hypergraphs. The two pathways are combined through a feature-wise, sample-adaptive arbitration mechanism, enabling the model to preserve informative disagreement while suppressing incidental modality noise. A progressive optimization strategy stabilizes the interdependent hierarchy before joint agreement-conflict learning. Experiments on benchmark datasets demonstrate the effectiveness of the proposed formulation, while component and robustness analyses validate the distinct roles of hierarchical composition, prototype-mediated semantic refinement, and agreement-conflict arbitration.

多模态意图识别原型学习冲突建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。