arXiv:2608.30726cs.AI2026-08

动态选择专家模型,让情感分析更精准。

Multimodal Adaptive Expert Selection with Text Routing and Ordinal Prototype Optimization for Sentiment Analysis

论文配图:Multimodal Adaptive Expert Selection with Text Routing and Ordinal Prototype Optimization for Sentiment Analysis
图 1 · 摘自论文原文
  • 用文本引导动态路由,智能激活音频视觉专家
  • 在CMU-MOSI/MOSEI上达到最新最佳性能
  • 适合需要精细情感识别的多模态研究者

多模态情感分析(MSA)是情感计算的核心任务,旨在融合语言内容与语音语调、面部微表情等非语言线索,理解复杂情绪状态。尽管近期解耦方法有所进展,但仍面临两大挑战:一是静态计算图对所有样本一视同仁,无法适应不同语义复杂度;二是通用对比目标忽略情感强度的内在序数关系。为此,我们提出MAESTRO框架,通过文本引导的混合专家机制,动态激活音频-视觉专家以增强跨模态特征表达;同时引入序数感知原型对比学习(O-PCL),通过距离惩罚项构建有序潜在空间,保持情绪强度的自然顺序。在CMU-MOSI和CMU-MOSEI数据集上的实验表明,MAESTRO实现当前最优性能,定性分析也验证了其动态路由的可解释性。

原文摘要 · Abstract (English)

Multimodal Sentiment Analysis (MSA) is a fundamental component of affective computing that aims to decipher complex emotional states by integrating verbal content with non-verbal cues including vocal intonation and facial micro-expressions. While recent disentanglement-based approaches have advanced the field, their potential is hindered by two methodological challenges. First, static computation graphs process all samples indiscriminately regardless of semantic complexity, which leads to suboptimal representation for diverse emotional expressions and contextual scenarios. Second, generic contrastive objectives often neglect the intrinsic ordinal hierarchy of sentiment intensities. To systematically address these limitations, we introduce Multimodal Adaptive Expert Selection with Text Routing and Ordinal prototype optimization (MAESTRO), a novel framework designed to dynamically orchestrate and refine multimodal representations. Drawing inspiration from an orchestra conductor, we design a Text-Guided Hybrid Mixture-of-Experts (MoE) mechanism. Unlike static fusion, this module utilizes linguistic context as a routing signal to dynamically activate specific audio-visual experts, thereby resolving cross-modal ambiguity through adaptive feature enhancement. Furthermore, to capture fine-grained sentiment gradations, we propose an Ordinal-aware Prototype Contrastive Learning (O-PCL). By incorporating distance-based penalties into the prototype learning objective, O-PCL enforces a structured latent space that preserves the natural order of emotion. Extensive experiments on the CMU-MOSI and CMU-MOSEI benchmarks demonstrate that MAESTRO achieves state-of-the-art performance, and qualitative analysis further confirms the interpretability of our dynamic routing paradigm.

情感分析多模态专家网络序数学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。