arXiv:2605.18884cs.LGcs.CV2026-05被引 4

用分层超球空间建模情绪树,提升多模态情绪识别精度

Navigating the Emotion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition

论文配图:Navigating the Emotion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition
图 1 · 摘自论文原文
  • 将情绪类别嵌入超球空间,按层级逐步检索证据
  • 在多个数据集上准确率超越现有方法,最高提升6.2%
  • 适合需要细粒度情绪分析的智能客服与心理评估系统

多模态情绪识别旨在融合文本、音频和视频信息以理解人类情感状态。尽管多模态大模型具备强大的推理能力,但通常将情绪类别视为独立标签,忽略了人类心理学中丰富的层次化分类结构。此外,缺乏外部上下文知识使其容易过度解读噪声信号,进一步加剧细粒度情绪分类的难度。为此,我们提出HyperEmo-RAG,一种基于检索增强生成的框架,利用结构化情绪知识库。该框架包含两项关键创新:1)分层超球体定位。鉴于情绪分类固有的树状结构,我们将层次化情绪标签与多模态样本联合嵌入连续的超球空间(庞加莱球),并设计分层束搜索推敲机制,从粗到细逐级检索样本;2)结构化证据注入。基于检索到的证据构建证据图,并通过树感知注意力机制与EmotionGraphFormer,将结构化知识作为显式认知上下文注入LLM,保持图结构信息完整性。在多个数据集上的实验表明,HyperEmo-RAG显著优于现有方法。

原文摘要 · Abstract (English)

Multimodal emotion recognition aims to integrate text, audio, and video sources to understand human affective states. Although multimodal large language models excel at multimodal reasoning, they typically treat emotion categories as independent labels, ignoring the rich hierarchical taxonomy of human psychology. Moreover, lacking external contextual knowledge makes them highly susceptible to over-interpreting noisy cues, further complicating fine-grained emotion classification. To address these issues, we propose \textbf{HyperEmo-RAG}, a retrieval-augmented generation framework that leverages a structured emotional knowledge base. Our framework introduces two key innovations. 1) Hierarchical hyperbolic grounding. Recognizing the inherent hierarchical tree structure of emotion taxonomies, we jointly embed hierarchical emotion labels and multimodal samples into a continuous hyperbolic space (Poincaré ball) and design a hierarchical beam-search deliberation process that progressively retrieves samples from coarse to fine-grained levels. 2) Structured evidence injection. Based on the retrieved evidence, we construct an evidence graph and inject the structured knowledge as explicit cognitive context into the LLM through a Tree-Aware Attention mechanism and an EmotionGraphFormer, preserving the integrity of graph-structured information. Experiments on multiple datasets demonstrate that HyperEmo-RAG significantly outperforms existing methods.

情绪识别超球空间RAG多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。