arXiv:2512.19379cs.LGcs.AI2025-12

首个印尼语多模态情感识别数据集,提升低资源语言情感分析性能

OmniMER: Auxiliary-Enhanced LLM Adaptation for Indonesian Multimodal Emotion Recognition

  • 通过文本、音频、视频三类辅助任务增强大模型对情绪线索的感知
  • 在印尼语数据集上实现0.582宏平均F1(情感分类)和0.454(情绪识别)
  • 适用于多模态情感分析、低资源语言研究及跨语言迁移任务

印度尼西亚语(超过2亿人使用)虽在东南亚社交媒体中占主导地位,但在多模态情感识别研究中仍被忽视。我们提出IndoMER,首个面向印尼语的多模态情感识别基准,包含来自203名说话者的1,944段视频片段,涵盖七类情绪,具有时间对齐的文本、音频和视觉标注。该数据集呈现真实挑战,包括跨模态不一致与受印尼文化沟通习惯影响的长尾分布。为此,我们提出OmniMER框架,基于Qwen2.5-Omni,引入三种模态特定的辅助感知任务:文本的情感关键词提取、视频的面部表情分析、音频的韵律分析。这些任务使模型在融合前更准确识别各模态中的情绪线索,减少低资源场景下的虚假关联依赖。在IndoMER上的实验显示,OmniMER在情感分类上达到0.582宏平均F1,情绪识别达0.454,分别优于基础模型7.6和22.1个绝对点。跨语言评估在中文CH-SIMS数据集上进一步验证了框架的泛化能力。数据集与代码已公开。

原文摘要 · Abstract (English)

Indonesian, spoken by over 200 million people, remains underserved in multimodal emotion recognition research despite its dominant presence on Southeast Asian social media platforms. We introduce IndoMER, the first multimodal emotion recognition benchmark for Indonesian, comprising 1,944 video segments from 203 speakers with temporally aligned text, audio, and visual annotations across seven emotion categories. The dataset exhibits realistic challenges including cross-modal inconsistency and long-tailed class distributions shaped by Indonesian cultural communication norms. To address these challenges, we propose OmniMER, a multimodal adaptation framework built upon Qwen2.5-Omni that enhances emotion recognition through three auxiliary modality-specific perception tasks: emotion keyword extraction for text, facial expression analysis for video, and prosody analysis for audio. These auxiliary tasks help the model identify emotion-relevant cues in each modality before fusion, reducing reliance on spurious correlations in low-resource settings. Experiments on IndoMER show that OmniMER achieves 0.582 Macro-F1 on sentiment classification and 0.454 on emotion recognition, outperforming the base model by 7.6 and 22.1 absolute points respectively. Cross-lingual evaluation on the Chinese CH-SIMS dataset further demonstrates the generalizability of the proposed framework. The dataset and code are publicly available. https://github.com/yanxm01/INDOMER

多模态情感识别低资源语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。