提出隐私保护的情感识别新任务,用去身份化音视频实现安全情绪理解。
DEEMO: De-identity Multimodal Emotion Recognition and Reasoning

- 构建去身份化多模态情感识别与推理框架,避免依赖人脸和语音身份信息。
- 在去身份情感识别中达到74.49%准确率和74.45% F1值,推理任务得分显著领先。
- 适合关注隐私保护、伦理AI与多模态认知的科研人员和应用开发者。
情感理解是关键但具有挑战性的任务。现有方法严重依赖身份敏感信息(如面部表情和语音),引发个人隐私担忧。为此,我们提出去身份多模态情感识别与推理(DEEMO)新任务,支持使用去标识化的音视频输入进行情感理解。DEEMO数据集包含两个子集:DEEMO-NFBL,包含丰富的非面部身体语言标注;DEEMO-MER,用于基于无身份线索的多模态情感识别与推理的指令数据集。该设计可在不泄露身份的前提下实现情感理解。此外,我们提出DEEMO-LLaMA,一种融合去标识化音频、视频与文本信息的多模态大语言模型,显著提升情感识别与推理能力。大量实验表明,DEEMO-LLaMA在两项任务上均达到当前最优性能,去身份情感识别准确率达74.49%,F1得分为74.45%;去身份情感推理中线索重叠达6.20,标签重叠达7.66。本工作推动了伦理AI发展,助力隐私保护的情感计算进步。
原文摘要 · Abstract (English)
Emotion understanding is a critical yet challenging task. Most existing approaches rely heavily on identity-sensitive information, such as facial expressions and speech, which raises concerns about personal privacy. To address this, we introduce the De-identity Multimodal Emotion Recognition and Reasoning (DEEMO), a novel task designed to enable emotion understanding using de-identified video and audio inputs. The DEEMO dataset consists of two subsets: DEEMO-NFBL, which includes rich annotations of Non-Facial Body Language (NFBL), and DEEMO-MER, an instruction dataset for Multimodal Emotion Recognition and Reasoning using identity-free cues. This design supports emotion understanding without compromising identity privacy. In addition, we propose DEEMO-LLaMA, a Multimodal Large Language Model (MLLM) that integrates de-identified audio, video, and textual information to enhance both emotion recognition and reasoning. Extensive experiments show that DEEMO-LLaMA achieves state-of-the-art performance on both tasks, outperforming existing MLLMs by a significant margin, achieving 74.49% accuracy and 74.45% F1-score in de-identity emotion recognition, and 6.20 clue overlap and 7.66 label overlap in de-identity emotion reasoning. Our work contributes to ethical AI by advancing privacy-preserving emotion understanding and promoting responsible affective computing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。