arXiv:2512.10975cs.LGcs.AI2025-12

用多个智能体模块化训练多模态情绪识别,提升灵活性与效率

Agent-Based Modular Learning for Multimodal Emotion Recognition in Human-Agent Systems

  • 每个模态和融合分类器作为独立智能体,由中央协调
  • 支持新模态接入与旧组件替换,训练计算开销更低
  • 适合需要持续更新感知能力的虚拟/实体交互系统

高效人机交互依赖对人类情绪状态的准确与自适应感知。尽管融合面部表情、语音和文本线索的多模态深度学习模型在情绪识别中表现优异,但其训练与维护常面临计算成本高、模态变化不灵活的问题。本文提出一种新型多智能体框架,用于训练多模态情绪识别系统:各模态编码器与融合分类器作为自主智能体,由中央监督者协同管理。该架构支持新模态(如通过emotion2vec提取的音频特征)的模块化集成、过时组件的无缝替换,并降低训练期间的计算负担。我们通过一个包含视觉、音频与文本模态的原型实现验证,其中分类器作为共享决策智能体。实验表明,该框架不仅提升了训练效率,还为具身与虚拟代理在人机交互场景中的感知模块设计提供了更灵活、可扩展、易维护的解决方案。

原文摘要 · Abstract (English)

Effective human-agent interaction (HAI) relies on accurate and adaptive perception of human emotional states. While multimodal deep learning models - leveraging facial expressions, speech, and textual cues - offer high accuracy in emotion recognition, their training and maintenance are often computationally intensive and inflexible to modality changes. In this work, we propose a novel multi-agent framework for training multimodal emotion recognition systems, where each modality encoder and the fusion classifier operate as autonomous agents coordinated by a central supervisor. This architecture enables modular integration of new modalities (e.g., audio features via emotion2vec), seamless replacement of outdated components, and reduced computational overhead during training. We demonstrate the feasibility of our approach through a proof-of-concept implementation supporting vision, audio, and text modalities, with the classifier serving as a shared decision-making agent. Our framework not only improves training efficiency but also contributes to the design of more flexible, scalable, and maintainable perception modules for embodied and virtual agents in HAI scenarios.

情绪识别多模态智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。