arXiv:2603.22306cs.AI2026-03被引 1

构建可持续更新的情感记忆系统,让AI理解情绪的长期上下文。

Memory Bear AI Memory Science Engine for Multimodal Affective Intelligence: A Technical Report

  • 将多模态情绪信息建模为可演化的情感记忆单元
  • 在噪声和缺模态下仍保持高准确率,优于现有系统
  • 适合需要连续情感理解的实际应用场景

真实交互中的情感判断很少是局部预测问题。情绪意义常依赖先前轨迹、累积上下文及多模态证据,而这些证据在当前可能弱、嘈杂或不完整。尽管多模态情绪识别(MER)已提升文本、语音与视觉信号融合,但多数系统仍针对短时推理优化,对持久情感记忆、长时依赖建模及输入不全时的鲁棒解释支持有限。本技术报告提出 Memory Bear AI 情感记忆科学引擎,一种以记忆为核心的多模态情感智能框架。该框架不将情绪视为瞬时输出标签,而是将其建模为记忆系统中结构化且动态演化的变量。通过结构化记忆形成、工作记忆聚合、长期固化、记忆驱动检索、动态融合校准和持续记忆更新,实现情感信息的保存、重激活与修正。多模态信号被转换为结构化情感记忆单元(EMUs),支持跨交互周期的情感动态追踪。实验显示,在基准与业务场景中均显著优于对比系统,尤其在噪声或缺模态条件下表现更稳健。该框架为从局部情绪识别迈向更连续、鲁棒且可部署的情感智能提供了实用路径。

原文摘要 · Abstract (English)

Affective judgment in real interaction is rarely a purely local prediction problem. Emotional meaning often depends on prior trajectory, accumulated context, and multimodal evidence that may be weak, noisy, or incomplete at the current moment. Although multimodal emotion recognition (MER) has improved the integration of text, speech, and visual signals, many existing systems remain optimized for short-range inference and provide limited support for persistent affective memory, long-horizon dependency modeling, and robust interpretation under imperfect input. This technical report presents the Memory Bear AI Memory Science Engine, a memory-centered framework for multimodal affective intelligence. Instead of treating emotion as a transient output label, the framework models affective information as a structured and evolving variable within a memory system. It organizes processing through structured memory formation, working-memory aggregation, long-term consolidation, memory-driven retrieval, dynamic fusion calibration, and continuous memory updating. At its core, multimodal signals are transformed into structured Emotion Memory Units (EMUs), enabling affective information to be preserved, reactivated, and revised across interaction horizons. Experimental results show consistent gains over comparison systems across benchmark and business-grounded settings, with stronger accuracy and robustness, especially under noisy or missing-modality conditions. The framework offers a practical step from local emotion recognition toward more continuous, robust, and deployment-relevant affective intelligence.

情感计算记忆机制多模态鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。