arXiv:2604.15823cs.CV2026-04

为机器人看视频时的情绪理解建立新数据集和模型,提升真实场景下的表现。

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions

论文配图:Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions
图 1 · 摘自论文原文
  • 构建首个面向第一视角屏幕观影情绪理解的数据集ESE。
  • 在真实观看场景下,模型性能比传统方法下降超40%(27.99→16.69)。
  • 适合研究具身智能、跨域情绪理解与多模态推理的学者使用。

具身机器人常通过第一视角屏幕界面观看电影,而非原始影视画面,导致视角扭曲、尺度变化、光照差异及环境干扰等域偏移问题。然而,现有电影情绪理解研究几乎全部基于影视画面,难以泛化到真实观看场景。为此,我们提出EgoScreen-Emotion(ESE),首个面向第一视角屏幕观影情绪理解的基准数据集。ESE包含224条在受控条件下拍摄的电影预告片,共生成28,667个时间对齐的关键帧,由多位评分者以可信度感知的多标签协议标注,以应对情绪模糊性。我们进一步构建了融合时序视觉证据、叙事摘要、压缩历史上下文与音频线索的多模态长程推理框架。跨域实验表明:在真实第一视角屏幕观测下,仅在影视画面训练的模型宏平均F1从27.99降至16.69。在ESE上训练显著提升了模型鲁棒性。本方法在性能上媲美强闭源多模态模型,凸显领域特定数据与长程多模态推理的重要性。

原文摘要 · Abstract (English)

Embodied robotic agents often perceive movies through an egocentric screen-view interface rather than native cinematic footage, introducing domain shifts such as viewpoint distortion, scale variation, illumination changes, and environmental interference. However, existing research on movie emotion understanding is almost exclusively conducted on cinematic footage, limiting cross-domain generalization to real-world viewing scenarios. To bridge this gap, we introduce EgoScreen-Emotion (ESE), the first benchmark dataset for egocentric screen-view movie emotion understanding. ESE contains 224 movie trailers captured under controlled egocentric screen-view conditions, producing 28,667 temporally aligned key-frames annotated by multiple raters with a confidence-aware multi-label protocol to address emotional ambiguity. We further build a multimodal long-context emotion reasoning framework that models temporal visual evidence, narrative summaries, compressed historical context, and audio cues. Cross-domain experiments reveal a severe domain gap: models trained on cinematic footage drop from 27.99 to 16.69 Macro-F1 when evaluated on realistic egocentric screen-view observations. Training on ESE substantially improves robustness under realistic viewing conditions. Our approach achieves competitive performance compared with strong closed-source multimodal models, highlighting the importance of domain-specific data and long-context multimodal reasoning.

情绪理解具身智能多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。