arXiv:2504.07660cs.CV2025-04中稿 · paper

统一表情检测与识别,端到端提升长视频表现

End-to-End Facial Expression Detection in Long Videos

  • 将表情定位与分类合并为单一端到端任务
  • 在三个数据集上超越基线,提升定位与识别效果
  • 发现专家标注与自报情绪标签差异,提示标注需改进

表情检测需同时确定表情发生时间及情感类别。现有方法通常分步处理,限制了真实场景下的性能与鲁棒性。本文提出FEDN,一个统一的端到端表情检测网络,将定位与识别整合为单一任务。FEDN引入两个时序注意力模块:段级注意力捕捉细粒度局部动态,滑动窗口注意力捕捉更广义的时间上下文。二者输出通过多尺度时序特征金字塔融合,实现对不同持续时间表情的精准定位。该统一框架支持任务间联合优化与共享表征学习。FEDN在三个公开基准上均优于强基线,在定位与识别任务中表现优异。此外,研究揭示了专家标注与自报情绪标签间此前未被注意的差异,凸显表情标注中的关键挑战,推动更精细标注协议的发展。

原文摘要 · Abstract (English)

Facial expression detection requires spotting when expressions occur and recognizing which emotional category they belong to. Despite their close relationships, existing approaches typically address these tasks separately, limiting performance and robustness in real-world settings. In this work, we propose FEDN, a Facial Expression Detection Network, which unifies spotting and recognition into a single detection task performed fully end-to-end. FEDN introduces two temporal attention modules, segment-level attention to capture fine-grained local dynamics and sliding window attention to capture the broader temporal context. Their output is combined in a multi-scale temporal feature pyramid, which enables spotting of expressions with varying duration. This unified framework enables joint optimization and shared representation learning across tasks. FEDN outperforms strong baselines in both spotting and detection on three public benchmarks, demonstrating the effectiveness of unifying spotting and recognition across multiple temporal scales. Additionally, we uncover a previously unreported discrepancy between expert-annotated and self-reported emotion labels, highlighting a key challenge in expression benchmarking and motivating the development of more nuanced annotation protocols.

表情检测端到端时序建模视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。