arXiv:2603.14750cs.CV2026-03被引 1

用人脸特征精修情感边界,让弱监督视频情感定位更准。

Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization

  • 通过双分支融合人脸特征,捕捉细微情感线索。
  • 对比学习增强模型对标注点附近帧的情感边界识别能力。
  • 将稀疏标注转化为平滑伪标签,提升跨场景泛化性能。

点级弱监督视频情感定位(P-WTSL)旨在仅使用时间戳级情感标注,自动检测未剪辑多模态视频中的情感相关片段,显著降低帧级标注成本。针对现有方法在情感边界定位不准确的问题,本文提出面向面部引导的情感边界增强网络(FSENet),一个统一框架,利用细粒度人脸特征辅助情感定位。首先引入面部引导情感发现(FSD)模块,通过双分支建模将人脸特征融入多模态交互,有效提取情感刺激线索;随后提出点感知情感语义对比(PSSC)策略,利用对比学习区分标注点附近的候选帧情感语义,增强边界识别能力;最后设计边界感知情感伪标签生成(BSPG)方法,将稀疏点标注转换为时序平滑的监督伪标签。在基准数据集上的大量实验与可视化验证了该框架的有效性,在全监督、视频级和点级弱监督设置下均达到最优性能,展现出FSENet在不同标注条件下的强泛化能力。

原文摘要 · Abstract (English)

Point-level weakly-supervised temporal sentiment localization (P-WTSL) aims to detect sentiment-relevant segments in untrimmed multimodal videos using timestamp sentiment annotations, which greatly reduces the costly frame-level labeling. To further tackle the challenges of imprecise sentiment boundaries in P-WTSL, we propose the Face-guided Sentiment Boundary Enhancement Network (\textbf{FSENet}), a unified framework that leverages fine-grained facial features to guide sentiment localization. Specifically, our approach \textit{first} introduces the Face-guided Sentiment Discovery (FSD) module, which integrates facial features into multimodal interaction via dual-branch modeling for effective sentiment stimuli clues; We \textit{then} propose the Point-aware Sentiment Semantics Contrast (PSSC) strategy to discriminate sentiment semantics of candidate points (frame-level) near annotation points via contrastive learning, thereby enhancing the model's ability to recognize sentiment boundaries. At \textit{last}, we design the Boundary-aware Sentiment Pseudo-label Generation (BSPG) approach to convert sparse point annotations into temporally smooth supervisory pseudo-labels. Extensive experiments and visualizations on the benchmark demonstrate the effectiveness of our framework, achieving state-of-the-art performance under full supervision, video-level, and point-level weak supervision, thereby showcasing the strong generalization ability of our FSENet across different annotation settings.

情感定位弱监督人脸特征边界增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。