arXiv:2605.25409cs.CV2026-05中稿 · the Workshop on Af…被引 1

提出首个精准定位视频中笑声的轻量级模型与两个新数据集。

MTLLFM: Multimodal-Temporal Laughter Localization: UR-FUNNY-Temporal and SMILE-Temporal Benchmarks with an Adaptive Multimodal Fusion Model

论文配图:MTLLFM: Multimodal-Temporal Laughter Localization: UR-FUNNY-Temporal and SMILE-Temporal Benchmarks with an Adaptive Multimodal Fusion Model
图 1 · 摘自论文原文
  • 用固定编码器+自适应模态门控,仅靠片段标签实现帧级笑声定位。
  • 在体育直播数据上达99%准确率,比Gemini 3 Flash提升显著。
  • 适合做情感分析、内容理解的开发者与研究者使用。

视频中笑声检测对情感计算与叙事理解至关重要,但现有方法多为粗粒度片段分类,难以捕捉短暂笑声的精确时间边界。本文提出两项互补贡献:一是构建了UR-FUNNY-Temporal和SMILE-Temporal两个全标注时序笑声数据集,覆盖超过11,053个视频(78.8小时),提供每段笑声的起止边界,并标注发言者/观众笑声、模态主导性(声学、视觉或两者)及强度等级;二是设计一种轻量级弱监督时序笑声定位框架,结合HuBERT与MAE固定编码器、时序Softmax池化与自适应模态门控,仅使用片段级标签即可学习细粒度时间定位,无需帧级标注。跨三个数据集实验表明,该方法显著优于多模态基座模型如Gemini 3 Flash,体育广播数据上达到99% F1和68.1%定位精度。消融实验证明各组件有效性。精确时间标签使下游笑声推理性能提升227%(CIDEr),令GPT-3.5超越GPT-4o。代码与数据集已开源。

原文摘要 · Abstract (English)

Detecting laughter in video is essential for affective computing and narrative understanding, yet existing approaches treat it as coarse clip-level classification, failing to capture precise temporal boundaries of brief, transient laughter events. We address this gap with two complementary contributions. First, we introduce UR-FUNNY-Temporal and SMILE-Temporal, fully annotated temporal laughter datasets extending two widely-used humor benchmarks. Our annotations cover over 11,053 videos (78.8 hours) and provide precise onset/offset boundaries for each laughter event, along with rich metadata distinguishing speaker vs. audience laughter, modality dominance (acoustic, visual, or both), and intensity levels. Second, we propose a lightweight weakly-supervised framework for temporal laughter localization. Our architecture combines fixed HuBERT and MAE encoders with temporal softmax pooling and adaptive modality gating, learning fine-grained temporal grounding from clip-level labels without requiring frame-level annotations during training. Experiments across three datasets demonstrate that our approach substantially outperforms multimodal foundation models including Gemini 3 Flash, achieving 99% F1 and 68.1% localization precision on sports broadcast data. Ablations validate each architectural component. Furthermore, our precise temporal tags improve downstream laughter reasoning by 227% on CIDEr, enabling GPT-3.5 to outperform GPT-4o. The code, UR-FUNNY-Temporal and SMILE-Temporal datasets are publicly available at https://github.com/WSCSports/MTLLFM-temporal-laughter-localization.

笑声检测时序定位多模态融合弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。