arXiv:2602.15381cs.IRcs.CV2026-02AAAI

自动从长片电影中提取搞笑片段,提升短视频内容创作效率

Automatic Funny Scene Extraction from Long-form Cinematic Videos

  • 融合视觉与文本线索进行场景分割,提升定位精度
  • 在OVSD数据集上比现有方法高18.3%的场景检测准确率
  • 适合影视平台做预告片剪辑与爆款内容生成

从长篇影视作品中自动提取吸引人且高质量的幽默片段,对制作吸睛视频预告和短平快内容至关重要,有助于提升流媒体平台用户参与度。长时长与复杂叙事给场景定位带来挑战,而幽默依赖多模态信息且风格微妙,进一步增加难度。本文提出端到端系统,实现长片电影中幽默片段的自动识别与排序,包含镜头检测、多模态场景定位与针对影视内容优化的幽默标记。关键创新包括结合视觉与文本线索的新型场景分割方法、通过引导三元组挖掘改进镜头表示,以及融合音频与文本的多模态幽默标记框架。系统在OVSD数据集上较现有最优方法提升18.3% AP,幽默检测F1得分为0.834。五部影视作品的评估显示,87%的提取片段确为幽默设计,98%场景定位准确。在预告片上的良好泛化能力表明该流程可显著提升内容创作效率,增强用户参与度,简化多样化影视格式的短内容生成。

原文摘要 · Abstract (English)

Automatically extracting engaging and high-quality humorous scenes from cinematic titles is pivotal for creating captivating video previews and snackable content, boosting user engagement on streaming platforms. Long-form cinematic titles, with their extended duration and complex narratives, challenge scene localization, while humor's reliance on diverse modalities and its nuanced style add further complexity. This paper introduces an end-to-end system for automatically identifying and ranking humorous scenes from long-form cinematic titles, featuring shot detection, multimodal scene localization, and humor tagging optimized for cinematic content. Key innovations include a novel scene segmentation approach combining visual and textual cues, improved shot representations via guided triplet mining, and a multimodal humor tagging framework leveraging both audio and text. Our system achieves an 18.3% AP improvement over state-of-the-art scene detection on the OVSD dataset and an F1 score of 0.834 for detecting humor in long text. Extensive evaluations across five cinematic titles demonstrate 87% of clips extracted by our pipeline are intended to be funny, while 98% of scenes are accurately localized. With successful generalization to trailers, these results showcase the pipeline's potential to enhance content creation workflows, improve user engagement, and streamline snackable content generation for diverse cinematic media formats.

视频生成多模态幽默检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。