提出多视角视频高光检测基准,从事件、情感、本质三维度提升真实场景识别能力
TRINITY: A Multi-Perspective Benchmark for Personal-Style Video Highlight Detection

- 将高光分为事件、情感、本质三维度,统一时间框架下建模
- 在Mr. HiSum和YouTube Highlights上分别提升7.15和10.82 mAP
- 适合做个性化视频剪辑、多模态理解的研究者
传统视频高光检测依赖狭隘的事件中心型显著性定义,在非受限个人视频中常失效,因高光具有异质性和视角依赖性。为此,我们提出TRINITY,一个将高光显著性分解为事件、情感与本质三个互补维度的多视角基准,构建于统一时间框架内。基于此多面视角,我们设计共享主干多分支架构,通过视图特定专家实现并行多视角预测。大量实验表明,该方法显著优于现有基线,在Mr. HiSum上达+7.15/+3.62 mAP(rho=15%/50%),在YouTube Highlights上达+10.82 mAP。结果验证了多视角建模在复杂现实场景中更鲁棒、更全面。基准与代码将在接受后发布,地址为https://huggingface.co/datasets/vanilladucky/TRINITY和https://github.com/vanilladucky/TRINITY。
原文摘要 · Abstract (English)
Traditional video highlight detection relies on a narrow, event-centric definition of saliency, which often fails to generalize to unconstrained personal videos where highlights are heterogeneous and perspective-dependent. To address this, we introduce TRINITY, a multi-perspective benchmark that decomposes highlight saliency into three complementary dimensions, Event, Emotion, and Nature, within a unified temporal framework. Leveraging this multi-faceted view, we propose a shared-backbone multi-branch architecture designed for parallel multi-perspective prediction via view-specific experts. Comprehensive experiments demonstrate that our method significantly outperforms state-of-the-art baselines, achieving gains of +7.15/+3.62 mAP (rho=15%/50%) on Mr. HiSum and +10.82 mAP on YouTube Highlights. These results validate that multi-perspective modeling provides a more robust and comprehensive formulation of video saliency, especially for complex real-world scenarios. The benchmark and relevant codes will be released upon acceptance. The benchmark is available at https://huggingface.co/datasets/vanilladucky/TRINITY and the code is available at https://github.com/vanilladucky/TRINITY.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。