构建首个跨时长视频描述一致性评测基准,验证模型长期一致性和生成质量。
CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales

- 提出多场景、跨时长的视频描述评测基准,支持音视频与纯视觉双模式
- 现有模型在长视频中描述准确率与主体指代一致性显著下降
- 该基准与下游任务表现强相关,适合评估视频理解与生成模型
准确且全面的视频描述以及一致的主体指代对下游理解与生成任务至关重要。然而,现有评测基准难以在多样时长和场景下客观、全面地评估这些特性,制约了视频描述模型的发展。为此,我们提出CapRiCorn-1K,一个综合性评测基准,用于评估视频描述质量与跨长时间尺度的主体指代一致性。该基准支持音频视觉与纯视觉两种设置,以适应不同需求。在CapRiCorn-1K上的大量实验表明,当前模型在生成准确、完整描述的同时保持主体指代一致性方面普遍表现不佳,且随着视频时长增加,整体描述质量与指代一致性均明显下降。值得注意的是,我们的评估指标与基于生成描述的下游理解与生成任务性能高度相关,进一步验证了其有效性。项目开源地址:https://github.com/xlchen0205/CapRiCorn-1K。
原文摘要 · Abstract (English)
Accurate and comprehensive video captions with consistent subject references are critical for downstream understanding and generation tasks. However, few existing benchmarks can objectively and comprehensively evaluate these properties across diverse durations and scenarios, thereby hindering the advancement of video captioning models. To bridge this gap, we propose CapRiCorn-1K, a comprehensive benchmark designed to evaluate both video captioning quality and subject referential consistency across long temporal horizons and diverse video domains. To accommodate varied evaluation needs, our benchmark supports both audiovisual and visual-only settings. Extensive experiments on CapRiCorn-1K reveal that current models generally struggle to generate accurate and comprehensive captions while maintaining consistent subject references. Moreover, as video duration increases, both the overall caption quality and subject referential consistency decline. Notably, our evaluation metrics exhibit strong correlations with the performance of downstream understanding and generation tasks conditioned on the generated captions, further validating their effectiveness. The project is available at https://github.com/xlchen0205/CapRiCorn-1K .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。