arXiv:2606.24636cs.AI2026-06被引 2

用时空锚点结构化推理,让AI精准描述电影拍摄手法。

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning

论文配图:CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning
图 1 · 摘自论文原文
  • 引入时空锚点,将镜头语言分解为可追踪的视觉单元
  • 在472对视频-字幕上超越现有模型,提升描述完整与准确
  • 适合影视生成、视频理解研究者使用

电影化字幕生成旨在用专业影视术语(如运镜、景别、景深、构图、拍摄角度)描述视频的拍摄方式,对精细视频理解与可控电影级视频生成至关重要,但现有多模态大模型对此探索不足。不同于基于问答的评估,该任务需要对多个影视维度进行统一的开放式描述。挑战在于:模型需从细微视觉线索中推断专业概念,且生成内容须全面准确。为此,我们提出CineCap框架,结合结构化推理与时空锚点,以及以完整性、准确性与门控覆盖为奖励的强化学习。前者将专业描述锚定于明确视觉证据,组织为紧凑原子推理单元用于监督微调;后者优化描述完整度与事实正确性的平衡。此外,我们构建了包含472组人工标注视频-字幕对的CineCap Bench基准。大量实验表明,CineCap持续优于主流专有与开源基线,建立了电影化字幕生成的新SOTA。代码、模型与数据集已公开于https://github.com/Hectormxy/CineCap.git。

原文摘要 · Abstract (English)

Cinematographic captioning aims to describe how a video is filmed using professional film-language concepts such as camera movement, shot size, depth of field, composition, and shooting angle. This capability is important for fine-grained video understanding and controllable movie-quality video generation, yet remains underexplored in existing multimodal large language models. Unlike question-answering-based evaluation of cinematic understanding, cinematographic captioning requires a unified open-form description over multiple cinematographic dimensions. This task is challenging for two main reasons: the model must infer professional cinematographic concepts from subtle visual evidence, and it must generate captions that are both comprehensive and accurate. Accordingly, we propose CineCap, a framework that combines structured reasoning with spatio-temporal anchors and reinforcement learning with comprehensiveness, accuracy, and gated coverage rewards. The former grounds professional cinematographic descriptions in explicit visual evidence and organizes them into compact atomic reasoning for supervised fine-tuning, while the latter improves the balance between descriptive completeness and factual correctness. In addition, we construct CineCap Bench, a benchmark of 472 manually annotated video-caption pairs for systematic evaluation. Extensive experiments show that CineCap consistently outperforms strong proprietary and open-source baselines, establishing a new state of the art for cinematographic captioning. The code, model checkpoint, and benchmark are publicly available in https://github.com/Hectormxy/CineCap.git.

视频字幕电影语言结构化推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。