首个聚焦人物叙事连贯性的多镜头视频评测基准,解决角色状态断层问题。
PersonaShot: Benchmarking Person-Centric Narrative Continuity in Multi-Shot Video Generation

- 构建三层次连贯性评估:单镜头内、跨镜头间、整体叙事轨迹。
- 16项指标覆盖物理、情感与电影语法,发现顶级模型仍存状态重置与情绪突变。
- 基于人类判断训练轻量评估器,适合研究生成连贯叙事的视频模型者。
视频生成正从单镜头片段转向多镜头叙事,以人物为核心叙事锚点。然而现有评测多关注角色外观或单镜头质量,缺乏对跨镜头物理与情感状态一致性的衡量,且极少提供针对性评估方法——物理连续性、面部动态与电影关系需不同视觉、时间与关系证据。为此,我们提出PersonaShot,首个以人物为中心的多镜头视频叙事连贯性评测基准。包含约1000个多镜头片段和16项指标,涵盖物理连续性、情感动态与电影语法。第一,建立三层次连贯性评估:单镜头内状态、跨镜头过渡、序列级轨迹。第二,设计人对齐专用评估器,将大模型多模态推理提炼为轻量化判别器,每类对应视觉、时间或关系证据,并与专家判断对齐。第三,系统评估揭示当前先进模型具备不同能力特征,且感知质量与跨镜头连贯性间存在明显差距:即使视觉效果出色,仍常出现状态重置、情绪突变与电影关系断裂。人类实验进一步验证评估器与专家判断高度一致。
原文摘要 · Abstract (English)
Video generation is rapidly evolving from single-shot clips to multi-shot narratives, where the human character serves as the core narrative anchor. However, existing benchmarks mainly assess character appearance or individual-shot quality, without measuring whether physical and emotional states remain coherent across cuts. They also rarely provide criterion-specific evaluation methods, although physical continuity, facial dynamics, and cinematic relations require different visual, temporal, and relational evidence. To address these limitations, we introduce PersonaShot, the first person-centric benchmark for narrative continuity in multi-shot video generation. PersonaShot contains approximately 1,000 multi-shot segments and 16 metrics spanning physical continuity, affective dynamics, and cinematic grammar. \textbf{\textit{1)} Narrative Continuity Benchmark:} We evaluate character coherence across three temporal levels: within-shot states, cross-shot transitions, and sequence-level trajectories. \textbf{\textit{2)} Human-Aligned Specialist Evaluators:} We distill reasoning from a large multimodal teacher into lightweight criterion-specific evaluators, each grounded in the visual, temporal, or relational evidence required by its metric, and align them with expert human judgments. \textbf{\textit{3)} Systematic Evaluation and Insights:} Our evaluation reveals distinct capability profiles across state-of-the-art models and a clear gap between perceptual quality and cross-shot narrative continuity. Even visually compelling videos frequently exhibit physical-state resets, abrupt affective shifts, and broken cinematic relations across shots. Human studies further demonstrate strong agreement between our evaluators and expert judgments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。