arXiv:2605.30090cs.CLcs.CV2026-05被引 3

针对长视频生成的多智能体诊断评估框架,精准定位生成瓶颈。

DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation

论文配图:DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation
图 1 · 摘自论文原文
  • 用80项元数据+7个用户画像+40个检查点多维度评估
  • 发现片段过渡质量仅0.256,远低于提示满足度0.71
  • 支持个性化评价,揭示传统评分忽略的失败模式

长视频生成正从短场景合成转向分钟级、多镜头、具叙事结构的复杂创作,包含电影化控制、音频与跨模态同步。然而,现有评估基准仍局限于局部视觉质量、短时序一致性或通用提示对齐,难以诊断流程缺陷和用户偏好差异。本文提出DirectorBench,一个面向长视频生成的个性化多智能体诊断基准。该框架基于80项结构化元数据、7个用户画像和40个检查点,从剧本、视觉、音频、跨模态及稳定性五个维度进行评估。不同于单一综合得分,DirectorBench可定位到检查点级别的瓶颈,并支持用户画像感知的评价。我们评估了4种长视频生成工作流、6个基础大语言模型及7个用户画像。结果显示,各单元间过渡质量平均仅为0.256,最优工作流达0.356,而提示级用户需求满足率平均为0.71。通过14名标注者的人类评估验证,DirectorBench能有效捕捉人类感知的质量差异,并揭示聚合评分所掩盖的工作流与用户画像相关的故障模式。这凸显了诊断性与个性化基准在长视频生成中的关键作用。

原文摘要 · Abstract (English)

Long-form video generation is rapidly moving from short, single-scene synthesis toward minute-long, multi-shot creation with narrative structure, cinematic control, audio, and cross-modal synchronization. However, evaluating such videos remains challenging, since existing benchmarks largely focus on local visual quality, short-horizon temporal consistency, or generic prompt alignment, and provide limited diagnosis of workflow failures and user-dependent preferences. We introduce DirectorBench, a personalized multi-agent diagnostic benchmark for long-form video generation. DirectorBench evaluates generated videos with respect to 80 structured metadata entries, 7 user profiles, and 40 checkpoint criteria across 5 dimensions: script, visual, audio, cross-modal, and stability. Instead of reducing quality to a single aggregate score, DirectorBench localizes checkpoint-level bottlenecks and supports profile-aware evaluation. We evaluate 4 long-form video generation workflows, 6 base LLMs, and 7 user profiles. Across workflows, DirectorBench reveals a between-unit bottleneck: transition quality averages only 0.256 and reaches 0.356 for the best workflow, while prompt-level user demand fulfillment averages 0.71. We further conduct human evaluation with 14 annotators to validate the alignment between DirectorBench and human judgment. The results show that DirectorBench captures human-perceptible quality differences and reveals workflow- and profile-dependent failure modes that are hidden by aggregate scoring. These findings highlight the importance of diagnostic and profile-aware benchmarking for long-form video generation.

视频生成评估基准多智能体个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。