arXiv:2605.06924cs.CVcs.AI2026-05被引 5

A²RD通过自迭代生成提升长视频一致性,解决语义漂移问题。

A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency

论文配图:A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency
图 1 · 摘自论文原文
  • 分段生成并循环自优化,结合记忆与自修正机制
  • 在1-10分钟视频上一致性和连贯性提升最高30%和20%
  • 适合需要长视频生成与稳定性的研究者和开发者

长视频的合成与连贯性仍面临根本挑战。现有方法在长时序下易出现语义漂移与叙事崩溃。本文提出A²RD,一种代理式自回归扩散架构,将创意生成与一致性保障解耦。A²RD将长视频合成建模为闭环过程,通过检索-生成-修正-更新循环实现逐段自优化。其核心包含三部分:(i) 多模态视频记忆,跨模态追踪视频进展;(ii) 自适应分段生成,根据自然演进切换生成模式以保证视觉一致性;(iii) 分层测试时自改进,于帧级与视频级自我修正,防止错误传播。我们进一步引入LVBench-C,一个包含非线性实体与环境变化的挑战性基准。在公开数据集及LVBench-C上,覆盖1至10分钟视频,A²RD在一致性上较最先进基线提升最高达30%,叙事连贯性提升20%。人工评估验证了这些提升,并指出运动与过渡平滑性显著改善。

原文摘要 · Abstract (English)

Synthesizing consistent and coherent long video remains a fundamental challenge. Existing methods suffer from semantic drift and narrative collapse over long horizons. We present A$^2$RD, an Agentic Auto-Regressive Diffusion architecture that decouples creative synthesis from consistency enforcement. A$^2$RD formulates long video synthesis as a closed-loop process that synthesizes and self-improves video segment-by-segment through a Retrieve--Synthesize--Refine--Update cycle. It comprises three core components: (i) Multimodal Video Memory that tracks video progression across modalities; (ii) Adaptive Segment Generation that switches among generation modes for natural progression and visual consistency; and (iii) Hierarchical Test-Time Self-Improvement that self-improves each segment at frame and video levels to prevent error propagation. We further introduce LVBench-C, a challenging benchmark with non-linear entity and environment transitions to stress-test long-horizon consistency. Across public and LVBench-C benchmarks spanning one- to ten-minute videos, A$^2$RD outperforms state-of-the-art baselines by up to 30% in consistency and 20% in narrative coherence. Human evaluations corroborate these gains while also highlighting notable improvements in motion and transition smoothness.

长视频生成扩散模型自迭代一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。