让视频生成智能体学会理解电影结构,提升创作质量。
AVA-Encoder: Towards Agent-Native Video Representation Learning

- 用知识图谱结构化视频内容,便于智能体理解和操作。
- 相比最强基线,视频重建准确率提升73.1%。
- 适合研究智能体视频理解与生成的开发者使用。
视频创作智能体仍缺乏从高质量影视作品中学习的有效方法,限制了其生成电影级视频的能力。核心挑战在于缺少既忠实于影片内容、又可直接用于智能体推理与操作的结构化视频表征。为此,我们提出面向智能体的视频自编码框架AVA-Encoder,通过智能体自进化驱动,将视频转化为电影知识图谱(Film KG)表示并实现重建。该知识图谱显式捕捉实体、事件、资产及其多模态关系,结构清晰,便于智能体查询与操控。重建残差驱动双环文本梯度优化框架,协同优化知识图谱与智能体视频编码器。大量实验表明,AVA-Encoder相较最强外部基线取得20.7个百分点的绝对提升(相对提升73.1%)。在仅策略控制设置下,其伪训练得到的智能体编码器策略也优于精心人工调优的策略,同时减少74.3%的帧级和70.1%的关键帧级系统提示词。我们开源了完整的AVA-Encoder框架、可靠的智能体视频重建基准,以及首个高质量电影知识图谱数据集。
原文摘要 · Abstract (English)
Video creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both faithful to film content and directly usable for agentic reasoning and manipulation. To address the challenge, we propose the Agentic Video Auto-Encoder (AVA-Encoder), a novel auto-encoding framework driven by agentic self-evolution to learn agent-native video representations. AVA-Encoder transforms a video into a Film Knowledge Graph (KG) representation and then reconstructs it back into video. This Film KG representation explicitly captures entities, events, assets, and their multimodal relationships in a structured form that can be easily understood, queried, and manipulated by agents. The reconstruction residual drives a dual-loop textual-gradient optimization framework that jointly improves the Film KG representation and the Agentic Video Encoder. Extensive experiments show that AVA-Encoder achieves a 20.7-percentage-point absolute gain, or a 73.1% relative improvement, over the strongest external baseline. In the controlled policy-only setting, its pseudo-trained Agentic Video Encoder policy also outperforms a carefully human-tuned policy while using 74.3% fewer shot-level and 70.1% fewer keyframe-level system-prompt tokens. We release the complete AVA-Encoder framework, a reliable agentic video reconstruction benchmark, and the first dataset of high-quality Film KG representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。