让主体固定,让场景自由:用查询感知路由提升长视频生成的动态表现
Tether the Subject, Release the Scene: Query-Aware Memory Routing for Long-Horizon Autoregressive Video Generation

- 分离主体与场景查询,按区域和历史新旧动态控制记忆访问
- 在10名标注者2400次盲测中,整体质量和场景推进评分均领先其他8种基线
- 无需训练,可适配冻结模型,适合需要持续变化背景的长视频生成任务
流式自回归视频模型通过历史记忆逐块生成长视频,现有方法通常对主体和场景查询采用相似的记忆访问策略,虽能稳定主体,却可能使背景、视角和场景结构被早期生成状态锁定,导致局部运动持续但场景停滞。这种现象称为记忆锚定的场景进展不足,仅靠一致性和运动指标难以发现。本文提出TetherMem,一种无需训练、基于查询感知的时空记忆路由机制,用于冻结的视频生成模型。TetherMem将主体与场景查询分离,并引入区域和年龄条件的先验,使主体查询保留身份相关记忆,而场景查询减少对主体历史和陈旧背景的依赖。在10名标注者进行的2400次盲测中,TetherMem在整体质量(0.780)和场景推进(0.769)上均获得最高估计偏好得分。在完整的30秒视频中,它能持续改变背景、视角和场景状态,同时保持主体可识别性和时间连续性。
原文摘要 · Abstract (English)
Streaming autoregressive video models generate long videos chunk by chunk, using historical memory to maintain consistency. Existing methods typically expose subject and scene queries to history through similar policies. This stabilizes the subject, but can also lock backgrounds, viewpoints, and scene structure to previously generated states even when local motion continues. We call this failure memory-anchored scene under-progression; consistency and motion metrics alone can miss it. We introduce TetherMem, a training-free, query-aware spatiotemporal memory router for frozen video generators. TetherMem separates subject and scene queries and modulates historical access with region- and age-conditioned priors: subject queries retain identity-bearing history, while scene queries reduce reliance on subject history and stale backgrounds. Across 2,400 blinded pairwise judgments from 10 annotators, TetherMem achieves the highest estimated expected preference among eight streaming long-video baselines for overall quality (0.780) and scene progression (0.769). On complete 30-second videos, it sustains changes in background, viewpoint, and scene state while preserving subject recognizability and temporal continuity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。