arXiv:2608.23383cs.CV2026-08

让视频生成能持续讲长故事、支持交互,角色和声音不变

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

论文配图:Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
图 1 · 摘自论文原文
  • 用跨镜头记忆+语音提取,保持角色外观和声音一致
  • 世界模型支持多种控制输入,实现视角自由交互
  • 新训练方法让长视频生成更稳定,质量领先

视频生成正从短片段转向长叙事和可交互世界,需维持角色身份、响应用户控制并长期稳定。我们提出JoyAI-Echo-1.5,一个统一的音视频生成系统,包含两个专用变体。长视频版本引入可组合的跨镜头记忆,聚合多帧视觉证据,并从语音过滤的全帧音频中提取说话人线索,实现文本、图像与记忆条件下的角色外观与声音身份持久性。世界模型版本将异构导航输入转化为校准的6-DoF相机轨迹,通过几何感知条件路径实现控制器无关的灵活视角交互。为支持高效长时生成,我们将双向音视频骨干网络改造为因果少步生成器,采用渐进式教师强制与自生成回放中的短/长时程自梯度强制。实验表明,在两项基准上均表现优异:长视频版本在跨镜头一致性、视觉质量、文本对齐与语音保真度上优于现有基线;世界模型版本在WBench上平均得分81.7,位列第一;在SANA-WM-Bench上达到领先的视觉质量与长时程持久性。结果表明,记忆机制、几何控制与回放感知训练共同构成生成连贯故事与持续演化的交互世界的实用基础。

原文摘要 · Abstract (English)

Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.

长视频生成交互世界音视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。