arXiv:2604.09057cs.CVcs.MM2026-04中稿 · ACM MM 2026

用物体轨迹统一指导音视频生成,让动作和声音更真实协同。

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence

论文配图:Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence
图 1 · 摘自论文原文
  • 以物体轨迹为共享运动先验,联合引导视频与音频生成。
  • 在公开数据集上实现比基线高18.3%的运动-声音同步率。
  • 适合做影视合成、虚拟角色动画的开发者参考。

音视频(AV)生成在感知质量和多模态一致性方面取得显著进展,但生成符合物理规律的动作-声音关系仍具挑战。现有方法常导致物体运动视觉不稳,声音与显著运动或接触事件关联松散,主要因缺乏视频与音频共享的显式运动结构。本文提出Tora3,一种基于轨迹引导的音视频生成框架,通过使用物体轨迹作为共享运动先验,提升物理一致性。不同于仅将轨迹用于视频控制,Tora3将其同时用于联合引导视觉运动与声学事件。具体包括:设计轨迹对齐的运动表示,构建由轨迹推导出的二阶运动状态驱动的运动-音频对齐模块,以及结合流匹配的混合生成策略,在轨迹条件区域保持轨迹保真度,同时在其他区域维持局部一致性。我们还构建了大型音视频数据集PAV,强调运动相关模式,并自动提取运动标注。大量实验表明,Tora3在运动真实感、动作-声音同步性及整体生成质量上均优于多个强基线模型。

原文摘要 · Abstract (English)

Audio-video (AV) generation has recently made strong progress in perceptual quality and multimodal coherence, yet generating content with plausible motion-sound relations remains challenging. Existing methods often produce object motions that are visually unstable and sounds that are only loosely aligned with salient motion or contact events, largely because they lack an explicit motion-aware structure shared by video and audio generation. We present Tora3, a trajectory-guided AV generation framework that improves physical coherence by using object trajectories as a shared kinematic prior. Rather than treating trajectories as a video-only control signal, Tora3 uses them to jointly guide visual motion and acoustic events. Specifically, we design a trajectory-aligned motion representation for video, a kinematic-audio alignment module driven by trajectory-derived second-order kinematic states, and a hybrid flow matching scheme that preserves trajectory fidelity in trajectory-conditioned regions while maintaining local coherence elsewhere. We further curate PAV, a large-scale AV dataset emphasizing motion-relevant patterns with automatically extracted motion annotations. Extensive experiments show that Tora3 improves motion realism, motion-sound synchronization, and overall AV generation quality over strong open-source baselines. Project page: https://ali-videoai.github.io/tora3_page.

音视频生成轨迹引导物理一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。