arXiv:2508.16512cs.CV2025-08被引 3

微调视频生成模型会牺牲动态准确性,影响自动驾驶仿真效果。

Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation

  • 发现微调使模型更关注视觉真实感而非动态行为精确性。
  • 在驾驶场景中,模型对物体运动的建模精度下降,但画面更逼真。
  • 用跨领域持续学习可平衡画质与动态准确性,适合仿真系统优化。

近期视频生成技术显著提升了视觉质量与时间连贯性,使其在自动驾驶仿真和“世界模型”等应用中愈发重要。本文研究现有微调方法对结构化驾驶数据集的影响,发现存在潜在权衡:虽然视觉保真度提高,但动态元素的空间建模准确性可能下降。我们归因于视觉质量与动态理解目标之间的对齐偏移。在具有多样化时空结构的数据集中,这两者高度相关;但在驾驶场景这种规律重复的环境中,模型可通过拟合主导运动模式提升画质,却未必保持细粒度动态行为。因此,微调促使模型优先追求表面真实感。进一步实验表明,简单的持续学习策略(如跨域回放)能有效兼顾空间准确性和强视觉质量,提供更优替代方案。

原文摘要 · Abstract (English)

Recent advancements in video generation have substantially improved visual quality and temporal coherence, making these models increasingly appealing for applications such as autonomous driving, particularly in the context of driving simulation and so-called "world models". In this work, we investigate the effects of existing fine-tuning video generation approaches on structured driving datasets and uncover a potential trade-off: although visual fidelity improves, spatial accuracy in modeling dynamic elements may degrade. We attribute this degradation to a shift in the alignment between visual quality and dynamic understanding objectives. In datasets with diverse scene structures within temporal space, where objects or perspective shift in varied ways, these objectives tend to highly correlated. However, the very regular and repetitive nature of driving scenes allows visual quality to improve by modeling dominant scene motion patterns, without necessarily preserving fine-grained dynamic behavior. As a result, fine-tuning encourages the model to prioritize surface-level realism over dynamic accuracy. To further examine this phenomenon, we show that simple continual learning strategies, such as replay from diverse domains, can offer a balanced alternative by preserving spatial accuracy while maintaining strong visual quality.

视频生成自动驾驶微调仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。