arXiv:2602.24096cs.CVcs.AI2026-02被引 7

用扩散模型在线提升机器人仿真画面真实感与一致性

DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhancer

  • 将预训练扩散模型转为单步时序增强器,支持实时运行
  • 在NeRF/3DGS渲染基础上显著减少视角畸变与光影不自然
  • 适合自动驾驶仿真研究与工业级系统部署

仿真对自动驾驶等自主机器人开发至关重要。神经重建方法如NeRF和3D Gaussian Splatting可从真实数据自动生成多样化场景,但渲染新视角时常出现伪影,且插入动态物体时难以实现真实光影融合,尤其当物体来自不同场景时。为此,我们提出DiffusionHarmonizer,一种在线生成增强框架,能将不完美渲染结果转化为时序一致、更逼真的输出。核心是一个由预训练多步扩散模型转换而来的单步时序条件增强器,可在单块GPU上实时运行于在线仿真器中。其有效训练依赖于定制的数据构建流水线,专门生成强调外观协调、伪影修正与光照真实性的合成-真实图像对。该系统在科研与生产环境中均显著提升了仿真保真度。

原文摘要 · Abstract (English)

Simulation is essential to the development and evaluation of autonomous robots such as self-driving vehicles. Neural reconstruction is emerging as a promising solution as it enables simulating a wide variety of scenarios from real-world data alone in an automated and scalable way. However, while methods such as NeRF and 3D Gaussian Splatting can produce visually compelling results, they often exhibit artifacts particularly when rendering novel views, and fail to realistically integrate inserted dynamic objects, especially when they were captured from different scenes. To overcome these limitations, we introduce DiffusionHarmonizer, an online generative enhancement framework that transforms renderings from such imperfect scenes into temporally consistent outputs while improving their realism. At its core is a single-step temporally-conditioned enhancer that is converted from a pretrained multi-step image diffusion model, capable of running in online simulators on a single GPU. The key to training it effectively is a custom data curation pipeline that constructs synthetic-real pairs emphasizing appearance harmonization, artifact correction, and lighting realism. The result is a scalable system that significantly elevates simulation fidelity in both research and production environments.

机器人仿真扩散模型神经渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。