arXiv:2501.09838cs.CVcs.AI2025-01中稿 · the 2025 WACV work…被引 3

跨模态扩散模型实现多视角图像生成,无需几何先验

CrossModalityDiffusion: Multi-Modal Novel View Synthesis with Unified Intermediate Representation

  • 用统一特征体表示融合多模态输入,建模场景结构
  • 在合成数据集上实现多模态、多视角图像准确生成
  • 适合遥感图像生成与跨模态场景理解研究者

地理空间成像利用多种传感模态(如光学、雷达、激光雷达)的数据,覆盖从地面无人机到卫星视角的多源信息。这些异构输入为场景理解提供了机遇,但缺乏精确真实数据时,准确解析几何结构仍具挑战。为此,我们提出CrossModalityDiffusion,一种模块化框架,可在无场景几何先验的情况下,生成不同模态和视角的图像。该框架采用模态特定编码器,将多输入图像转换为与相机位置相关的几何感知特征体,特征体所在空间作为统一的跨模态表示基础。通过体素渲染技术,将特征体重叠并从新视角投影为特征图像,再作为条件输入给模态特定的扩散模型,以合成目标模态的新图像。本文证明,联合训练各模块可确保框架内所有模态的一致几何理解。我们在合成的ShapeNet车辆数据集上验证了其有效性,展示了跨多模态和视角生成准确一致图像的能力。

原文摘要 · Abstract (English)

Geospatial imaging leverages data from diverse sensing modalities-such as EO, SAR, and LiDAR, ranging from ground-level drones to satellite views. These heterogeneous inputs offer significant opportunities for scene understanding but present challenges in interpreting geometry accurately, particularly in the absence of precise ground truth data. To address this, we propose CrossModalityDiffusion, a modular framework designed to generate images across different modalities and viewpoints without prior knowledge of scene geometry. CrossModalityDiffusion employs modality-specific encoders that take multiple input images and produce geometry-aware feature volumes that encode scene structure relative to their input camera positions. The space where the feature volumes are placed acts as a common ground for unifying input modalities. These feature volumes are overlapped and rendered into feature images from novel perspectives using volumetric rendering techniques. The rendered feature images are used as conditioning inputs for a modality-specific diffusion model, enabling the synthesis of novel images for the desired output modality. In this paper, we show that jointly training different modules ensures consistent geometric understanding across all modalities within the framework. We validate CrossModalityDiffusion's capabilities on the synthetic ShapeNet cars dataset, demonstrating its effectiveness in generating accurate and consistent novel views across multiple imaging modalities and perspectives.

跨模态生成扩散模型三维重建遥感图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。