arXiv:2505.05644cs.CVeess.IV2025-05被引 3

用统一Transformer融合多种月球数据,实现高精度3D重建与反照率估计。

The Moon's Many Faces: A Single Unified Transformer for Multimodal Lunar Reconstruction

  • 设计单一Transformer模型,跨模态学习灰度图、高程图等四类月球数据。
  • 在月面3D重建和光照反照率估计任务中表现优异,物理合理性强。
  • 适用于行星科学中的大规模三维建模,未来可扩展更多模态。

多模态学习是多个学科的新兴研究方向,但在行星科学中应用极少。本文提出一种单一、统一的Transformer架构,通过训练学习灰度图像、数字高程模型(DEMs)、表面法线和反照率图之间的共享表示。该架构支持任意输入模态到任意目标模态的灵活转换。实验表明,该基础模型能学习到这四种模态间物理上合理的关联。我们进一步发现,基于图像的月面3D重建与反照率估计(形状与反照率从阴影)可被建模为多模态学习问题。结果证明多模态学习在解决此问题上的潜力,并为大规模行星三维重建提供新方法。未来引入更多输入模态将进一步提升性能,实现光度归一化和配准等任务。

原文摘要 · Abstract (English)

Multimodal learning is an emerging research topic across multiple disciplines but has rarely been applied to planetary science. In this contribution, we propose a single, unified transformer architecture trained to learn shared representations between multiple sources like grayscale images, Digital Elevation Models (DEMs), surface normals, and albedo maps. The architecture supports flexible translation from any input modality to any target modality. Our results demonstrate that our foundation model learns physically plausible relations across these four modalities. We further identify that image-based 3D reconstruction and albedo estimation (Shape and Albedo from Shading) of lunar images can be formulated as a multimodal learning problem. Our results demonstrate the potential of multimodal learning to solve Shape and Albedo from Shading and provide a new approach for large-scale planetary 3D reconstruction. Adding more input modalities in the future will further improve the results and enable tasks such as photometric normalization and co-registration.

月球重建多模态Transformer3D建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。