arXiv:2602.12177cs.CV2026-02被引 1

统一处理多源遥感数据的压缩编码器,提升生成模型效率

EO-VAE: Towards A Multi-sensor Tokenizer for Earth Observation Data

  • 用动态超网络统一建模多种传感器数据
  • 在TerraMesh上重建质量优于现有方法
  • 适合遥感生成模型研究者使用

当前先进的图像与视频生成模型依赖于将高维输入压缩为高效潜在表示的分词器。尽管这一范式已革新了RGB数据生成,但地球观测(EO)数据因传感器规格多样、光谱通道不一而面临独特挑战。本文提出EO-VAE,一种面向地球观测领域的多传感器变分自编码器,作为该领域的基础分词器。不同于以往对每种模态分别训练分词器的方法,EO-VAE通过动态超网络实现单一模型对灵活通道组合的编码与重建。在TerraMesh数据集上的实验表明,相比TerraMind分词器,EO-VAE实现了更优的重建保真度,为遥感领域的潜在生成建模建立了稳健基准。

原文摘要 · Abstract (English)

State-of-the-art generative image and video models rely heavily on tokenizers that compress high-dimensional inputs into more efficient latent representations. While this paradigm has revolutionized RGB generation, Earth observation (EO) data presents unique challenges due to diverse sensor specifications and variable spectral channels. We propose EO-VAE, a multi-sensor variational autoencoder designed to serve as a foundational tokenizer for the EO domain. Unlike prior approaches that train separate tokenizers for each modality, EO-VAE utilizes a single model to encode and reconstruct flexible channel combinations via dynamic hypernetworks. Our experiments on the TerraMesh dataset demonstrate that EO-VAE achieves superior reconstruction fidelity compared to the TerraMind tokenizers, establishing a robust baseline for latent generative modeling in remote sensing.

遥感生成模型变分自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。