arXiv:2507.11129cs.CVcs.AI2025-07ICCV被引 2

统一建模多模态场景,解决模态间差异带来的表示难题

MMOne: Representing Multiple Modalities in One Scene

  • 用新型模态指示器捕捉各模态独特属性
  • 通过多模态分解将混合高斯拆分为单模态成分
  • 可扩展至新模态,适合多模态感知研究者

人类通过多模态线索理解与互动环境。学习多模态场景表示有助于深入理解物理世界。然而,模态间固有差异导致属性不一致和粒度不一致两大挑战。为此,我们提出通用框架MMOne,实现单场景内多模态统一表示,并可轻松扩展至新增模态。具体而言,设计含新颖模态指示器的模态建模模块以捕获各模态特性;提出多模态分解机制,基于模态差异将多模态高斯分解为单模态高斯。通过解耦共享与模态特定成分,实现更紧凑高效的多模态场景表示。大量实验表明,该方法持续提升各模态表示能力,且具备良好的可扩展性。代码已开源:https://github.com/Neal2020GitHub/MMOne。

原文摘要 · Abstract (English)

Humans perceive the world through multimodal cues to understand and interact with the environment. Learning a scene representation for multiple modalities enhances comprehension of the physical world. However, modality conflicts, arising from inherent distinctions among different modalities, present two critical challenges: property disparity and granularity disparity. To address these challenges, we propose a general framework, MMOne, to represent multiple modalities in one scene, which can be readily extended to additional modalities. Specifically, a modality modeling module with a novel modality indicator is proposed to capture the unique properties of each modality. Additionally, we design a multimodal decomposition mechanism to separate multi-modal Gaussians into single-modal Gaussians based on modality differences. We address the essential distinctions among modalities by disentangling multimodal information into shared and modality-specific components, resulting in a more compact and efficient multimodal scene representation. Extensive experiments demonstrate that our method consistently enhances the representation capability for each modality and is scalable to additional modalities. The code is available at https://github.com/Neal2020GitHub/MMOne.

多模态场景表示高斯溅射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。