用全景图提升大模型对物理世界的密集理解能力
Dense360: Dense Understanding from Omnidirectional Panoramas
- 设计了适用于全景图的新型位置编码ERP-RoPE,解决投影畸变问题
- 构建了包含500万实体级描述的16万张全景数据集,规模领先
- 提出首个全景图文理解评测基准Dense360-Bench,推动领域发展
多模态大语言模型需全面视觉输入以实现对物理世界的密集理解。现有模型依赖有限视场(如70度)输入,表现受限。本文首次探索从全向全景图中实现密集理解,构建了首个带有可靠性评分的全景图数据集,包含16万张全景图、500万实体级描述、100万唯一指代表达和10万条实体引导的全景场景描述。相比多视角方案,等距圆柱投影(ERP)能提供更完整、紧凑且连续的场景表示。但ERP带来两个关键挑战:纬线方向的空间不连续性,以及纬度相关的信息密度变化。为此,我们提出ERP-RoPE位置编码方案,专门应对全景投影特性。同时,提出Dense360-Bench——首个用于评估全景图像描述与定位的基准,建立全面评测框架,推动全景视觉-语言理解发展。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) require comprehensive visual inputs to achieve dense understanding of the physical world. While existing MLLMs demonstrate impressive world understanding capabilities through limited field-of-view (FOV) visual inputs (e.g., 70 degree), we take the first step toward dense understanding from omnidirectional panoramas. We first introduce an omnidirectional panoramas dataset featuring a comprehensive suite of reliability-scored annotations. Specifically, our dataset contains 160K panoramas with 5M dense entity-level captions, 1M unique referring expressions, and 100K entity-grounded panoramic scene descriptions. Compared to multi-view alternatives, panoramas can provide more complete, compact, and continuous scene representations through equirectangular projections (ERP). However, the use of ERP introduces two key challenges for MLLMs: i) spatial continuity along the circle of latitude, and ii) latitude-dependent variation in information density. We address these challenges through ERP-RoPE, a position encoding scheme specifically designed for panoramic ERP. In addition, we introduce Dense360-Bench, the first benchmark for evaluating MLLMs on omnidirectional captioning and grounding, establishing a comprehensive framework for advancing dense visual-language understanding in panoramic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。