arXiv:2512.05131cs.CVcs.AI2025-12被引 2

用视觉语言引导的智能体,高效重建3D场景

AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance

  • 采用前馈3D重建与视觉语言模型结合的新框架
  • 在稀疏视角下实现领先精度,减少冗余观测
  • 适合需要高效3D建模的机器人与自动驾驶应用

主动3D重建使智能体能自主选择观测视角,以高效获取精确完整的场景几何信息,而非依赖预先采集的图像被动重建。然而,现有方法多依赖手工设计的几何启发式规则,常导致冗余观测且对重建质量提升有限。为此,我们提出AREA3D,一种融合前馈3D重建模型与视觉语言引导的主动重建智能体。该框架将视角不确定性建模与底层前馈重建器解耦,实现无需在线优化的精确不确定性估计。同时,集成的视觉语言模型提供高层语义指导,促使智能体选择更具信息量且多样化的视角,超越纯几何线索。在场景级与物体级基准上的大量实验表明,AREA3D在稀疏视角条件下达到当前最优重建精度。代码将公开于:https://github.com/TianlingXu/AREA3D。

原文摘要 · Abstract (English)

Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rather than passively reconstructing scenes from pre-collected images. However, existing active reconstruction methods often rely on hand-crafted geometric heuristics, which can lead to redundant observations without substantially improving reconstruction quality. To address this limitation, we propose AREA3D, an active reconstruction agent that leverages feed-forward 3D reconstruction models and vision-language guidance. Our framework decouples view-uncertainty modeling from the underlying feed-forward reconstructor, enabling precise uncertainty estimation without expensive online optimization. In addition, an integrated vision-language model provides high-level semantic guidance, encouraging informative and diverse viewpoints beyond purely geometric cues. Extensive experiments on both scene-level and object-level benchmarks demonstrate that AREA3D achieves state-of-the-art reconstruction accuracy, particularly in the sparse-view regime. Code will be made available at: https://github.com/TianlingXu/AREA3D .

3D重建视觉语言智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。