无需训练,用可编辑的记忆体让大模型零样本推理3D场景
Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models

- 将3D场景建模为可编辑的图文记忆体,通过组合空间工具与现成多模态模型交互
- 在Compose3D上实现超越微调方法的多跳推理能力,尤其依赖动态生成空间操作
- 适合需要灵活推理新物体、空地布局的开放场景应用,如智能设计助手
3D场景理解涵盖自由空间推理、物体定位、假设物体插入、复杂几何关系分析,并需整合外部工具与数据。现有方法通常依赖大规模3D-语言训练或仅聚焦物体定位与简单空间关系。我们主张,广义泛化能力可在推理时实现,无需特定3D训练。提出Flame3D——一种无训练框架,将场景表示为可编辑的视觉-文本3D记忆体,通过可组合的空间工具向现成多模态大模型(MLLM)暴露。该框架允许代理在推理时合成自定义空间程序,支持对布局、空域及未出现物体的开放推理。外部数据与修正可注入记忆体而无需重新训练。在ScanQA上表现媲美微调3D-LMM,在定制的组合式空间推理基准Compose3D上评估多跳推理能力。结果表明固定工具不足,代理在推理时合成空间操作的能力至关重要。这引发思考:未来3D理解应聚焦更丰富的场景记忆与表达性组合抽象?
原文摘要 · Abstract (English)
3D scene understanding spans reasoning about free space, object grounding, hypothetical object insertions, complex geometric relationships, and integrating all of these with external tools and data sources. Existing 3D understanding methods typically rely on large-scale 3D-language training or focus on object grounding and simple spatial relationships. We argue that the broad generalization that motivates 3D-language training can be achieved at inference time, without 3D-specific training. We propose Flame3D, a training-free framework that represents scenes as editable visual-textual 3D memories and exposes them to an off-the-shelf MLLM through composable spatial tools. Flame3D also lets the agent synthesize custom spatial programs at inference time, enabling open-ended reasoning over layouts, empty space, and objects not yet present in the scene. External data and corrections can be added to the memory without retraining. In addition to showing competitive performance to finetuned 3D-LMM methods on ScanQA, we study multi-hop 3D reasoning capabilities of Flame3D by evaluating it on a curated compositional spatial-reasoning benchmark, Compose3D. We find that fixed tools fall short and that the agent's ability to synthesize spatial operations at inference time is essential. These results invite the question: should future progress in 3D scene understanding focus on richer scene memories and expressive compositional abstractions?
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。