通过错误驱动的场景编辑,提升大模型在3D环境中的语言定位能力。
Error-Driven Scene Editing for 3D Grounding in Large Language Models
- 基于错误诊断生成精准场景修改,构建针对性训练数据
- 在多个3D定位任务中实现4-6%的性能提升
- 适合研究3D视觉语言模型与空间理解的学者
尽管3D大语言模型取得进展,但在将语言准确关联到3D环境的视觉与空间元素方面仍受限。这主要源于训练数据侧重语言推理而非空间理解,因3D资源稀缺导致固有的定位偏差未被解决。为此,我们提出将3D场景编辑作为核心机制,通过细粒度空间操作生成视觉反事实样本,以缓解偏差,且无需昂贵的场景重建或大规模3D数据收集。为使编辑更具针对性,我们引入DEER-3D——一种错误驱动框架,通过“分解、诊断、编辑、重训练”的结构化循环,识别定位失败并生成目标反事实监督信号。具体而言,给定一次定位失败,DEER-3D首先识别谓词层级错误(如属性或空间关系),随后执行最小化的谓词对齐编辑(如重新着色或重定位),并构造明确针对失败谓词的问答对,形成靶向反事实训练样本。我们在多个3D定位与场景理解基准上评估该编辑流程,结果表明通过迭代优化,在所有接地数据集上均实现4-6%的性能提升。DEER-3D验证了靶向、错误驱动的场景编辑在弥合语言推理与空间定位之间鸿沟方面的有效性。
原文摘要 · Abstract (English)
Despite recent progress in 3D-LLMs, they remain limited in accurately grounding language to visual and spatial elements in 3D environments. This limitation stems in part from training data that focuses on language reasoning rather than spatial understanding due to scarce 3D resources, leaving inherent grounding biases unresolved. To address this, we propose 3D scene editing as a key mechanism to generate visual counterfactuals that mitigate these biases through fine-grained spatial manipulation, without requiring costly scene reconstruction or large-scale 3D data collection. Furthermore, to make these edits targeted and directly address the specific weaknesses of the model, we introduce DEER-3D, an error-driven framework that diagnoses grounding failures and generates targeted counterfactual training supervision via a structured "Decompose, Diagnose, Edit, and Retrain" loop. Specifically, given a grounding failure, DEER-3D first identifies the predicate-level error (e.g., attribute or spatial relation). It then performs minimal predicate-aligned scene edits, such as recoloring or repositioning, and constructs aligned question-answer pairs that explicitly target the failed predicate, forming targeted counterfactual training examples. We evaluate our editing pipeline across multiple benchmarks for 3D grounding and scene understanding tasks, consistently demonstrating improvements across all grounding datasets through iterative refinement (4-6% gains). DEER-3D underscores the effectiveness of targeted, error-driven scene editing in bridging linguistic reasoning with spatial grounding in 3D LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。