arXiv:2503.11806cs.CV2025-03ICCV被引 5

用户一键修正3D场景布局局部错误,提升复杂场景建模精度。

Human-in-the-Loop Local Corrections of 3D Scene Layouts via Infilling

  • 将局部修正任务建模为自然语言处理中的'补全'问题。
  • 在保持全局预测性能基础上,显著增强局部修正能力。
  • 支持用户迭代修正,适应训练数据外的复杂布局。

我们提出一种新型人机协同方法,从第一视角获取人类反馈以估计3D场景布局。通过引入新颖的局部修正任务,用户可识别局部错误并触发模型自动修正。基于先进的SceneScript框架(利用结构化语言进行3D场景布局估计),我们将该问题重构为自然语言处理中的'补全'任务。训练了一个多任务版本的SceneScript,既维持了全局预测性能,又显著提升了局部修正能力。将其集成至人机协同系统中,用户可通过低门槛的“一键修复”工作流迭代优化场景布局。该系统允许最终结果偏离训练分布,从而更准确地建模复杂布局。

原文摘要 · Abstract (English)

We present a novel human-in-the-loop approach to estimate 3D scene layout that uses human feedback from an egocentric standpoint. We study this approach through introduction of a novel local correction task, where users identify local errors and prompt a model to automatically correct them. Building on SceneScript, a state-of-the-art framework for 3D scene layout estimation that leverages structured language, we propose a solution that structures this problem as "infilling", a task studied in natural language processing. We train a multi-task version of SceneScript that maintains performance on global predictions while significantly improving its local correction ability. We integrate this into a human-in-the-loop system, enabling a user to iteratively refine scene layout estimates via a low-friction "one-click fix'' workflow. Our system enables the final refined layout to diverge from the training distribution, allowing for more accurate modelling of complex layouts.

3D布局人机协同补全任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。