用文字描述任意物体,就能补全被遮挡部分的完整外观。
Open-World Amodal Appearance Completion
- 通过文本查询实现无需训练的开放世界补全
- 支持直接词汇和抽象描述,泛化能力更强
- 生成可无缝融入3D重建与图像编辑的透明图层
理解并重建被遮挡物体是一项挑战性任务,尤其在类别与场景多样且不可预测的开放世界中。传统方法通常局限于封闭的物体类别集合,难以应对复杂场景。我们提出开放世界无模态外观补全(Open-World Amodal Appearance Completion),一种无需训练的框架,可通过灵活的文本查询输入扩展补全能力。该方法能针对直接词汇和抽象查询,通用地重构任意物体的完整外观,我们称之为推理式无模态补全。系统结合分割、遮挡分析与修复,生成包含透明度信息的RGBA元素,支持3D重建与图像编辑等应用。大量实验表明,该方法在新物体与新遮挡模式上具有优异泛化能力,为开放世界无模态补全设立了新基准。代码与数据集将在论文录用后发布。
原文摘要 · Abstract (English)
Understanding and reconstructing occluded objects is a challenging problem, especially in open-world scenarios where categories and contexts are diverse and unpredictable. Traditional methods, however, are typically restricted to closed sets of object categories, limiting their use in complex, open-world scenes. We introduce Open-World Amodal Appearance Completion, a training-free framework that expands amodal completion capabilities by accepting flexible text queries as input. Our approach generalizes to arbitrary objects specified by both direct terms and abstract queries. We term this capability reasoning amodal completion, where the system reconstructs the full appearance of the queried object based on the provided image and language query. Our framework unifies segmentation, occlusion analysis, and inpainting to handle complex occlusions and generates completed objects as RGBA elements, enabling seamless integration into applications such as 3D reconstruction and image editing. Extensive evaluations demonstrate the effectiveness of our approach in generalizing to novel objects and occlusions, establishing a new benchmark for amodal completion in open-world settings. The code and datasets will be released after paper acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。