用手势指导物体添加,让设计更直观精准。
AbracADDbra: Touch-Guided Object Addition by Decoupling Placement and Editing Subtasks
- 分两步走:先用视觉语言模型根据手势定位,再用扩散模型生成物体和掩码。
- 在新基准上表现优于随机放置和通用模型,生成效果真实自然。
- 适合想快速创作、追求操作直观的设计师或创意用户。
基于指令的物体添加常受限于纯文本提示的模糊性或掩码输入的繁琐性。为此,我们提出 AbracADDbra,一种利用直观手势先验实现简洁指令空间定位的用户友好框架。其高效解耦架构首先通过视觉语言变换器进行触控引导定位,再由扩散模型联合生成物体与实例掩码,实现高保真融合。为推动标准化评估,我们构建了 Touch2Add 基准。大量实验表明,我们的定位模型显著优于随机放置和通用视觉语言模型基线,验证了框架生成高质量编辑的能力。进一步分析显示初始定位精度与最终编辑质量强相关,支持解耦策略的有效性。本工作为更易用、高效的创意工具开辟了新路径。
原文摘要 · Abstract (English)
Instruction-based object addition is often hindered by the ambiguity of text-only prompts or the tedious nature of mask-based inputs. To address this usability gap, we introduce AbracADDbra, a user-friendly framework that leverages intuitive touch priors to spatially ground succinct instructions for precise placement. Our efficient, decoupled architecture uses a vision-language transformer for touch-guided placement, followed by a diffusion model that jointly generates the object and an instance mask for high-fidelity blending. To facilitate standardized evaluation, we contribute the Touch2Add benchmark for this interactive task. Our extensive evaluations, where our placement model significantly outperforms both random placement and general-purpose VLM baselines, confirm the framework's ability to produce high-fidelity edits. Furthermore, our analysis reveals a strong correlation between initial placement accuracy and final edit quality, validating our decoupled approach. This work thus paves the way for more accessible and efficient creative tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。