用可视图结构让复杂图像编辑更直观,无需猜提示词。
SceneCraft: Interactive System for Image Editing via Scene Graph

- 将图像转为可交互的场景图,直接操作物体关系
- 用户操作后自动生成精准编辑提示,减少试错
- 支持多模型并行生成,结果更稳定高质量
生成式AI的进步使自然语言驱动的图像编辑成为可能,但现有系统在包含多个交互物体的复杂场景中表现不佳,因其严重依赖用户构建精确文本提示。为解决结构化控制缺失的问题,我们提出SceneCraft,一种新颖的交互式框架,通过将图像表示为可编辑的场景图,连接用户意图与模型执行。用户不再需反复尝试猜测文本提示,而是直接在可视化图上进行空间和关系操作。这些图结构修改被自动转化为精确、上下文感知的编辑提示,有效消除语言歧义。为确保结果鲁棒且多样,结构化提示被分发至多个最先进的生成模型。在多种编辑场景下的评估显示,SceneCraft提供了更直观的控制方式,显著降低手动提示工程的认知负担,生成结果在质量与保真度方面均获用户一致更高评价。
原文摘要 · Abstract (English)
Recent advances in generative AI have enabled natural language-driven image editing, yet existing systems often fail in complex scenes with multiple interacting objects because they rely heavily on users crafting precise text prompts. To address the absence of structured control, we propose SceneCraft, a novel interactive framework that bridges user intent and model execution by representing images as editable scene graphs. Instead of guessing text prompts through trial and error, users interact directly with a visual graph to perform complex spatial and relational operations. These graph modifications are automatically translated into precise, context-aware editing prompts, effectively eliminating linguistic ambiguity. To ensure robust and diverse results, structured prompts are dispatched to multiple state-of-the-art generative models. Evaluations across diverse editing scenarios show that SceneCraft provides a more intuitive control mechanism, significantly reducing the cognitive burden of manual prompt engineering while generating outputs that users consistently rate as higher in quality and fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。