arXiv:2502.18116cs.CVcs.AI2025-02ACL被引 16

用大模型+贝叶斯优化实现无需标注的精准图像编辑

Bayesian Optimization for Controlled Image Editing via LLMs

  • 通过贝叶斯优化自动调参,让大模型理解自然语言指令
  • 在不微调模型前提下,编辑准确率显著优于现有方法
  • 适合希望无代码控制图像生成的设计师与研究者

在快速发展的图像生成领域,实现对生成内容的精确控制并保持语义一致性仍是关键挑战,尤其体现在定位技术与模型微调需求方面。为此,我们提出 BayesGenie,一种无需预训练或微调的即插即用方法,将大语言模型(LLMs)与贝叶斯优化结合,实现用户友好的精准图像编辑。该方法支持通过自然语言描述修改图像,无需手动标记区域,同时保持原图语义完整。得益于其模型无关设计,BayesGenie 在多种 LLM(如 Claude3 与 GPT-4)上均表现出强适应性。我们采用改进的贝叶斯优化策略自动调整推理参数,实现低干预下的高精度编辑。大量实验表明,该框架在编辑准确性和语义保留方面显著优于现有方法。

原文摘要 · Abstract (English)

In the rapidly evolving field of image generation, achieving precise control over generated content and maintaining semantic consistency remain significant limitations, particularly concerning grounding techniques and the necessity for model fine-tuning. To address these challenges, we propose BayesGenie, an off-the-shelf approach that integrates Large Language Models (LLMs) with Bayesian Optimization to facilitate precise and user-friendly image editing. Our method enables users to modify images through natural language descriptions without manual area marking, while preserving the original image's semantic integrity. Unlike existing techniques that require extensive pre-training or fine-tuning, our approach demonstrates remarkable adaptability across various LLMs through its model-agnostic design. BayesGenie employs an adapted Bayesian optimization strategy to automatically refine the inference process parameters, achieving high-precision image editing with minimal user intervention. Through extensive experiments across diverse scenarios, we demonstrate that our framework significantly outperforms existing methods in both editing accuracy and semantic preservation, as validated using different LLMs including Claude3 and GPT-4.

图像编辑大模型贝叶斯优化自然语言控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。