arXiv:2601.01915cs.CV2026-01被引 3

无需训练,对话式操控图像编辑,精准调用现有工具。

TalkPhoto: A Versatile Training-Free Conversational Assistant for Intelligent Image Editing

  • 用提示词引导开源大模型分析指令,分层调用已有编辑方法。
  • 零训练实现多任务编辑,比传统方法少用 token 且效果更好。
  • 适合希望快速集成新编辑功能的开发者或设计师使用。

得益于大语言模型强大的语言理解能力,现有基于指令的图像编辑方法引入多模态大语言模型(MLLMs)以促进指令与图像间的信息交互,提升编辑的可控性与灵活性。然而,这些框架通常需构建多指令数据集来训练模型以处理多种编辑任务,不仅耗时耗力,且效果不佳。本文提出 TalkPhoto,一种无需训练的多功能图像编辑框架,通过对话交互实现精确图像操作。我们通过设计的提示模板指导开源大模型,在接收指令后分析用户需求,并分层调用现有的先进编辑方法,全程无需额外训练。此外,我们实现了即插即用、高效的编辑方法调用机制,使复杂且未见过的编辑任务可无缝集成到当前框架中,获得稳定且高质量的编辑结果。大量实验表明,该方法不仅在调用准确性上更优、令牌消耗更低,且在各类图像编辑任务中均达到更高质量。

原文摘要 · Abstract (English)

Thanks to the powerful language comprehension capabilities of Large Language Models (LLMs), existing instruction-based image editing methods have introduced Multimodal Large Language Models (MLLMs) to promote information exchange between instructions and images, ensuring the controllability and flexibility of image editing. However, these frameworks often build a multi-instruction dataset to train the model to handle multiple editing tasks, which is not only time-consuming and labor-intensive but also fails to achieve satisfactory results. In this paper, we present TalkPhoto, a versatile training-free image editing framework that facilitates precise image manipulation through conversational interaction. We instruct the open-source LLM with a specially designed prompt template to analyze user needs after receiving instructions and hierarchically invoke existing advanced editing methods, all without additional training. Moreover, we implement a plug-and-play and efficient invocation of image editing methods, allowing complex and unseen editing tasks to be integrated into the current framework, achieving stable and high-quality editing results. Extensive experiments demonstrate that our method not only provides more accurate invocation with fewer token consumption but also achieves higher editing quality across various image editing tasks.

图像编辑对话系统零训练大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。