让普通用户用自然语言精准编辑图像视频,无需技术背景。
Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era
- 用自然语言指令驱动视觉编辑,打通用户意图与操作
- 整合100+文献,覆盖从GAN到扩散模型的多类方法
- 适合内容创作者、教育者及跨领域非专业用户
大语言模型(LLMs)和多模态学习的快速发展重塑了数字内容创作与编辑方式。传统视觉工具依赖专业知识,使用门槛高。近年来,基于指令的编辑技术使用户能通过自然语言实现直观交互,将人类意图精准映射至复杂编辑操作。本综述梳理了超过100篇相关研究,涵盖从生成对抗网络到扩散模型的方法,重点探讨多模态融合在细粒度内容控制中的应用。应用场景包括时尚设计、3D场景编辑与视频生成,显著提升可访问性并契合人类直觉。文章对比现有工作,强调LLM赋能的编辑优势,指出关键挑战以推动后续研究。目标是让强大视觉编辑能力普及至娱乐、教育等多个行业。感兴趣读者可访问仓库:https://github.com/tamlhp/awesome-instruction-editing。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) and multimodal learning has transformed digital content creation and manipulation. Traditional visual editing tools require significant expertise, limiting accessibility. Recent strides in instruction-based editing have enabled intuitive interaction with visual content, using natural language as a bridge between user intent and complex editing operations. This survey provides an overview of these techniques, focusing on how LLMs and multimodal models empower users to achieve precise visual modifications without deep technical knowledge. By synthesizing over 100 publications, we explore methods from generative adversarial networks to diffusion models, examining multimodal integration for fine-grained content control. We discuss practical applications across domains such as fashion, 3D scene manipulation, and video synthesis, highlighting increased accessibility and alignment with human intuition. Our survey compares existing literature, emphasizing LLM-empowered editing, and identifies key challenges to stimulate further research. We aim to democratize powerful visual editing across various industries, from entertainment to education. Interested readers are encouraged to access our repository at https://github.com/tamlhp/awesome-instruction-editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。