arXiv:2412.10566cs.CV2024-12

让AI理解模糊指令,精准生成跨维度视觉编辑方案

EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing

  • 通过反思式推理框架解析用户意图,结合参考图像生成精准编辑指令
  • 在3万条带理由的示例上训练,实现与人类意图高度对齐
  • 适用于图像、视频、3D、4D等多维编辑,适合内容创作与交互设计

从模糊或部分指定的指令中编辑复杂视觉内容,仍是视觉语言建模的核心挑战。现有模型虽能上下文理解内容,但常无法推断参考图像或场景中的潜在意图,导致编辑结果不一致或错位。我们提出编辑视觉语言模型(EVLM),通过结合参考视觉信息,将模糊指令转化为精确、上下文感知的编辑提示。其核心创新在于一种反思式推理框架,利用反射感知的KL散度目标优化(RKTO)对齐人类标注的推理理由,将主观意图转化为结构化可执行输出。通过链式思维(CoT)与RKTO联合优化,无需二元监督即可捕捉精细编辑偏好。在包含3万条带人类标注理由质量的示例数据集上训练,EVLM在图像、视频、3D和4D编辑任务中均显著提升与人类意图的一致性,生成连贯高质量的编辑指令,为多模态编辑与推理提供可扩展基础。

原文摘要 · Abstract (English)

Editing complex visual content from ambiguous or partially specified instructions remains a core challenge in vision-language modeling. Existing models can contextualize content but often fail to infer the underlying intent within a reference image or scene, leading to inconsistent or misaligned edits. We introduce the Editing Vision-Language Model (EVLM), a system that interprets ambiguous instructions in conjunction with reference visuals to produce precise, context-aware editing prompts. EVLM's key innovation is a reflective reasoning framework that translates subjective user intent into structured, actionable outputs by aligning with human-rated rationales through Reflection-Aware KL-Divergence Target Optimization (RKTO). By combining Chain-of-Thought (CoT) reasoning with RKTO alignment, EVLM captures fine-grained editing preferences without relying on binary supervision. Trained on a dataset of 30,000 CoT examples with human-annotated rationale quality, EVLM achieves substantial gains in alignment with human intent. Experiments across image, video, 3D, and 4D editing tasks show that EVLM generates coherent and high-quality instructions, providing a scalable foundation for multimodal editing and reasoning.

视觉编辑多模态意图理解推理框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。