arXiv:2503.13500cs.LGcs.AI2025-03ICLR被引 6

LIGER让长序列视觉指令自动纠错,图像更连贯准确。

Long-horizon Visual Instruction Generation with Logic and Attribute Self-reflection

  • 每步生成时结合历史提示与视觉记忆,保持前后一致
  • 用编辑工具修正属性、逻辑、物体冗余等错误,提升准确性
  • 适合需要精准视觉引导的复杂任务,如教学或流程演示

长时序任务的视觉指令对直观理解复杂概念和长期记忆至关重要。直接使用文生图模型生成一系列图像而忽略前序步骤上下文,会导致图像不一致,增加认知负担。且生成图像常缺失物体或属性(如颜色、形状、状态)不准确。为此,我们提出LIGER——首个无需训练的长时序指令生成框架,具备逻辑与属性自反思能力。LIGER首先利用历史提示与前序步骤的视觉记忆生成每一步的草稿图像,实现长时序任务中图像的一致性。随后,通过多种图像编辑工具修正草稿图像中的属性错误、逻辑矛盾、物体冗余及身份不一致等问题。该自反思机制显著提升了图像的逻辑合理性与对象属性正确性。为验证生成图像是否有助于人类理解,我们人工构建了一个包含多种长时序任务的新基准,其人类标注的真值表达式反映了人类对图像应具表现力的标准。实验表明,相较于基线方法,LIGER生成的视觉指令更具完整性。

原文摘要 · Abstract (English)

Visual instructions for long-horizon tasks are crucial as they intuitively clarify complex concepts and enhance retention across extended steps. Directly generating a series of images using text-to-image models without considering the context of previous steps results in inconsistent images, increasing cognitive load. Additionally, the generated images often miss objects or the attributes such as color, shape, and state of the objects are inaccurate. To address these challenges, we propose LIGER, the first training-free framework for Long-horizon Instruction GEneration with logic and attribute self-Reflection. LIGER first generates a draft image for each step with the historical prompt and visual memory of previous steps. This step-by-step generation approach maintains consistency between images in long-horizon tasks. Moreover, LIGER utilizes various image editing tools to rectify errors including wrong attributes, logic errors, object redundancy, and identity inconsistency in the draft images. Through this self-reflection mechanism, LIGER improves the logic and object attribute correctness of the images. To verify whether the generated images assist human understanding, we manually curated a new benchmark consisting of various long-horizon tasks. Human-annotated ground truth expressions reflect the human-defined criteria for how an image should appear to be illustrative. Experiments demonstrate the visual instructions generated by LIGER are more comprehensive compared with baseline methods.

视觉指令长时序生成自反思图像一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。