针对图像中人-物交互编辑难题,提出新基准与自修正框架,提升动态关系生成准确性。
Taming I2V models for Image HOI Editing: A Cognitive Benchmark and Agentic Self-Correcting Framework

- 构建分层级的HOI-Edit基准,用视觉语言模型问答评估交互真实性。
- 发现I2V模型因具备时序生成能力,可复现失败过程以定位错误原因。
- 设计自修正框架SCPE,通过迭代提示优化生成视频,精准还原目标交互。
当前图像编辑方法擅长处理静态属性,却难以应对复杂的人-物交互(HOI)问题。现有基准将HOI与静态属性混为一谈,依赖全局指标,无法同时评估动态交互有效性与人-物对的保留情况。为此,我们首次提出HOI-Edit,一个包含三个认知层级的综合性基准,其自动化度量工具HOI-Eval通过让视觉语言模型在含语义锚定的人-物对图像上推理后问答,可靠评估实例级交互。鉴于任务本质是重构动态关系,我们评测了图像到视频(I2V)模型,发现其固有的时序生成能力使其天然适合动态编辑。关键在于,除性能优越外,该能力还提供“失败过程回放”,实现独特可诊断性。因此,我们提出SCPE(Self-Correcting Process Editing),一种新型代理式自修正框架,通过迭代优化提示约束I2V生成,使生成视频更准确呈现目标HOI。最终编辑结果由这些视频提取帧获得。在HOI-Edit上,SCPE在交互表现上达到与前沿编辑模型如Nano Banana相当的水平。代码已开源:https://github.com/oceanflowlab/HOI-Edit。
原文摘要 · Abstract (English)
Current image editing methods excel at static attributes but fail at complex Human-Object Interactions (HOI), a critical challenge unaddressed by existing benchmarks that conflate HOI with static attributes, relying on global metrics incapable of simultaneously assessing dynamic interaction validity and entangled human-object pair preservation. Thus, we first introduce HOI-Edit, a comprehensive benchmark with three progressive cognitive levels, which features an automated metric HOI-Eval that reliably evaluates instance-level interaction by letting VLM Q&A after thinking with images containing grounded Human-Object pairs. Considering the task's essence of remodeling dynamic relationships, we benchmark Image-to-Video (I2V) models, finding them inherently suited for dynamic editing due to their temporal generation capabilities. Crucially, beyond superior performance, this capability provides a "replay of the failure process," offering unique diagnosability into why errors occur. We thus propose SCPE (Self-Correcting Process Editing), a novel, agentic self-correcting framework that constrains the generation of I2V models through iteratively refined prompts, enabling the generated videos to more accurately present the target HOI. Extracted frames from these videos are the final editing results. On HOI-Edit, SCPE achieves performance competitive with state-of-the-art (SOTA) editing models like Nano Banana on interaction. Code is available at https://github.com/oceanflowlab/HOI-Edit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。