测试视觉语言模型对真实动作的可逆理解能力,揭示当前模型在动作推理上的不足。
Do-Undo Bench: Reversibility for Action Understanding in Image Generation
- 要求模型先模拟真实动作效果,再还原到原始状态,检验因果理解能力。
- 现有模型在可逆动作任务上表现不佳,平均准确率低于40%。
- 适合研究多模态生成、动作理解与真实世界动态建模的学者使用。
我们提出Do-Undo任务与基准,以填补视觉语言模型在理解并生成由真实动作驱动的合理场景变换方面的关键空白。与依赖提示驱动图像生成和编辑的先前工作不同,我们的训练假设要求模型模拟真实动作的结果,再将其逆向恢复至原始状态。这种正向-反向的双重要求测试的是真正的因果理解,而非风格或语义编辑。我们从真实场景中收集高质量的可逆动作数据集,以实现稳健的动作定位。实验表明,当前模型在动作可逆性任务上表现较差,凸显了评估动作理解的必要性。Do-Undo为评估和推进需推理真实世界动态的多模态系统中的动作感知生成提供了直观的测试平台。
原文摘要 · Abstract (English)
We introduce the Do-Undo task and benchmark to address a critical gap in vision-language models: understanding and generating plausible scene transformations driven by real-world actions. Unlike prior work that relies on prompt-based image generation and editing to perform action-conditioned image manipulation, our training hypothesis requires models to simulate the outcome of a real-world action and then reverse it to the original state. This forward-reverse requirement tests genuine cause-and-effect understanding rather than stylistic or semantic edits. We curate a high-quality benchmark of reversible actions from real-world scenarios to enable robust action grounding. Our experiments reveal that current models struggle with action reversibility, highlighting the need to evaluate action understanding. Do-Undo provides an intuitive testbed for evaluating and advancing action-aware generation in multimodal systems that must reason about real-world dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。