arXiv:2602.06663cs.CV2026-02被引 4

评测大模型在电脑任务规划中的图像生成能力

PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks

  • 构建新基准PlanViz,评估大模型完成电脑任务图像生成与编辑的能力
  • 针对路线规划、工作图绘制、网页界面显示三类日常任务进行测试
  • 提出计划评分机制,量化生成结果的准确性与效率

统一多模态模型(UMMs)在生成自然图像和多模态推理方面表现出色,但在支持与日常生活密切相关的电脑任务规划方面仍缺乏探索。电脑使用任务中的图像生成与编辑需要空间推理和过程理解能力,目前尚不清楚UMMs是否具备这些能力。为此,我们提出PlanViz,一个专门用于评估图像生成与编辑在电脑任务中表现的新基准。为实现评估目标,我们聚焦于日常生活中频繁出现且需规划的子任务:路线规划、工作图绘制和网页/用户界面展示。通过人工标注问题与参考图像,并采用质量控制流程确保数据质量;为实现细致精确的评估,提出了任务自适应评分方法PlanScore,可衡量生成图像的正确性、视觉质量和效率。实验揭示了当前模型的关键局限与未来研究机遇。

原文摘要 · Abstract (English)

Unified multimodal models (UMMs) have shown impressive capabilities in generating natural images and supporting multimodal reasoning. However, their potential in supporting computer-use planning tasks, which are closely related to our lives, remain underexplored. Image generation and editing in computer-use tasks require capabilities like spatial reasoning and procedural understanding, and it is still unknown whether UMMs have these capabilities to finish these tasks or not. Therefore, we propose PlanViz, a new benchmark designed to evaluate image generation and editing for computer-use tasks. To achieve the goal of our evaluation, we focus on sub-tasks which frequently involve in daily life and require planning. Specifically, three representative sub-tasks are designed: route planning, work diagramming, and web&UI displaying. We address challenges in data quality ensuring by curating human-annotated questions and reference images, and a quality control process. For detailed and exact evaluation, a task-adaptive score, PlanScore, is proposed. The score helps understanding the correctness, visual quality and efficiency of generated images. Through experiments, we highlight key limitations and opportunities for future research on this topic.

图像生成多模态任务规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。