构建开放基准评估人机协作的视觉指令理解与编辑能力。
CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences

- 提出CV-Arena基准,覆盖16类真实图像任务,含12K高分辨率指令对。
- 采用人机协同评估协议,模型在物理合理性与细节保留上仍有明显不足。
- 适合研究专业级视觉任务的模型泛化与闭环推理,尤其关注指令遵循能力。
指令引导的图像编辑正成为视觉工作通用接口,但现有基准多局限于外观修改,未能涵盖专业流程中的多样化任务。本文将指令式计算机视觉问题求解定义为更广义的图像编辑:给定真实输入图像和自然语言指令,系统需生成符合请求变换且满足显式保留、几何、物理及可用性约束的输出。为此,我们构建了开源基准CV-Arena,包含12,000对高分辨率真实图像指令对,覆盖16种基于指令的视觉任务类型,通过CogRetriever双轨检索与筛选流程(含定向网络搜索、智能查询优化、验证与可追溯性)构建。为在保持人类判断力的前提下实现大规模评估,提出主动埃洛(Active Elo)协议,结合逻辑门控、多维度视觉语言模型评估器CV-Judge,用于剔除明显失败案例并处理高置信度对比;高质量近似对比则交由专家评审。混合人类与AI监督后,通过可靠性加权埃洛更新聚合评分。对21个系统(含专有、开源与代理模型)的全面评估揭示,当前模型在指令遵循、物理推理、结构控制与细粒度细节保留方面仍存在显著差距。进一步开发轻量级代理模型CV-Agent,整合规划、编辑与验证,证明闭环推理是实现专业级指令跟随视觉编辑的可行方向。
原文摘要 · Abstract (English)
Instruction-guided image editing is becoming a general interface for visual work, yet existing benchmarks still focus largely on narrow appearance edits and do not fully capture the diversity of real-image tasks in professional workflows. Here, we define instructional computer vision problem solving as a broader formulation of image editing: given a real input image and a natural-language instruction, a system must produce an edited output that realizes the requested transformation while satisfying explicit preservation, geometric, physical, and usability constraints. We introduce CV-Arena, an open benchmark designed to evaluate this capability at professional scales. CV-Arena contains 12K high-resolution real-image instruction pairs spanning 16 instruction-based visual task types, constructed using CogRetriever, a dual-track retrieval-and-curation pipeline that combines targeted web search, agentic query refinement, verification, and traceability. To evaluate models at scale while preserving human fidelity, we propose Active Elo, a human-AI collaborative preference protocol that leverages CV-Judge, a logic-gated, multi-dimensional VLM evaluator, to reject clear failures and resolve high-confidence comparisons; and to route close, high-quality comparisons to expert raters. Mixed human and AI supervision is then aggregated through reliability-weighted Elo updates. Our comprehensive evaluation of 21 systems, including proprietary, open-source, and agentic models, on CV-Arena reveals persistent gaps in instruction adherence, physical reasoning, structural control, and fine-grained detail preservation. We further develop CV-Agent, a lightweight agentic model that combines planning, editing, and verification, and demonstrate that closed-loop reasoning is a promising direction for professional-grade instruction-following visual editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。