arXiv:2602.00105cs.CVcs.AI2026-02

评测前沿图像编辑模型的真实可靠性,揭示低价模型因重试成本更高反而更贵。

HYPE-EDIT-1: Benchmark for Measuring Reliability in Frontier Image Editing Models

  • 构建100任务基准,每任务生成10次输出评估成功率与重试成本。
  • 真实有效成本范围0.66-1.42美元/成功编辑,低单价模型总成本更高。
  • 适合关注工业级图像编辑落地成本的设计师和开发者使用。

公开演示的图像编辑模型通常展示最佳结果;真实工作流需承担重试和人工审核时间成本。我们提出HYPE-EDIT-1,一个包含100个参考式营销/设计编辑任务的基准,采用二分类通过/失败判定。每个任务生成10次独立输出,以估算单次尝试通过率、通过率@10、在重试上限下的预期尝试次数,以及融合模型价格与人工审核时间的有效成功成本。我们发布50个公开任务,并保留50个私有任务用于服务器端评估,同时提供标准化JSON格式和工具支持视觉语言模型与人工评判。在评估的模型中,单次尝试通过率介于34%-83%,有效成本每成功编辑为0.66-1.42美元。那些单图价格较低的模型,因需更多重试和人工审核,总有效成本反而更高。

原文摘要 · Abstract (English)

Public demos of image editing models are typically best-case samples; real workflows pay for retries and review time. We introduce HYPE-EDIT-1, a 100-task benchmark of reference-based marketing/design edits with binary pass/fail judging. For each task we generate 10 independent outputs to estimate per-attempt pass rate, pass@10, expected attempts under a retry cap, and an effective cost per successful edit that combines model price with human review time. We release 50 public tasks and maintain a 50-task held-out private split for server-side evaluation, plus a standardized JSON schema and tooling for VLM and human-based judging. Across the evaluated models, per-attempt pass rates span 34-83 percent and effective cost per success spans USD 0.66-1.42. Models that have low per-image pricing are more expensive when you consider the total effective cost of retries and human reviews.

图像编辑基准评测成本分析可靠性的

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。