arXiv:2508.06152cs.CV2025-08被引 1

构建用户导向的图文生成评估基准,区分可量化与抽象语义评价。

VISTAR:A User-Centric and Role-Driven Benchmark for Text-to-Image Evaluation

  • 分两层评估:物理属性用确定性指标,抽象语义用约束型视觉语言模型打分。
  • 通过1.5万次人工对比验证2845个提示,抽象语义评分准确率达85.9%。
  • 引入7类用户角色,揭示无通用最优模型,助力场景化部署决策。

我们提出VISTAR,一个以用户为中心、多维度的文本到图像生成评估基准,解决现有评估指标的局限性。VISTAR采用两级混合范式:对可量化的物理属性(如文本渲染、光照)使用确定性、可脚本化指标;对抽象语义(如风格融合、文化契合度)引入新型分层加权正负问答(HWPQ)机制,借助受限的视觉语言模型进行评估。基于120位专家参与的德尔菲研究,定义了7种用户角色和9个评估角度,构建包含2,845个经验证提示的基准,经超过15,000次人工成对比较验证。所提指标与人类判断高度一致(>75%),其中HWPQ在抽象语义任务上达到85.9%准确率,显著优于传统VQA基线。对主流模型的全面评估表明,不存在普遍最优模型,角色加权得分会重新排序排名,为特定领域部署提供可操作指导。所有资源均已公开,推动可复现的T2I评估。

原文摘要 · Abstract (English)

We present VISTAR, a user-centric, multi-dimensional benchmark for text-to-image (T2I) evaluation that addresses the limitations of existing metrics. VISTAR introduces a two-tier hybrid paradigm: it employs deterministic, scriptable metrics for physically quantifiable attributes (e.g., text rendering, lighting) and a novel Hierarchical Weighted P/N Questioning (HWPQ) scheme that uses constrained vision-language models to assess abstract semantics (e.g., style fusion, cultural fidelity). Grounded in a Delphi study with 120 experts, we defined seven user roles and nine evaluation angles to construct the benchmark, which comprises 2,845 prompts validated by over 15,000 human pairwise comparisons. Our metrics achieve high human alignment (>75%), with the HWPQ scheme reaching 85.9% accuracy on abstract semantics, significantly outperforming VQA baselines. Comprehensive evaluation of state-of-the-art models reveals no universal champion, as role-weighted scores reorder rankings and provide actionable guidance for domain-specific deployment. All resources are publicly released to foster reproducible T2I assessment.

图文生成评估基准用户角色语义评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。