arXiv:2605.06170cs.CV2026-05

动态生成提示词,防止模型过拟合,让图像生成评估更公平可靠。

DynT2I-Eval: A Dynamic Evaluation Framework for Text-to-Image Models

论文配图:DynT2I-Eval: A Dynamic Evaluation Framework for Text-to-Image Models
图 1 · 摘自论文原文
  • 构建可控制的语义空间,自动组合新提示词
  • 持续刷新提示集,降低对固定数据集的依赖
  • 适合评测文本到图像模型的真实泛化能力

现有文本到图像(T2I)评估基准多依赖固定提示集,易因重复使用导致过拟合和基准污染。本文提出 DynT2I-Eval,一个全自动动态评估框架。该框架从长篇描述中构建结构化视觉语义空间,将提示分解为可控制维度(如主体、逻辑约束、环境、构图),支持任务特定空间下的持续新提示生成与难度感知采样。评估涵盖文本对齐性、感知质量和美学表现。异构输出通过提示条件下的成对比较统一,结合动态调度器、微批次聚合与加权贝叶斯更新,确保在提示分布变化和模型注入情况下仍保持稳定在线排行榜。独立采样的提示流实验表明,持续刷新提示能有效降低提示集特异性调优的影响。仿真与消融实验进一步验证了该排名框架在冷启动收敛、后期发现能力与长期排名保真度之间达到良好平衡。

原文摘要 · Abstract (English)

Existing text-to-image (T2I) benchmarks largely rely on fixed prompt sets, leaving them vulnerable to overfitting and benchmark contamination once publicly released and repeatedly reused. In this work, we propose DynT2I-Eval, a fully automated dynamic evaluation framework for T2I models. It constructs a structured visual semantic space from long-form descriptions, decomposing prompts into controllable dimensions (e.g., subject, logical constraint, environment, and composition). This enables the continuous generation of fresh prompts via task-specific spaces and difficulty-aware sampling. DynT2I-Eval evaluates model performance across text alignment, perceptual quality, and aesthetics. Heterogeneous outputs are unified into prompt-conditioned pairwise comparisons, allowing a dynamic scheduler, micro-batch aggregation, and weighted Bayesian updates to maintain a stable online leaderboard despite changing prompt distributions and model injection. Experiments with independently sampled prompt streams demonstrate that continually refreshed prompts provide a robust evaluation protocol, reducing the impact of prompt-set-specific tuning. Simulations and ablations further confirm that the proposed ranking framework achieves a strong balance among cold-start convergence, late-entry discovery, and long-run ranking fidelity.

图像生成评估框架动态测试文本到图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。