为文本生成图像设计适配不同评估能力的标注方法,提升评价可靠性。
Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation

- 根据评估技能特性定制标注策略,而非统一使用同一套方法。
- 相比传统方式,标注一致性提升,跨模型评价更稳定。
- 可自动化的流程支持细粒度、空间定位的反馈,适合大规模评估。
文本到图像(T2I)生成技术快速发展,模型性能差距缩小,可靠评估愈发关键。现有评估多采用统一标注方式(如李克特量表或二分类问答),但未考虑不同评估技能的本质差异。本文提出技能对齐标注,使标注策略匹配各评估技能的内在特征。系统对比显示,该方法能产生更一致的评估信号,显著提升标注者间一致性,并增强跨模型评估稳定性。最后,我们构建自动化流水线实现该评估协议,支持可扩展、细粒度且具备空间定位反馈的评估。研究证明,优化评估基础可提升可靠性与效率,无需单纯增加人工标注投入。希望推动评估协议作为模型可靠评估核心组成部分的进一步研究。
原文摘要 · Abstract (English)
Text-to-image (T2I) generation has advanced rapidly, making reliable evaluation critical as performance differences between models narrow. Existing evaluation practices typically apply uniform annotation mechanisms, such as Likert-scale or binary question answering (BQA), across heterogeneous evaluation skills, despite fundamental differences in their nature. In this work, we revisit T2I evaluation through the lens of skill-aligned annotation, where annotation strategies reflect the underlying characteristics of each evaluation skill. We systematically compare skill-aligned annotation against uniform baselines and show that it produces more consistent evaluation signals, with higher inter-annotator agreement and improved stability across models. Finally, we present an automated pipeline that instantiates the proposed evaluation protocol, enabling scalable and fine-grained evaluation with spatially grounded feedback. Our work highlights that improving the foundations of image evaluation can increase reliability and efficiency without simply scaling annotation effort. We hope this motivates further research on refining evaluation protocols as a central component of reliable model assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。