arXiv:2603.08090cs.CVcs.AI2026-03被引 1

新基准DSH-Bench可精细评估文本生成图像中主体的还原效果。

DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation

论文配图:DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
图 1 · 摘自论文原文
  • 构建分层主体分类体系,覆盖58类细粒度主体。
  • 提出新指标SICS,与人工评估相关性提升9.4%。
  • 提供诊断建议,助力模型训练优化。

主体驱动的文本到图像生成已取得显著进展,旨在根据用户指令合成包含目标主体的新图像。然而,模型评估仍面临重大挑战。现有基准存在三大缺陷:1)主体图像多样性与全面性不足;2)对不同主体难度级别和提示场景下的性能评估粒度不够;3)缺乏可操作的洞察与诊断指导。为此,我们提出DSH-Bench,一个综合基准,通过四项创新实现对主体驱动T2I模型的多视角系统分析:1)采用分层分类采样机制,确保58个细粒度类别全覆盖;2)创新性地划分主体难度等级与提示场景,实现精细化能力评估;3)提出新的主体身份一致性评分(SICS),在量化主体保真度时相比现有方法与人工评估相关性提高9.4%;4)基于基准得出全面诊断洞察,为未来模型训练范式与数据构建策略优化提供关键指导。通过对19个领先模型的广泛实证评估,DSH-Bench揭示了现有方法中此前未被察觉的局限,明确了未来研究发展方向。

原文摘要 · Abstract (English)

Significant progress has been achieved in subject-driven text-to-image (T2I) generation, which aims to synthesize new images depicting target subjects according to user instructions. However, evaluating these models remains a significant challenge. Existing benchmarks exhibit critical limitations: 1) insufficient diversity and comprehensiveness in subject images, 2) inadequate granularity in assessing model performance across different subject difficulty levels and prompt scenarios, and 3) a profound lack of actionable insights and diagnostic guidance for subsequent model refinement. To address these limitations, we propose DSH-Bench, a comprehensive benchmark that enables systematic multi-perspective analysis of subject-driven T2I models through four principal innovations: 1) a hierarchical taxonomy sampling mechanism ensuring comprehensive subject representation across 58 fine-grained categories, 2) an innovative classification scheme categorizing both subject difficulty level and prompt scenario for granular capability assessment, 3) a novel Subject Identity Consistency Score (SICS) metric demonstrating a 9.4\% higher correlation with human evaluation compared to existing measures in quantifying subject preservation, and 4) a comprehensive set of diagnostic insights derived from the benchmark, offering critical guidance for optimizing future model training paradigms and data construction strategies. Through an extensive empirical evaluation of 19 leading models, DSH-Bench uncovers previously obscured limitations in current approaches, establishing concrete directions for future research and development.

图像生成评估基准主体保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。