arXiv:2608.16765cs.CVcs.AI2026-08中稿 · ACM Multimedia 202…

提出新基准,拆解多参考图像生成的底层能力并精准定位模型短板。

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

论文配图:TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
图 1 · 摘自论文原文
  • 用四种原子操作构建可组合的提示公式,量化任务复杂度。
  • 9个主流模型中,属性绑定能力最弱,最佳仅0.74分。
  • 适合研究生成模型细粒度能力与故障诊断的研究者。

尽管统一多模态模型在多参考图像生成方面取得进展,但现有基准仍围绕预定义任务类型(如“主体组合”)组织,难以适应该组合场景,导致覆盖碎片化、复杂度不可控且缺乏诊断价值。我们发现多样化的多参考任务共享一组基本操作,提出以能力为导向的视角,形式化定义四种算子:锚定(f)、解耦(g)、应用(⊕)和组合(C)。任意多参考提示均可表示为这些算子的组合公式,其结构复杂度由算子槽位数衡量。基于此,我们构建了TRACE-Bench,包含约1,600个评估样本,涵盖1–8个槽位,基于631个公式模板和约4,000张跨艺术风格与现实主题的参考图像。公式结构直接支持算子对齐的性能评分与递归故障定位的诊断树分析。对9个领先模型的评估揭示了全局评分无法察觉的洞见:主要瓶颈在于解耦(g)与属性绑定(⊕),而非场景级组合(C),即使最优模型在属性保真度上也仅得0.74分。

原文摘要 · Abstract (English)

Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor ($f$), Disentangle ($g$), Apply ($\oplus$), and Compose ($C$). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement ($g$) and attribute binding ($\oplus$) rather than scene-level composition ($C$), with even the best model scoring only 0.74 on attribute fidelity. Project page: https://amuseum-whr.github.io/TraceBench

图像生成能力分解模型诊断多参考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。