arXiv:2606.15867cs.CV2026-06被引 1

评测多主体参考图像生成,揭示现有模型在多人物互动时严重失效。

CogCanvas: A Benchmark for Evaluating Multi-Subject Reference-Based Image Generation

论文配图:CogCanvas: A Benchmark for Evaluating Multi-Subject Reference-Based Image Generation
图 1 · 摘自论文原文
  • 构建1952张精选图像,涵盖百位名人、百余物品与背景场景。
  • 五种主流模型在五人组合下物体绑定失败率接近100%。
  • 引入新评估指标,精准衡量背景一致性与人物属性绑定。

多主体参考图像生成需同时保留多个身份特征、绑定人物与物品/服饰,并符合指定背景场景,而当前扩散模型表现脆弱。现有基准仅评估单一维度,无法联合衡量多身份组合、人物-物体交互、背景锚定及空间合理性。我们提出CogCanvas,包含1,952张精心筛选的参考图像,覆盖100位名人、115件独特物品与服饰,以及29个真实世界背景(如地标),并据此构建1,361个组合提示,涵盖2至5人组。其筛选流程结合DINOv2去重、两阶段美学过滤,以及自动化生成结构化交互与位置图作为真值监督。CogCanvas支持三项任务:参考式多人物-物体生成(主任务)、文本到图像组合生成与参考检索,统一采用六轴评估协议。我们提出两项专为多参考设定设计的指标:BG-Sim通过DINOv3特征相似性评估SAM三掩码区域的背景保真度;Attr-VQA利用多模态大模型验证个体属性绑定与人物间交互是否匹配结构图。对五种SOTA方法的测试显示,随着组数从2增至5,所有模型性能显著下降,三人以上时物体/服饰绑定近乎完全失败。

原文摘要 · Abstract (English)

Multi-subject reference-based image generation requires jointly preserving multiple human identities, binding per-person objects and fashion items, and respecting a specified background scene, a regime where current diffusion models remain brittle. Existing benchmarks evaluate only one axis at a time and none jointly captures multi-identity composition with human-object interaction, background grounding, and spatial plausibility. We introduce CogCanvas, a benchmark of 1,952 curated reference images spanning 100 celebrity identities, 115 distinctive objects and fashion items, and 29 real-world background scenes including landmarks, from which we construct 1,361 compositional prompts covering 2-5 person group sizes. The curation pipeline combines DINOv2-based deduplication, two-stage aesthetic filtering, and automated derivation of structured interaction and position graphs that serve as ground-truth supervision. CogCanvas supports three tasks, reference-based multi-human-object generation (primary), text-to-image compositional generation, and reference retrieval, under a unified six-axis evaluation protocol. We introduce two metrics tailored to the multi-reference setting: BG-Sim, which scores background fidelity on SAM 3-masked regions via DINOv3 feature similarity, and Attr-VQA, which uses a multimodal LLM to verify per-subject attribute binding and inter-person interactions against the structured graphs. Benchmarking five SOTA methods reveals that every model degrades substantially as group size grows from 2 to 5, with near-complete failure on object/fashion binding beyond three subjects.

图像生成多主体评估基准扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。