arXiv:2603.21937cs.CV2026-03被引 2

构建多主体生成属性错绑的评测基准,揭示现有模型隐藏错误。

MultiBind: A Benchmark for Attribute Misbinding in Multi-Subject Generation

  • 基于真实多人照片构建带标注的多主体生成数据集
  • 提出分维度混淆评估法,精准识别属性错绑类型
  • 适合研究可控图像生成与多主体一致性问题的学者

主体驱动的图像生成正需对单张图像中多个实体实现细粒度控制。在多参考工作流中,用户可提供多个主体图像、背景参考图及带实体索引的长提示词以控制场景内多人。此情境下关键失败模式为跨主体属性错绑:属性被错误保留、编辑或迁移至错误主体。现有评测与指标主要关注整体保真度或单主体自相似性,难以诊断此类错误。本文提出 MultiBind,一个基于真实多人照片构建的基准数据集。每个实例包含按槽位排序的主体裁剪图、掩码与边界框、标准化主体参考图、修复后的背景参考图,以及由结构化标注生成的密集实体索引提示词。我们还提出一种维度级混淆评估协议,通过人脸身份、外观、姿态、表情等专用模块匹配生成主体与真实槽位,并计算槽位间相似度。通过减去对应真实相似度矩阵,方法可分离自退化与真正跨主体干扰,暴露漂移、交换、主导、混合等可解释的失败模式。在现代多参考生成器上的实验表明,MultiBind 能揭示传统重建指标遗漏的绑定错误。

原文摘要 · Abstract (English)

Subject-driven image generation is increasingly expected to support fine-grained control over multiple entities within a single image. In multi-reference workflows, users may provide several subject images, a background reference, and long, entity-indexed prompts to control multiple people within one scene. In this setting, a key failure mode is cross-subject attribute misbinding: attributes are preserved, edited, or transferred to the wrong subject. Existing benchmarks and metrics largely emphasize holistic fidelity or per-subject self-similarity, making such failures hard to diagnose. We introduce MultiBind, a benchmark built from real multi-person photographs. Each instance provides slot-ordered subject crops with masks and bounding boxes, canonicalized subject references, an inpainted background reference, and a dense entity-indexed prompt derived from structured annotations. We also propose a dimension-wise confusion evaluation protocol that matches generated subjects to ground-truth slots and measures slot-to-slot similarity using specialists for face identity, appearance, pose, and expression. By subtracting the corresponding ground-truth similarity matrices, our method separates self-degradation from true cross-subject interference and exposes interpretable failure patterns such as drift, swap, dominance, and blending. Experiments on modern multi-reference generators show that MultiBind reveals binding failures that conventional reconstruction metrics miss.

图像生成多主体控制评测基准属性错绑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。