构建多主体交互图像生成的评测基准,解决真人像与互动关系难还原的问题。
MIBE: Multi-subject Interaction Benchmark and Evaluator for Personalized Image Generation

- 设计分层数据集,包含6万对标注样本和4千对人工评估黄金数据
- 新评估器MIE在未见过模型上仍达88.4%准确率,超越传统指标
- 适合研究个性化图像生成、多主体交互建模的学者使用
多主体个性化图像生成需精确呈现所有参考身份及其指定互动关系。现有模型常遗漏主体、失真外观或错配互动。现有评价指标针对单主体优化,随主体数量增加,排名区分度下降且与人类偏好脱节。为此,我们提出统一框架MIBE,包含多主体交互基准(MIB)与评估器(MIE)。MIB采用解耦数据设计,包含6万对由VLM标注的银质数据集用于指标训练,以及4千对双盲人工评估的黄金数据集,覆盖多种生成模型,银质数据集跨VLM偏好一致率达95.1%。基于银质数据集,我们构建轻量级参考条件评估器MIE,采用双头排序与诊断目标训练。MIE在黄金数据集上整体成对准确率达0.922,已见模型为0.982,未见模型为0.884。相比多种基线指标(包括CLIP与DINO变体),MIE展现更强跨生成器泛化能力,证明诊断监督可维持排名区分性与人类一致性。
原文摘要 · Abstract (English)
Multi-subject personalized image generation requires the precise rendering of all requested reference identities and their specified interactions based on a guiding prompt. However, state-of-the-art models still struggle with this process, frequently omitting subjects, failing to preserve reference appearances, or misattributing interactions. Furthermore, existing metrics designed primarily for single-subject fidelity cannot reliably capture these errors, suffering severe degradation in ranking separability and failing to align with human preference as the subject count increases. To address this gap, we introduce Multi-subject Interaction Benchmark and Evaluator (MIBE), a unified framework comprising a Multi-subject Interaction Benchmark (MIB) and a Multi-subject Interaction Evaluator (MIE). MIB systematically covers diverse relation types and scene complexities through a decoupled data regime. This consists of a 60K-pair VLM-labeled Silver Set for scalable metric training and a 4K-pair double-blind Human Evaluation Gold Set covering a diverse range of state-of-the-art generators, with the Silver Set reaching 95.1% cross-VLM preference agreement. To demonstrate the utility of this benchmark, we present MIE, a lightweight, reference-conditioned evaluator trained exclusively on the Silver Set with a dual-head ranking and diagnosis objective. MIE exhibits strong cross-generator generalization on the Gold Set, achieving 0.922 overall pairwise accuracy against human preference, including 0.982 on seen generators and 0.884 on unseen generators. By outperforming a broad spectrum of baseline metrics, including CLIP and DINO variants, MIE demonstrates that diagnostic supervision can preserve ranking separability and human alignment where traditional evaluators collapse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。