arXiv:2607.27616cs.CV2026-07

评测多人互动编辑的生理合理性,发现现有模型常出现肢体穿插、缺失等问题。

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

论文配图:MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing
图 1 · 摘自论文原文
  • 构建2500样本视频挖掘三元组数据集,覆盖14类互动和4种接触密度。
  • 提出双维度评估指标:身体完整性与接触几何匹配度,最高得分仅0.65和0.72。
  • 实验证明人工评分优于零样本视觉语言模型,且结果稳定可靠。

文本到图像及个性化编辑模型已能轻松生成高保真单人图像。然而,在将多名指定人物放入共享互动动作(如拥抱、背负、搏斗)时仍存在严重缺陷:肢体融合、虚构肢体、身体穿插等。现有评估普遍忽略这些解剖与几何问题,而基于视觉语言模型的评判清单在互动任务上得分饱和,但人类仍可轻易识别错误。本文提出MPIE-Bench,一个包含2500个样本的视频挖掘编辑三元组基准,覆盖405个场景、14种互动类别及四种接触密度(C0-C3)。同时提出MPIE-Eval,其两个新维度通过冻结的公开多人网格重建模型评估接触时间几何:解剖性要求每个类人实体均由完整重建的身体解释;互动性要求身体间穿透与表面距离匹配指令所要求的接触。在十种编辑器中,网格解剖性最高为0.65,网格互动性最高为0.72,无一同时表现优异;而视觉语言模型对相同图像评分均高于0.95。五名评审者研究证实,这两个维度比零样本视觉语言模型更贴近人类判断,且在所有权重与阈值消融下排名保持一致。

原文摘要 · Abstract (English)

Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.

图像编辑多人互动评估基准解剖合理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。