arXiv:2601.10369cs.CV2026-01

用分层选择的多模态大模型,精准评估人体姿态编辑的真假与质量。

Fine-Grained Human Pose Editing Assessment via Layer-Selective MLLMs

  • 基于分层敏感性分析,选出最适合姿态评估的特征层。
  • 在1700个样本上实现高精度真假判断与多维度质量评分。
  • 适合关注生成内容真实性与细节质量的研究者使用。

文本引导的人体姿态编辑在AIGC应用中备受关注,但仍存在结构异常和生成伪影问题。现有评估指标常将真实性检测与质量评估分离,难以提供针对姿态的细粒度洞察。为此,我们构建了HPE-Bench基准,包含17个先进编辑模型产生的1,700个标准化样本,提供真实性标签与多维质量评分。同时提出一种基于分层选择式多模态大模型(MLLM)的统一框架,通过对比LoRA微调与新颖的层敏感性分析(LSA)机制,识别最优姿态评估特征层。该框架在真实性检测与多维质量回归任务中均表现优异,有效弥合了检测与评估之间的鸿沟。

原文摘要 · Abstract (English)

Text-guided human pose editing has gained significant traction in AIGC applications. However,it remains plagued by structural anomalies and generative artifacts. Existing evaluation metrics often isolate authenticity detection from quality assessment, failing to provide fine-grained insights into pose-specific inconsistencies. To address these limitations, we introduce HPE-Bench, a specialized benchmark comprising 1,700 standardized samples from 17 state-of-the-art editing models, offering both authenticity labels and multi-dimensional quality scores. Furthermore, we propose a unified framework based on layer-selective multimodal large language models (MLLMs). By employing contrastive LoRA tuning and a novel layer sensitivity analysis (LSA) mechanism, we identify the optimal feature layer for pose evaluation. Our framework achieves superior performance in both authenticity detection and multi-dimensional quality regression, effectively bridging the gap between forensic detection and quality assessment.

姿态编辑多模态大模型评估基准真实性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。