提出新评估方法,让机器评价更贴近人类对概念定制的偏好。
Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
- 将评估拆解为细粒度维度,用多模态大模型逐项打分。
- 在单概念到多人交互的多种任务上表现优于现有方法。
- 适合研究生成模型评估、个性化内容定制的学者使用。
概念定制的评估极具挑战性,需全面衡量生成结果与提示词及概念图像的一致性。评估多个概念比单一概念更复杂,不仅需关注每个概念本身,还需分析概念间的相互作用。尽管人类能直观判断生成图像质量,但现有指标往往过于狭窄或笼统,导致与人类偏好不一致。为此,我们提出分解式GPT评分(D-GPTScore),将评估标准细分为多个维度,并利用多模态大语言模型进行各维度评估。同时,我们发布人类偏好对齐的概念定制基准(CC-AlignBench),包含单概念与多概念任务,支持从个体动作到多人交互的多阶段评估。实验表明,该方法在该基准上显著优于现有方法,与人类偏好相关性更高。本工作确立了概念定制评估的新标准,并指出了未来研究的关键挑战。相关数据集与工具已开源:https://github.com/ReinaIshikawa/D-GPTScore。
原文摘要 · Abstract (English)
Evaluating concept customization is challenging, as it requires a comprehensive assessment of fidelity to generative prompts and concept images. Moreover, evaluating multiple concepts is considerably more difficult than evaluating a single concept, as it demands detailed assessment not only for each individual concept but also for the interactions among concepts. While humans can intuitively assess generated images, existing metrics often provide either overly narrow or overly generalized evaluations, resulting in misalignment with human preference. To address this, we propose Decomposed GPT Score (D-GPTScore), a novel human-aligned evaluation method that decomposes evaluation criteria into finer aspects and incorporates aspect-wise assessments using Multimodal Large Language Model (MLLM). Additionally, we release Human Preference-Aligned Concept Customization Benchmark (CC-AlignBench), a benchmark dataset containing both single- and multi-concept tasks, enabling stage-wise evaluation across a wide difficulty range -- from individual actions to multi-person interactions. Our method significantly outperforms existing approaches on this benchmark, exhibiting higher correlation with human preferences. This work establishes a new standard for evaluating concept customization and highlights key challenges for future research. The benchmark and associated materials are available at https://github.com/ReinaIshikawa/D-GPTScore.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。