arXiv:2605.19776cs.CV2026-05

融合偏好与评分构建美学评估新基准,提升模型性能。

Preferences Order, Ratings Anchor: From Fused Expert Aesthetic Ground Truth to Self-Distillation

论文配图:Preferences Order, Ratings Anchor: From Fused Expert Aesthetic Ground Truth to Self-Distillation
图 1 · 摘自论文原文
  • 设计双协议标注框架,统一采集偏好与评分数据。
  • 融合信号后专家判断一致率达90%以上,显著提升评估精度。
  • 仅需单次推理即可超越多数开源模型,适合实际部署。

图像美学评估(IAA)中,成对偏好与点式评分是主流标注方式,但现有基准多仅采用其一,未能在可控条件下验证二者互补性。本文提出PPaint,一个匹配的双协议基准,由15位领域专家(每类5人)对150幅中国画在五个美学维度上同时进行双协议标注,通过局部密集偏好设计收集45,900条成对专家判断及对应评分。匹配设计揭示:偏好生成更一致的序数排序,评分则锚定绝对分数尺度。通过两种独立的偏好转评分方法融合信号,所得融合专家真值使两种构造结果趋于一致。该原则延伸至无标签视觉语言模型训练:PSDistill利用Elo参考池将模型成对判断转化为校准伪分数,并以置信度加权排序优化训练,实现单次推断即得美学评分器。仅在单一画种训练下,蒸馏后的Qwen3-VL-8B在三个类别上平均SRCC从0.504提升至0.709,优于所有开源基线,接近闭源Gemini-3.1-Pro(差值<0.04),且在APDDv2上验证跨域迁移能力。数据集与代码将全部开源。

原文摘要 · Abstract (English)

Pairwise preferences and pointwise ratings are the two dominant annotation protocols in image aesthetic assessment (IAA), yet existing benchmarks adopt only one, leaving their complementarity unmeasured under controlled conditions. We introduce PPaint, a matched dual-protocol benchmark in which 15 domain experts, 5 per category, annotate 150 Chinese paintings under both protocols across five aesthetic dimensions, collecting 45,900 pairwise expert judgments through a locally dense preference design alongside the matched ratings. The matched design reveals complementary strengths: preferences yield more consistent ordinal rankings, while ratings anchor the absolute score scale. Fusing both signals via two independent preference-to-score methods yields a fused expert ground truth on which the two constructions converge to nearly identical scores. The same preference-to-score principle extends to label-free VLM training. PSDistill converts VLM pairwise judgments into calibrated pseudo-scores via an Elo reference pool, and trains the same VLM with confidence-weighted ranking optimization to produce a single-pass aesthetic scorer. Trained on a single painting category, the distilled Qwen3-VL-8B improves mean SRCC from 0.504 to 0.709 across all three categories, outperforming all open-source baselines including the dedicated aesthetic model ArtiMuse and matching closed-source Gemini-3.1-Pro within 0.04 SRCC at single-pass inference cost, with cross-domain transfer further validated on APDDv2. We will release the full PPaint dataset and training code.

美学评估多模态自蒸馏视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。