arXiv:2605.09030cs.CVcs.LG2026-05

发现风格相似度评分失效,提出诊断方法并改进评估可靠性

When Style Similarity Scores Fail: Diagnosing Raw CSD Cosine in Artist-Style Evaluation

论文配图:When Style Similarity Scores Fail: Diagnosing Raw CSD Cosine in Artist-Style Evaluation
图 1 · 摘自论文原文
  • 提出'判别间隙'诊断法,检验风格余弦是否可绝对区分同艺术家作品
  • 在91位艺术家中有15位的原始评分无法可靠判断风格一致
  • 采用CSLS+位置插值可将验证准确率提升至0.905,适合艺术风格评估者

对比风格描述器(CSD)输出空间中的原始余弦相似度,现被广泛视为文生图与风格模仿评估中风格保真度的绝对、校准分数。本文引入‘判别间隙’——一种无需原型、无需阈值、基于语料内部的诊断方法,用于检验对比风格余弦能否在候选艺术家语料中实现‘相同 vs. 不同’的绝对解释。在包含1799幅作品、91位艺术家的公开语料上,原始CSD余弦在成对层面有23/91位艺术家出现负点估计间隙(经自助法验证仅2/91稳健),在聚合池评分模式下也有15/91位艺术家出现负间隙。使用冻结主干网络上的CSLS读出可将负间隙数量降至4/91;结合位置嵌入插值至336像素后,跨25个艺术家不相交划分的无监督配对验证AUC从0.883提升至0.905。我们称此基于诊断驱动的读出协议为CSD+,并非新编码器。跨主干检查在CLIP-ViT-L/14、SigLIP-large和DINOv2-Large上重现相同失败模式,表明该残差反映的是四种主干共有的局限性,而非仅限于CSD。实践建议:报告CSD余弦前,应在目标语料上运行诊断;若失败,最小修正方式为使用CSLS。

原文摘要 · Abstract (English)

Raw cosine in the 768-dimensional output space of the Contrastive Style Descriptor (CSD) is now widely read as an absolute, calibrated style-fidelity score for text-to-image and style-imitation evaluation. We introduce the discrimination gap, a corpus-internal, prototype-free and threshold-free diagnostic that tests whether contrastive style cosines admit an absolute same-versus-different interpretation on a candidate artist corpus. On a 1799-artwork, 91-artist public-domain corpus, raw CSD cosine yields negative point-estimate gaps for $23/91$ artists at the pairwise level ($2/91$ robust under bootstrap) and for $15/91$ in the aggregated-pool scoring regime style-fidelity evaluations typically use. CSLS readout on the frozen backbone reduces the aggregated negative-gap count to $4/91$; combined with positional-embedding interpolation to $336$ pixels it raises unsupervised pair-verification AUC from $0.883$ to $0.905$ across $25$ artist-disjoint splits. We refer to this diagnostic-driven readout protocol on the frozen backbone (CSLS as default, pos-interp $336$ as the stronger optional setting) as CSD+, not a new encoder.A cross-backbone check on CLIP-ViT-L/14, SigLIP-large and DINOv2-Large reproduces the same shared-tradition failure pattern, providing evidence that the residual reflects a shared limitation of the four backbones we tested rather than a CSD-specific artefact. Practical implication: before reporting CSD cosine as an absolute style-fidelity score, run the diagnostic on the candidate corpus; CSLS is the minimal correction when it fails.

风格评估对比学习模型诊断图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。