对比6种视觉语言模型对形状情感特征的判断一致性,发现结果参差不齐。
Do Vision-Language Models Agree on the Affective Qualities of Shape? A Cross-Model Audit for Generative Design Interfaces

- 用情感词对排序3D物体,检验模型对形状美感理解是否一致
- 平均相关性0.36,高于随机水平但低于几何特征基准
- 模型间共识受类别语义方向匹配度影响,非单纯形状差异
生成式设计界面越来越多地引入语义控制,如“更优雅”或“更极简”,通常由视觉语言模型(VLM)编码。一个实际问题是:当前最先进的VLM是否在相同概念上对物体表示一致?我们审计了6个VLM,通过将未贴图的3D物体沿Kansei形容词对进行排序来评估,其中Kansei描述产品形态带来的情感印象,每个维度定义为两个极点文本表征的差异。几何词对作为正向对照,无关形容词对建立经验零假设。在ShapeNet数据库的10个类别中,情感轴的收敛性高于零假设(平均成对秩相关系数0.36 vs. 0.14),但低于几何上限(0.44)。模型间的一致性部分且极不均匀:在所有类别共享的三个轴上,平均收敛度从书架的0.21到罐子的0.51不等。收敛性主要取决于类别表征变异是否与所评估语义方向一致,而非形状总体变化程度。跨模型收敛并不意味着与人类判断一致。基于这些发现,我们实现了一个用户界面原型,展示如何利用该审计指导特定物体类别应暴露哪些Kansei描述符作为控制项,哪些应隐藏。
原文摘要 · Abstract (English)
Generative design interfaces increasingly expose semantic controls that let users steer output with concepts such as "more elegant" or "more minimalist," typically encoded by a vision-language model (VLM). A practical question is whether state-of-the-art VLMs represent objects consistently in terms of the same concept. We audit 6 VLMs by ranking untextured 3D objects along Kansei adjective pairs, where Kansei describes affective impressions of product form, with each axis defined as the difference between the text representations of its two poles. Geometric pairs serve as positive controls, and pairs of unrelated adjectives establish an empirical null. Across 10 categories of ShapeNet database, affective axes converge above the null (mean pairwise rank correlation 0.36 vs. 0.14) but below the geometric ceiling (0.44). The agreement between models is partial and highly uneven: on the three axes shared by all categories, mean convergence ranges from 0.21 for bookshelves to 0.51 for jars. Convergence depends primarily on whether a category's representational variation aligns with the semantic direction being evaluated, rather than simply on how much the objects vary in shape overall. Cross-model convergence does not imply agreement with human judgments. Based on our findings, we implement a UI prototype that shows how the audit can inform which Kansei descriptors to expose as controls for a given object class and which to withhold.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。