用多属性隐空间可视化腕表,支持风格探索与混合筛选。
A Multi-Attribute Latent Space for Visual Analysis of Watches

- 分属性构建表盘颜色、设计图谱,结合表款类型组织语义结构。
- 融合多特征的嵌入空间使用户可直观对比不同风格与功能组合。
- 交互界面支持按例搜索、细节查看,适合收藏者与设计新手使用。
我们提出一种设计原理、嵌入模型及交互式视觉分析系统,用于通过异构视觉与语义属性探索大规模腕表收藏。现有电商和目录界面虽支持元数据过滤,但缺乏对视觉相似性、风格替代及复合审美-功能标准的开放探索支持。因此,我们为表盘颜色与设计分别构建属性图谱,并以表款类型作为显式语义组织器。表盘通过U-Net分割,表款类型由Vision Transformer预测,颜色采用共享的CIELAB参考色板表示,表盘结构则用基于梯度的图像描述符刻画。我们扩展UMAP,通过联合属性特异性邻域图与统一概率目标,并引入类别感知布局项,将全局类型结构与局部视觉邻域分离。最终生成的映射在交互界面中实现空间导航、元数据过滤、细节查看与示例搜索插入。通过参数分析、运行时测量及专家与新手的定性试点研究评估,结果表明系统有助于发现与比较,但也揭示了可扩展性评估、示例搜索验证的局限性,以及需更广泛领域研究的必要性。本文明确讨论这些限制,并为跨异质视觉集合的多属性隐空间可视化提供设计启示。
原文摘要 · Abstract (English)
We present a design rationale, embedding model, and interactive visual-analysis system for exploring large wristwatch collections through heterogeneous visual and semantic attributes. The system addresses a common limitation of catalog and e-commerce interfaces: users can filter by metadata, but they receive little support for open-ended exploration of visual similarity, stylistic alternatives, and mixed aesthetic-functional criteria. We therefore represent watches with separate attribute graphs for dial color and dial design, while using watch type as an explicit semantic organizer. Dials are segmented with a U-Net, watch types are predicted with a Vision Transformer, colors are represented through a shared CIELAB reference palette, and dial structure is described with a gradient-based image descriptor. We extend UMAP by combining attribute-specific neighborhood graphs in a unified probabilistic objective and by adding a class-aware layout term that separates global type structure from local visual neighborhoods. The resulting map is exposed in an interactive interface with spatial navigation, metadata filtering, detail inspection, and search-by-example insertion. We evaluate the approach through parameter analysis, runtime measurements, and a qualitative pilot study with watch experts and novices. The results suggest that the system supports discovery and comparison, while also revealing limitations in scalability assessment, search-by-example validation, and the need for broader domain studies. We explicitly discuss these limitations and derive design implications for multi-attribute latent-space visualization across heterogeneous visual collections.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。