让时尚编辑风格可被看见,通过图像识别品牌、年代与色彩传统
FASH-iCNN: Making Editorial Fashion Identity Inspectable Through Multimodal CNN Probing

- 用多模态卷积网络分析时装秀图,提取品牌、年代和色彩传统
- 仅凭服装图像就能以88.6%准确率识别年代,误差仅2.2年
- 发现纹理和亮度是编辑风格的核心信号,颜色影响较小
时尚AI系统常隐式编码特定品牌、编辑及历史时期的审美逻辑却未公开。我们提出FASH-iCNN,基于1991-2024年间15个品牌的87,547张Vogue秀场图像训练,使这种文化逻辑可被探查。给定一件服装照片,系统可识别其所属品牌(14家,top-1准确率78.2%)、年代(34年,top-1准确率58.3%,平均误差2.2年)及色彩传统。仅用服装图像即实现高精度识别。探针分析显示:移除颜色仅导致品牌识别准确率下降10.6个百分点,而移除纹理则下降37.6个百分点,表明纹理与亮度是编辑身份的主要载体。FASH-iCNN将编辑文化视为信号而非噪声,揭示每个预测背后所蕴含的品牌、编辑与历史时刻。
原文摘要 · Abstract (English)
Fashion AI systems routinely encode the aesthetic logic of specific houses, editors, and historical moments without disclosing it. We present FASH-iCNN, a multimodal system trained on 87,547 Vogue runway images across 15 fashion houses spanning 1991-2024 that makes this cultural logic inspectable. Given a photograph of a garment, the system recovers which house produced it, which era it belongs to, and which color tradition it reflects. A clothing-only model identifies the fashion house at 78.2% top-1 across 14 houses, the decade at 88.6% top-1, and the specific year at 58.3% top-1 across 34 years with a mean error of just 2.2 years. Probing which visual channels carry this signal reveals a sharp dissociation: removing color costs only 10.6pp of house identity accuracy, while removing texture costs 37.6pp, establishing texture and luminance as the primary carriers of editorial identity. FASH-iCNN treats editorial culture as the signal rather than background noise, identifying which houses, eras, and color traditions shaped each output so that users can see not just what the system predicts but which houses, editors, and historical moments are encoded in that prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。