Loom用语义与场合先验生成搭配得体的完整穿搭,速度超快且质量显著提升。
Loom: Hybrid Retrieval-Scoring Outfit Recommendation with Semantic Material Compatibility and Occasion-Aware Embedding Priors
- 融合嵌入检索与多信号评分,按服装类别约束找搭配件。
- 全系统得分0.179,错误率降42%,比随机基线提升3.3倍。
- 无需人工材质分类,用语义推断厚薄兼容性,适合时尚推荐研究者。
我们提出Loom,一个结合神经嵌入检索与结构化评分的穿搭推荐系统,从时尚目录中生成完整连贯的穿搭。给定一件锚点衣物,Loom通过服饰类别约束的近似最近邻搜索,在FashionCLIP嵌入空间中检索互补单品,再利用包含六种信号的多目标函数评分:嵌入相似度、色彩协调性、正式度一致性、场合契合度、风格方向性与穿搭内多样性。提出两项技术克服纯学习或纯规则方法的局限:(1)语义材质权重,利用CLIP嵌入几何结构推断衣物厚重程度以实现叠穿兼容性,无需人工材质分类;(2)氛围/反氛围场合先验,将场合描述文本编码为CLIP空间中的锚向量,通过差异亲和度评分单品。在620件商品的目录上进行消融实验显示,各组件均有显著贡献:全系统平均穿搭得分为0.179,硬性错误率为9.3%;而类别约束随机基线得分为0.054,错误率16.0%。方向重排序是唯一不可替代组件:移除后得分降至0.052,接近随机水平。系统可在消费级硬件上5秒内生成三种风格各异的穿搭。
原文摘要 · Abstract (English)
We present Loom, an outfit recommendation system that combines neural embedding retrieval with structured domain scoring to generate complete, coherent outfits from fashion catalogs. Given an anchor clothing item, Loom retrieves complementary pieces via slot-constrained approximate nearest neighbor search over FashionCLIP embeddings, then scores candidate outfits using a multi-objective function that integrates six signals: embedding similarity, color harmony, formality consistency, occasion coherence, style direction, and within-outfit diversity. We introduce two techniques that address limitations of purely learned or purely rule-based approaches: (1) semantic material weight, which uses CLIP embedding geometry to infer garment heaviness for layer compatibility without hand-coded material taxonomies; and (2) vibe/anti-vibe occasion priors, which embed prose descriptions of occasion contexts as anchor vectors in CLIP space and score items by differential affinity. Ablation experiments on a catalog of 620 items show that each component contributes measurably to outfit quality: the full system achieves a mean outfit score of 0.179 with a 9.3% hard violation rate, compared to 0.054 score and 16.0% violations for a category-constrained random baseline, a 3.3x improvement in score and 42% reduction in violations. Direction reranking is the single indispensable component: removing it drops score to 0.052, essentially equal to random. The system generates three stylistically distinct outfits in under 5 seconds on commodity hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。