构建万级时尚数据集,实现多任务统一的对话式智能穿搭理解
OmniFashion: Towards Generalist Fashion Intelligence via Multi-Task Vision-Language Learning
- 基于全局到局部属性标注,构建百万级时尚数据集FashionX
- 在多任务检索与对话任务中表现优异,跨任务泛化能力强
- 适合需要统一理解与交互的时尚推荐系统研发者
时尚智能涵盖检索、推荐、识别与对话等多种任务,但受限于监督信号碎片化和时尚标注不完整,难以形成一致的视觉-语义结构,阻碍了现有视觉语言模型作为通用时尚智能中枢的发展。为此,我们构建了包含百万级样本的FashionX数据集,全面标注服饰搭配中的可见单品,并从全局到局部层级组织属性信息。在此基础上,提出OmniFashion统一框架,通过统一的时尚对话范式整合多样任务,支持多任务推理与交互对话。在多个子任务与检索基准测试中,OmniFashion展现出强任务准确率与跨任务泛化能力,验证了其通往可扩展、对话驱动的通用时尚智能的可行性。
原文摘要 · Abstract (English)
Fashion intelligence spans multiple tasks, i.e., retrieval, recommendation, recognition, and dialogue, yet remains hindered by fragmented supervision and incomplete fashion annotations. These limitations jointly restrict the formation of consistent visual-semantic structures, preventing recent vision-language models (VLMs) from serving as a generalist fashion brain that unifies understanding and reasoning across tasks. Therefore, we construct FashionX, a million-scale dataset that exhaustively annotates visible fashion items within an outfit and organizes attributes from global to part-level. Built upon this foundation, we propose OmniFashion, a unified vision-language framework that bridges diverse fashion tasks under a unified fashion dialogue paradigm, enabling both multi-task reasoning and interactive dialogue. Experiments on multi-subtasks and retrieval benchmarks show that OmniFashion achieves strong task-level accuracy and cross-task generalization, highlighting its offering of a scalable path toward universal, dialogue-oriented fashion intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。