arXiv:2606.13427cs.CV2026-06

构建越南传统服饰检索基准,支持草图+文字联合查询。

VietFashion: Benchmarking Sketch-Text Composed Image Retrieval for Cultural Outfits

论文配图:VietFashion: Benchmarking Sketch-Text Composed Image Retrieval for Cultural Outfits
图 1 · 摘自论文原文
  • 结合手绘草图与文本描述,实现文化服饰的多模态检索。
  • 包含21,000张真实感图像及对应标注,覆盖650个初始草图。
  • 引入多目标检索机制,适合研究文化语义与细粒度设计意图。

文化服饰对视觉检索系统构成独特挑战,因其身份依赖于细微的结构与象征性细节,现有标准AI模型难以捕捉。我们提出VietFashion,一个聚焦于越南传统服饰Ao Dai的草图-文本联合图像检索基准。该数据集初始包含650张手绘草图,通过生成模型扩展至超过21,000张写实图像,并配有对齐的图文描述。文本提示基于时尚杂志提取,确保真实性与多样性。为反映设计意图的固有模糊性,VietFashion采用多目标检索设置,即单个查询可对应多个有效结果。我们建立了标准化评估协议,并对当前先进多模态检索方法进行了基准测试。实验揭示了在建模细粒度文化语义与跨模态组合方面的显著性能差距,确立了VietFashion作为细粒度时尚检索的挑战性基准。数据集已公开:https://hng0303.github.io/VietFashion。

原文摘要 · Abstract (English)

Cultural garments pose a unique challenge for visual retrieval systems, as their identity often depends on subtle structural and symbolic details that are poorly captured by standard AI models. We introduce VietFashion, a new benchmark for sketch-text composed image retrieval centered on the Ao Dai, a traditional Vietnamese garment. VietFashion enables designers and researchers to retrieve culturally meaningful outfits using a combination of hand-drawn sketches, which convey garment structure, and textual descriptions, which encode cultural semantics. The dataset is initialized with 650 sketches and expanded using generative models to produce over 21,000 photorealistic images with aligned captions. Textual prompts that describe detailed outfit attributes, which are extracted from fashion magazines to ensure authenticity and diversity. To better reflect the inherent ambiguity of design intent, VietFashion adopts a multi-target retrieval setting, where a single query may correspond to multiple valid results. We establish standardized evaluation protocols and benchmark state-of-the-art composed image retrieval methods. Experimental results reveal significant performance gaps in modeling fine-grained cultural semantics and multi-modal composition, positioning VietFashion as a challenging benchmark for fine-grained fashion retrieval. The dataset is publicly available at: https://hng0303.github.io/VietFashion.

图像检索文化服饰多模态草图生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。