用复杂描述和配色查询精准检索美甲设计。
NaiLIA: Multimodal Nail Design Retrieval Based on Dense Intent Descriptions and Palette Queries
- 结合详细意图描述与配色选择进行多模态对齐检索。
- 在10625张图像上实现超越现有方法的检索精度。
- 适合需要表达细腻审美意图的个性化设计用户。
本文聚焦于基于密集意图描述的美甲设计图像检索任务,此类描述涵盖用户对涂鸦元素、预制造装饰、视觉特征、主题及整体印象等多层次需求。此外,用户可通过颜色选取器提供零个或多个颜色作为配色查询,以表达细微连续的色彩变化。现有视觉-语言基础模型难以有效融合此类复杂描述与配色信息。为此,我们提出NaiLIA,一种支持密集意图描述与配色查询联合对齐的多模态检索方法。该方法引入基于置信度得分的松弛损失,使未标注图像也能与描述对齐。为评估性能,我们构建了包含10,625张图像的基准数据集,图像来自多元文化背景人群,由200多名标注者提供长篇密集意图描述。实验表明,NaiLIA显著优于标准方法。
原文摘要 · Abstract (English)
We focus on the task of retrieving nail design images based on dense intent descriptions, which represent multi-layered user intent for nail designs. This is challenging because such descriptions specify unconstrained painted elements and pre-manufactured embellishments as well as visual characteristics, themes, and overall impressions. In addition to these descriptions, we assume that users provide palette queries by specifying zero or more colors via a color picker, enabling the expression of subtle and continuous color nuances. Existing vision-language foundation models often struggle to incorporate such descriptions and palettes. To address this, we propose NaiLIA, a multimodal retrieval method for nail design images, which comprehensively aligns with dense intent descriptions and palette queries during retrieval. Our approach introduces a relaxed loss based on confidence scores for unlabeled images that can align with the descriptions. To evaluate NaiLIA, we constructed a benchmark consisting of 10,625 images collected from people with diverse cultural backgrounds. The images were annotated with long and dense intent descriptions given by over 200 annotators. Experimental results demonstrate that NaiLIA outperforms standard methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。