arXiv:2409.15804cs.CL2024-09

为奢侈品行业打造专用命名实体识别模型,解决术语混乱与数据稀缺问题。

NER-Luxury: Named entity recognition for the fashion and luxury domain

  • 构建36类奢侈品实体分类体系,设计专属标注规范
  • 创建超4万句带层级结构的标注语料库
  • 专有模型性能优于通用大模型,适合产业级应用

本文针对英语奢侈品行业命名实体识别面临的多重挑战展开研究,包括实体歧义、多个子领域中的法语专业术语、ESG方法论数据匮乏,以及从小型奢侈品牌到大型集团的公司结构差异。为此,我们提出包含36种以上实体类型的分类体系,采用面向奢侈品领域的标注方案,并构建了一个超过4万句、具备清晰层级分类的语料库。同时,我们训练了五个针对时装、美妆、腕表、珠宝、香氛、化妆品及整体奢侈品的监督微调模型(NER-Luxury),兼顾美学特征与量化表现。在额外实验中,通过定量实证评估,我们的模型在命名实体识别性能上优于当前主流开源大语言模型,凸显了在现有机器学习流程中引入定制化NER模型的优势。

原文摘要 · Abstract (English)

In this study, we address multiple challenges of developing a named-entity recognition model in English for the fashion and luxury industry, namely the entity disambiguation, French technical jargon in multiple sub-sectors, scarcity of the ESG methodology, and a disparate company structures of the sector with small and medium-sized luxury houses to large conglomerate leveraging economy of scale. In this work, we introduce a taxonomy of 36+ entity types with a luxury-oriented annotation scheme, and create a dataset of more than 40K sentences respecting a clear hierarchical classification. We also present five supervised fine-tuned models NER-Luxury for fashion, beauty, watches, jewelry, fragrances, cosmetics, and overall luxury, focusing equally on the aesthetic side and the quantitative side. In an additional experiment, we compare in a quantitative empirical assessment of the NER performance of our models against the state-of-the-art open-source large language models that show promising results and highlights the benefits of incorporating a bespoke NER model in existing machine learning pipelines.

命名实体识别奢侈品文本标注领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。