用ECLASS标准增强语义搜索,提升电子元器件检索准确率
ECLASS-Augmented Semantic Product Search for Electronic Components
- 引入ECLASS层级语义构建产品表征,弥合用户意图与描述的词汇鸿沟
- 在专家查询上达到94.3%的命中率@5,远超BM25的31.4%
- 适合工业自动化、智能代理等需精准检索元器件的场景
高效获取工业产品数据是工厂自动化和基于大模型的智能体工作流的关键,要求工程师和自主代理从高度结构化的目录中识别合适元器件。然而,自然语言查询与属性主导的产品描述之间存在词汇不匹配,导致传统检索方法(如BM25)效果受限。本文系统评估了大模型辅助的密集检索在工业电子元器件语义搜索中的应用,并研究将ECLASS标准的层级语义融入基于嵌入的检索。结果表明,结合重排序的密集检索显著优于传统词法方法和基础模型网络搜索基线。特别地,该方法在专家查询上实现94.3%的Hit_Rate@5,远高于BM25的31.4%,同时在效果与效率上均超越基础模型基线。此外,通过ECLASS语义增强产品表征,在各类配置下均带来一致性能提升,证明标准化层级元数据能有效建立用户意图与稀疏产品描述之间的语义桥梁。
原文摘要 · Abstract (English)
Efficient semantic access to industrial product data is a key enabler for factory automation and emerging LLM-based agent workflows, where both human engineers and autonomous agents must identify suitable components from highly structured catalogs. However, the vocabulary mismatch between natural-language queries and attribute-centric product descriptions limits the effectiveness of traditional retrieval approaches, e.g., BM25. In this work, we present a systematic evaluation of LLM-assisted dense retrieval for semantic product search on industrial electronic components, and investigate the integration of hierarchical semantics from the ECLASS standard into embedding-based retrieval. Our results show that dense retrieval combined with re-ranking substantially outperforms classical lexical methods and foundation model web-search baselines. In particular, the proposed approach achieves a Hit_Rate@5 of 94.3 %, compared to 31.4 % for BM25 on expert queries, while also exceeding foundation model baselines in both effectiveness and efficiency. Furthermore, augmenting product representations with ECLASS semantics yields consistent performance gains across configurations, demonstrating that standardized hierarchical metadata provides a crucial semantic bridge between user intent and sparse product descriptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。