用集成模型自动构建德语化工领域语义搜索评估数据集
Automated Collection of Evaluation Dataset for Semantic Search in Low-Resource Domain Language
- 通过集成多个弱文本编码器生成查询并重评相关性
- 在人工标注相关性上达到更高一致性和准确率
- 适合资源稀缺领域的语义搜索系统开发者参考
特定领域语言常因专业术语多而属于低资源语言。在狭窄领域内收集测试数据耗时且需具备领域知识的人员参与标注。本研究针对流程工业领域德语这一低资源语言,提出端到端的自动化标注流程,用于生成语义搜索的评估数据。为克服德语化学领域文本编码器不足的问题,采用基于通用知识数据集训练的多个“弱”编码器集成方法。通过融合不同模型生成的文档候选与相关性评分,结合大模型输出的评分,实现查询-文档对的一致性判断。实验表明,该集成方法在人标注的相关性上显著提升,优于单个模型,在评价者间一致性与准确性指标上均表现更优。结果说明,集成学习可有效适配专业、低资源语言的语义搜索系统,为领域专用场景下的资源限制提供可行解决方案。
原文摘要 · Abstract (English)
Domain-specific languages that use a lot of specific terminology often fall into the category of low-resource languages. Collecting test datasets in a narrow domain is time-consuming and requires skilled human resources with domain knowledge and training for the annotation task. This study addresses the challenge of automated collecting test datasets to evaluate semantic search in low-resource domain-specific German language of the process industry. Our approach proposes an end-to-end annotation pipeline for automated query generation to the score reassessment of query-document pairs. To overcome the lack of text encoders trained in the German chemistry domain, we explore a principle of an ensemble of "weak" text encoders trained on common knowledge datasets. We combine individual relevance scores from diverse models to retrieve document candidates and relevance scores generated by an LLM, aiming to achieve consensus on query-document alignment. Evaluation results demonstrate that the ensemble method significantly improves alignment with human-assigned relevance scores, outperforming individual models in both inter-coder agreement and accuracy metrics. These findings suggest that ensemble learning can effectively adapt semantic search systems for specialized, low-resource languages, offering a practical solution to resource limitations in domain-specific contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。