arXiv:2412.03575cs.IRcs.AI2024-12被引 2

用大模型生成标注数据,让小模型高效精准完成矿点记录链接。

Leveraging Large Language Models for Generating Labeled Mineral Site Record Linkage Data

  • 用大模型自动生成训练数据,解决标注数据稀缺问题。
  • 相比传统方法,F1得分提升45%以上,推理速度加快近18倍。
  • 全流程自动化,适合缺乏专家标注资源的现实场景。

记录链接通过识别指向同一实体的记录来整合异构数据源。在矿点记录场景中,准确的记录链接对识别和绘制矿产资源分布至关重要。正确关联同一矿床的记录有助于明确矿化区域的空间范围,促进资源识别与数据归档。矿点记录链接属于空间记录链接范畴,因记录包含地理位置和非空间属性的表格信息。由于数据异质性强、规模庞大,该任务极具挑战性。以往研究虽使用预训练判别语言模型(PLM)于空间实体链接,但通常需大量人工标注的真值数据进行微调,而获取真值数据耗时费力且成本高昂。尽管大型生成式语言模型(LLMs)在自然语言处理任务中表现优异,包括记录链接,但其高推理开销与资源需求限制了应用。本文提出一种利用LLM生成训练数据并微调PLM的方法,以填补训练数据缺口,同时保持PLM的高效性。实验表明,该方法相较基于真值数据的传统PLM方法,F1得分提升超45%,推理时间减少近18倍。此外,我们提供无需人工干预的自动化流程,凸显该方法在应对记录链接挑战上的潜力。

原文摘要 · Abstract (English)

Record linkage integrates diverse data sources by identifying records that refer to the same entity. In the context of mineral site records, accurate record linkage is crucial for identifying and mapping mineral deposits. Properly linking records that refer to the same mineral deposit helps define the spatial coverage of mineral areas, benefiting resource identification and site data archiving. Mineral site record linkage falls under the spatial record linkage category since the records contain information about the physical locations and non-spatial attributes in a tabular format. The task is particularly challenging due to the heterogeneity and vast scale of the data. While prior research employs pre-trained discriminative language models (PLMs) on spatial entity linkage, they often require substantial amounts of curated ground-truth data for fine-tuning. Gathering and creating ground truth data is both time-consuming and costly. Therefore, such approaches are not always feasible in real-world scenarios where gold-standard data are unavailable. Although large generative language models (LLMs) have shown promising results in various natural language processing tasks, including record linkage, their high inference time and resource demand present challenges. We propose a method that leverages an LLM to generate training data and fine-tune a PLM to address the training data gap while preserving the efficiency of PLMs. Our approach achieves over 45\% improvement in F1 score for record linkage compared to traditional PLM-based methods using ground truth data while reducing the inference time by nearly 18 times compared to relying on LLMs. Additionally, we offer an automated pipeline that eliminates the need for human intervention, highlighting this approach's potential to overcome record linkage challenges.

记录链接大模型生成矿产数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。