面向生物医学领域的数字图书馆,用大模型降低文本标注成本。
A Library Perspective on Supervised Text Processing in Digital Libraries: An Investigation in the Biomedical Domain
- 用大模型和远监督生成训练数据,减少人工标注
- 在8个生物医学基准上验证了方法的实用性
- 兼顾准确率与部署成本,适合实际应用者参考
维护大量文本资源的数字图书馆希望为下游任务(如构建知识图谱、文档语义增强或新访问路径)进一步丰富内容。这些任务需要文本处理,例如识别实体、提取语义关系或对文档分类。然而,对于数字图书馆而言,实现可靠的监督工作流颇具挑战,因需手工制作训练数据并训练可靠模型。现有研究多关注基准上的最高准确率,而本文从数字图书馆实践者视角出发,同时考虑准确率与应用成本之间的权衡。我们探索通过远监督及大语言模型(如ChatGPT、LLama、Olmo)生成训练数据,并讨论最终流水线的设计。研究聚焦于关系抽取与文本分类,在8个生物医学基准上进行验证。
原文摘要 · Abstract (English)
Digital libraries that maintain extensive textual collections may want to further enrich their content for certain downstream applications, e.g., building knowledge graphs, semantic enrichment of documents, or implementing novel access paths. All of these applications require some text processing, either to identify relevant entities, extract semantic relationships between them, or to classify documents into some categories. However, implementing reliable, supervised workflows can become quite challenging for a digital library because suitable training data must be crafted, and reliable models must be trained. While many works focus on achieving the highest accuracy on some benchmarks, we tackle the problem from a digital library practitioner. In other words, we also consider trade-offs between accuracy and application costs, dive into training data generation through distant supervision and large language models such as ChatGPT, LLama, and Olmo, and discuss how to design final pipelines. Therefore, we focus on relation extraction and text classification, using the showcase of eight biomedical benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。