混合抽样方法提升科学领域信息抽取效果
STAYKATE: Hybrid In-Context Example Selection Combining Representativeness Sampling and Retrieval-based Approach -- A Case Study on Science Domains
- 结合代表性采样与检索策略,动态优化上下文示例选择
- 在三个科学数据集上表现优于传统方法,尤其对难识别实体类型提升明显
- 适合需要高效标注的科研场景,尤其在数据稀缺时
大型语言模型具备上下文学习能力,为科学信息抽取提供了新路径,尤其适用于训练数据不足且标注成本高的场景。由于上下文示例的选择直接影响性能,设计高效的样本选取方法至关重要。本文提出STAYKATE,一种静态-动态混合选择方法,融合主动学习中的代表性采样与主流检索式方法的优势。在三个领域特定数据集上的实验表明,STAYKATE显著优于传统监督方法及现有选择策略,尤其在其他方法难以处理的实体类型上表现突出。
原文摘要 · Abstract (English)
Large language models (LLMs) demonstrate the ability to learn in-context, offering a potential solution for scientific information extraction, which often contends with challenges such as insufficient training data and the high cost of annotation processes. Given that the selection of in-context examples can significantly impact performance, it is crucial to design a proper method to sample the efficient ones. In this paper, we propose STAYKATE, a static-dynamic hybrid selection method that combines the principles of representativeness sampling from active learning with the prevalent retrieval-based approach. The results across three domain-specific datasets indicate that STAYKATE outperforms both the traditional supervised methods and existing selection methods. The enhancement in performance is particularly pronounced for entity types that other methods pose challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。