用无监督方法从医学文本中发现实体间关系,减少对标注数据的依赖。
Reduction of Supervision for Biomedical Knowledge Discovery
- 基于依存树和注意力机制设计无监督算法,逐步降低对标注数据的依赖。
- 在生物医学基准数据集上验证,模型在噪声标签下仍保持良好性能。
- 适合资源有限但需快速构建知识图谱的研究场景。
知识发现受限于文献数量激增和高质量标注数据稀缺。为应对信息过载,需采用自动化方法提取与处理知识。如何在监督程度与模型效果间取得平衡是关键挑战:监督方法虽性能更优,但依赖人工标注,耗时且难扩展。本文研究在非结构化文本中识别生物医学实体(如疾病、蛋白质)间的语义关系,同时最大限度减少对标注数据的依赖。提出一套基于依存树和注意力机制的无监督算法,并结合多种点对点二分类方法,在弱监督到全无监督的渐进设置下评估其对噪声标签的鲁棒性。在多个生物医学基准数据集上的评估表明,该方法能有效实现从弱监督向完全无监督的过渡,揭示点对点分类技术在低标注场景下的潜力。综合对比提供了对这些技术有效性的深入理解,为构建可适应、数据高效的发现系统指明了方向。
原文摘要 · Abstract (English)
Knowledge discovery is hindered by the increasing volume of publications and the scarcity of extensive annotated data. To tackle the challenge of information overload, it is essential to employ automated methods for knowledge extraction and processing. Finding the right balance between the level of supervision and the effectiveness of models poses a significant challenge. While supervised techniques generally result in better performance, they have the major drawback of demanding labeled data. This requirement is labor-intensive and time-consuming and hinders scalability when exploring new domains. In this context, our study addresses the challenge of identifying semantic relationships between biomedical entities (e.g., diseases, proteins) in unstructured text while minimizing dependency on supervision. We introduce a suite of unsupervised algorithms based on dependency trees and attention mechanisms and employ a range of pointwise binary classification methods. Transitioning from weakly supervised to fully unsupervised settings, we assess the methods' ability to learn from data with noisy labels. The evaluation on biomedical benchmark datasets explores the effectiveness of the methods. Our approach tackles a central issue in knowledge discovery: balancing performance with minimal supervision. By gradually decreasing supervision, we assess the robustness of pointwise binary classification techniques in handling noisy labels, revealing their capability to shift from weakly supervised to entirely unsupervised scenarios. Comprehensive benchmarking offers insights into the effectiveness of these techniques, suggesting an encouraging direction toward adaptable knowledge discovery systems, representing progress in creating data-efficient methodologies for extracting useful insights when annotated data is limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。