arXiv:2509.11191cs.CLcs.IR2025-09中稿 · publication at the…被引 1

提出随机对抗训练,提升生物医学信息抽取的效率与效果。

RanAT4BIE: Random Adversarial Training for Biomedical Information Extraction

  • 结合随机采样与对抗训练,提升模型泛化能力。
  • 在多个生物医学任务上优于基线模型,计算成本更低。
  • 适合追求高效高精度的生物医学NLP研究者使用。

我们提出随机对抗训练(RAT),一种新型框架,成功应用于生物医学信息抽取(BioIE)任务。基于PubMedBERT架构,研究首先验证了传统对抗训练在提升预训练语言模型性能方面的有效性。尽管对抗训练在各项指标上带来显著提升,但其计算开销较大。为此,我们提出RAT作为高效解决方案,将随机采样机制与对抗训练原则有机结合,在增强模型泛化性和鲁棒性的同时,大幅降低计算成本。通过全面评估,RAT在多种BioIE任务中表现优于基线模型,展现出其在生物医学自然语言处理中的变革潜力,为模型性能与计算效率提供了平衡方案。

原文摘要 · Abstract (English)

We introduce random adversarial training (RAT), a novel framework successfully applied to biomedical information extraction (BioIE) tasks. Building on PubMedBERT as the foundational architecture, our study first validates the effectiveness of conventional adversarial training in enhancing pre-trained language models' performance on BioIE tasks. While adversarial training yields significant improvements across various performance metrics, it also introduces considerable computational overhead. To address this limitation, we propose RAT as an efficiency solution for biomedical information extraction. This framework strategically integrates random sampling mechanisms with adversarial training principles, achieving dual objectives: enhanced model generalization and robustness while significantly reducing computational costs. Through comprehensive evaluations, RAT demonstrates superior performance compared to baseline models in BioIE tasks. The results highlight RAT's potential as a transformative framework for biomedical natural language processing, offering a balanced solution to the model performance and computational efficiency.

生物医学对抗训练信息抽取效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。