通过语义扰动评估医学文献分类模型的鲁棒性,提升跨数据分布的准确率。
Robust Biomedical Publication Type and Study Design Classification with Knowledge-Guided Perturbations

- 引入可控语义扰动评估分类器在分布外场景下的表现
- 结合实体掩码与领域对抗训练,使模型更依赖方法学特征
- 适合关注模型可解释性与泛化能力的研究者
准确且一致地对生物医学文献进行发表类型和研究设计分类,对支持循证综合与知识发现至关重要。以往自动化分类研究多聚焦于扩展标签覆盖、丰富特征表示和提升域内准确性,评估通常基于与训练数据同分布的数据。尽管预训练生物医学语言模型在此类设置下表现优异,但为追求域内准确率而优化的模型可能依赖表面词汇或数据集特定线索,导致在分布外场景下鲁棒性下降。本研究提出基于受控语义扰动的评估框架,用于检验发表类型分类器的鲁棒性,并探索结合实体掩码与领域对抗训练的鲁棒性优化策略,以减少对虚假主题相关性的依赖。结果表明,当鲁棒性目标被设计为选择性抑制非任务定义特征的同时保留关键方法学信号时,鲁棒性与域内准确率之间的常见权衡可被缓解。改进源于两个互补机制:(1) 当输入中存在明确方法学线索时,模型更依赖这些线索;(2) 减少对领域特定主题特征的依赖。这些发现凸显了特征层面鲁棒性分析的重要性,提示进一步细化掩码与对抗目标以更精准抑制主题信息,或可进一步提升鲁棒性。数据、代码与模型已公开于:https://github.com/ScienceNLP-Lab/MultiTagger-v2/tree/main/ICHI
原文摘要 · Abstract (English)
Accurately and consistently indexing biomedical literature by publication type and study design is essential for supporting evidence synthesis and knowledge discovery. Prior work on automated publication type and study design indexing has primarily focused on expanding label coverage, enriching feature representations, and improving in-domain accuracy, with evaluation typically conducted on data drawn from the same distribution as training. Although pretrained biomedical language models achieve strong performance under these settings, models optimized for in-domain accuracy may rely on superficial lexical or dataset-specific cues, resulting in reduced robustness under distributional shift. In this study, we introduce an evaluation framework based on controlled semantic perturbations to assess the robustness of a publication type classifier and investigate robustness-oriented training strategies that combine entity masking and domain-adversarial training to mitigate reliance on spurious topical correlations. Our results show that the commonly observed trade-off between robustness and in-domain accuracy can be mitigated when robustness objectives are designed to selectively suppress non-task-defining features while preserving salient methodological signals. We find that these improvements arise from two complementary mechanisms: (1) increased reliance on explicit methodological cues when such cues are present in the input, and (2) reduced reliance on spurious domain-specific topical features. These findings highlight the importance of feature-level robustness analysis for publication type and study design classification and suggest that refining masking and adversarial objectives to more selectively suppress topical information may further improve robustness. Data, code, and models are available at: https://github.com/ScienceNLP-Lab/MultiTagger-v2/tree/main/ICHI
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。