arXiv:2607.13430cs.CL2026-07

小模型经对齐后生成药盒说明书更准,效果超大模型。

Exploring Post-Training Alignment of Small Language Models for Biomedical Data-to-Text Generation: A Case Study of Medication Leaflet

  • 用Qwen小模型+多种对齐方法生成医学文本
  • ORPO比微调更好,GRPO跨数据集最稳
  • 适合医疗文本生成、资源有限的研究者

将复杂生物医学数据转化为患者友好的叙述是现代生物医学信息学的核心。本研究对比了在药物说明书数据集上,基于Qwen的小语言模型(SLMs)使用监督微调(SFT)、直接偏好优化(DPO)、几率比偏好优化(ORPO)和组相对策略优化(GRPO)等后训练方法的表现。为评估跨数据集泛化能力,我们还收集了来自openFDA的药品标签数据。采用标准词汇重叠指标(如ROUGE)和语义相似性度量进行评估。实验结果表明:(1) 对齐后的小模型性能优于专有模型GPT-5;(2) ORPO优于SFT基线;(3) GRPO在所有测试对齐方法中表现出最强的跨数据集鲁棒性,甚至超过GPT-5。

原文摘要 · Abstract (English)

Translating complex biomedical data into patient-friendly narratives is central to modern biomedical informatics. This study presents a comparative analysis of training small language models (SLMs) in specialized biomedical datato-text generation tasks. We explore widely adopted post-training methods including supervised fine-tuning (SFT), direct preference optimization (DPO), odds ratio preference optimization (ORPO), and group relative policy optimization (GRPO) with Qwen-based SLMs on a medicine package leaflets dataset. To assess cross-dataset generalizability, we also curated drug label data from openFDA. We evaluate models using both standard lexical overlap metrics like ROUGE as well as semantic similarity measures. Across our experiments, the results show that (1) the aligned SLMs outperform proprietary models like GPT-5; (2) ORPO outperforms the SFTbaselines; (3) GRPO yields the most robust cross-dataset performance among the alignment methods tested as well as GPT-5.

小模型医疗生成对齐方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。