arXiv:2507.14096cs.CLcs.AI2025-07被引 2

用大模型把专业医学摘要转成通俗语言,评估效果并找短板。

Lessons from the TREC Plain Language Adaptation of Biomedical Abstracts (PLABA) track

  • 设计双任务评测:全篇重写与难词替换,结合专家人工评估。
  • 顶尖模型事实准确度接近人类,但简洁性仍不足,自动评分不靠谱。
  • 适合医疗科普、AI可解释性研究者参考,尤其关注真实应用风险。

目的:近年来语言模型在将面向专业人士的生物医学文献转化为通俗语言方面展现出潜力,使患者和照护者更容易理解。然而,这些模型的不可预测性,加上该领域潜在的高危害性,意味着必须进行严格的评估。本赛事的目标是推动相关研究,并提供对最先进系统的高质量评估。方法:我们在2023年和2024年文本检索会议(TREC)上举办了生物医学摘要通俗化(PLABA)赛道。任务包括完整重写摘要(任务1)和识别并替换难词(任务2)。针对任务1的自动评估,我们构建了四组由专业人士撰写的参考文本。任务1和任务2的提交均接受了生物医学专家的详细人工评价。结果:来自12个国家的12支团队参与,模型涵盖多层感知机到大规模预训练变换器。在任务1的人工评估中,表现最优的模型在事实准确性和完整性上达到人类水平,但在简洁性和简短性上仍有差距。基于参考的自动指标与人工判断的相关性较差。在任务2中,系统在识别难词及判断替换方式上表现不佳;但在生成替换词时,基于大语言模型的系统在人工评估中表现出色,准确性、完整性和简洁性均较高,唯独简短性不足。结论:PLABA赛道表明大语言模型在为公众改编生物医学文献方面具有前景,但也暴露了其缺陷,并凸显了改进自动基准工具的必要性。

原文摘要 · Abstract (English)

Objective: Recent advances in language models have shown potential to adapt professional-facing biomedical literature to plain language, making it accessible to patients and caregivers. However, their unpredictability, combined with the high potential for harm in this domain, means rigorous evaluation is necessary. Our goals with this track were to stimulate research and to provide high-quality evaluation of the most promising systems. Methods: We hosted the Plain Language Adaptation of Biomedical Abstracts (PLABA) track at the 2023 and 2024 Text Retrieval Conferences. Tasks included complete, sentence-level, rewriting of abstracts (Task 1) as well as identifying and replacing difficult terms (Task 2). For automatic evaluation of Task 1, we developed a four-fold set of professionally-written references. Submissions for both Tasks 1 and 2 were provided extensive manual evaluation from biomedical experts. Results: Twelve teams spanning twelve countries participated in the track, with models from multilayer perceptrons to large pretrained transformers. In manual judgments of Task 1, top-performing models rivaled human levels of factual accuracy and completeness, but not simplicity or brevity. Automatic, reference-based metrics generally did not correlate well with manual judgments. In Task 2, systems struggled with identifying difficult terms and classifying how to replace them. When generating replacements, however, LLM-based systems did well in manually judged accuracy, completeness, and simplicity, though not in brevity. Conclusion: The PLABA track showed promise for using Large Language Models to adapt biomedical literature for the general public, while also highlighting their deficiencies and the need for improved automatic benchmarking tools.

医学文本大模型通俗化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。