评估大模型在医学综述中对年龄差异的识别能力,发现存在显著偏见。
LLMs Do Not See Age: Assessing Demographic Bias in Automated Systematic Review Synthesis

- 构建年龄分层数据集DemogSummary,评估大模型摘要中的年龄信息保留度
- 成人相关摘要的年龄信息保留最差,少数群体更易出现幻觉
- 提出新评价指标DSS,推动公平性评估在生物医学NLP中的应用
临床干预常与年龄密切相关:某些药物或治疗对成人安全,却可能对儿童有害或对老年人无效。随着语言模型越来越多地融入生物医学证据合成流程,这些系统是否能保留关键的人口学差异仍不明确。为此,我们评估了先进语言模型在生成生物医学研究摘要时对年龄相关信息的保留能力。我们构建了新的年龄分层数据集DemogSummary,涵盖儿童、成人和老年群体。评估了三种主流摘要型大模型:Qwen(开源)、Longformer(开源)和GPT-4.1 Nano(专有)。使用标准指标和新提出的“人口学显著性评分”(DSS),量化年龄实体保留程度与幻觉情况。结果揭示模型间及不同年龄群体间存在系统性偏差:以成人为主的摘要中人口学保真度最低,被低估群体更易产生幻觉。研究凸显当前大模型在忠实且无偏摘要方面的局限,呼吁建立注重公平性的评估框架与摘要流程。
原文摘要 · Abstract (English)
Clinical interventions often hinge on age: medications and procedures safe for adults may be harmful to children or ineffective for older adults. However, as language models are increasingly integrated into biomedical evidence synthesis workflows, it remains uncertain whether these systems preserve such crucial demographic distinctions. To address this gap, we evaluate how well state-of-the-art language models retain age-related information when generating abstractive summaries of biomedical studies. We construct DemogSummary, a novel age-stratified dataset of systematic review primary studies, covering child, adult, and older adult populations. We evaluate three prominent summarisation-capable LLMs, Qwen (open-source), Longformer (open-source) and GPT-4.1 Nano (proprietary), using both standard metrics and a newly proposed Demographic Salience Score (DSS), which quantifies age-related entity retention and hallucination. Our results reveal systematic disparities across models and age groups: demographic fidelity is lowest for adult-focused summaries, and under-represented populations are more prone to hallucinations. These findings highlight the limitations of current LLMs in faithful and bias-free summarisation and point to the need for fairness-aware evaluation frameworks and summarisation pipelines in biomedical NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。