用大模型生成个性化科普摘要,提升理解但可能引入偏见和幻觉。
ReLay: Personalized LLM-Generated Plain-Language Summaries for Better Understanding, but at What Cost?

- 基于用户需求动态生成个性化科普摘要
- 个性化使理解度和满意度提升,但错误率也上升
- 适合关注可解释性与安全性的医疗信息传播研究者
通俗语言摘要(PLS)旨在让非专业读者理解科研内容,但通常采用统一风格,忽略读者差异。在健康领域,理解偏差可能影响实际决策。大语言模型(LLMs)为个性化PLS带来新可能,但其有效性、最佳策略及安全边界尚不明确。我们提出ReLay数据集,包含50名普通参与者在静态(专家撰写)与交互式(LLM个性化)场景下的300组参与者-摘要配对,涵盖用户特征、健康信息需求、搜索行为、理解结果、交互日志与质量评分。利用该数据集,评估了五种LLM在两种个性化方法下的表现。结果显示,个性化显著提升理解度与感知质量,但也加剧了用户偏见强化与幻觉风险,揭示了个性化与安全之间的权衡。研究强调需开发既有效又可信的个性化方法,以服务多元化的普通受众。
原文摘要 · Abstract (English)
Plain Language Summaries (PLS) aim to make research accessible to lay readers, but they are typically written in a one-size-fits-all style that ignores differences in readers' information needs and comprehension. In health contexts, this limitation is particularly important because misunderstanding scientific information can affect real-world decisions. Large language models (LLMs) offer new opportunities for personalizing PLS, but it remains unclear whether personalization helps, which strategies are most effective, and how to balance personalization with safety. We introduce ReLay, a dataset of 300 participant--PLS pairs from 50 lay participants in both static (expert-written) and interactive (LLM-personalized) settings. ReLay includes user characteristics, health information needs, information-seeking behavior, comprehension outcomes, interaction logs, and quality ratings. We use ReLay to evaluate five LLMs across two personalization methods. Personalization improves comprehension and perceived quality, but it also raises the risk of reinforcing user biases and introducing hallucinations, revealing a trade-off between personalization and safety. These findings highlight the need for personalization methods that are both effective and trustworthy for diverse lay audiences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。