构建面向不同读者的医学摘要数据集,提升科普可读性
PERCS: Persona-Guided Controllable Biomedical Summarization Dataset
- 按用户医疗素养分四类人群生成定制化摘要
- 医生审核确保内容准确与人设匹配,可读性差异明显
- 适合医学传播、AI辅助写作研究者使用
自动医学文本简化对提升健康素养至关重要,能让复杂生物医学研究更易被大众理解。然而现有资源多假设单一通用受众,忽视了不同群体间医疗素养和信息需求的显著差异。为此,我们提出PERCS(Persona-guided Controllable Summarization)数据集,包含四种人物角色:普通公众、预医学生、非医学研究者和医学专家的生物医学摘要及对应摘要。每条摘要均由医生评审,确保事实准确性与人物设定一致,并采用详细错误分类体系进行验证。技术验证显示,不同人物角色在可读性、词汇使用和内容深度上存在明显差异。我们还基于PERCS对四个大语言模型进行了基准测试,采用自动评估指标衡量全面性、可读性和忠实度,为后续研究提供基线结果。该数据集、标注指南与评估材料已公开,支持面向特定人物的医学沟通与可控摘要研究。
原文摘要 · Abstract (English)
Automatic medical text simplification plays a key role in improving health literacy by making complex biomedical research accessible to diverse readers. However, most existing resources assume a single generic audience, overlooking the wide variation in medical literacy and information needs across user groups. To address this limitation, we introduce PERCS (Persona-guided Controllable Summarization), a dataset of biomedical abstracts paired with summaries tailored to four personas: Laypersons, Premedical Students, Non-medical Researchers, and Medical Experts. These personas represent different levels of medical literacy and information needs, emphasizing the need for targeted, audience-specific summarization. Each summary in PERCS was reviewed by physicians for factual accuracy and persona alignment using a detailed error taxonomy. Technical validation shows clear differences in readability, vocabulary, and content depth across personas. Along with describing the dataset, we benchmark four large language models on PERCS using automatic evaluation metrics that assess comprehensiveness, readability, and faithfulness, establishing baseline results for future research. The dataset, annotation guidelines, and evaluation materials are publicly available to support research on persona-specific communication and controllable biomedical summarization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。