arXiv:2603.19082cs.CL2026-03

首个临床笔记健康素养数据集,助力自动识别患者医疗理解能力

A Dataset and Resources for Identifying Patient Health Literacy Information from Clinical Notes

  • 从真实临床笔记中构建标注数据集,结合人工与大模型主动学习
  • 含589份跨9类笔记的标注数据,支持低/正常/高三级健康素养判断
  • 适用于医疗信息化、临床辅助决策系统开发者

健康素养是影响患者结局的关键因素,但现有筛查工具在项目数量、题型和维度上差异较大,难以结构化录入电子病历。从非结构化临床笔记中自动检测提供了可行替代方案,因这些笔记常包含更丰富、更具上下文的信息。然而进展受限于缺乏标注资源。本文提出HEALIX,首个公开可用的基于真实临床笔记的健康素养标注数据集,通过社工笔记采样、关键词过滤和大模型驱动的主动学习进行构建。该数据集包含589份笔记,覆盖9类笔记类型,标注了低、正常、高三个健康素养等级。为验证其价值,我们在四个开源大语言模型上测试了零样本和少样本提示策略。

原文摘要 · Abstract (English)

Health literacy is a critical determinant of patient outcomes, yet current screening tools are not always feasible and differ considerably in the number of items, question format, and dimensions of health literacy they capture, making documentation in structured electronic health records difficult to achieve. Automated detection from unstructured clinical notes offers a promising alternative, as these notes often contain richer, more contextual health literacy information, but progress has been limited by the lack of annotated resources. We introduce HEALIX, the first publicly available annotated health literacy dataset derived from real clinical notes, curated through a combination of social worker note sampling, keyword-based filtering, and LLM-based active learning. HEALIX contains 589 notes across 9 note types, annotated with three health literacy labels: low, normal, and high. To demonstrate its utility, we benchmarked zero-shot and few-shot prompting strategies across four open source large language models (LLMs).

健康素养临床笔记标注数据集LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。