arXiv:2507.17717cs.CLcs.AI2025-07EMNLP被引 5

用医生反馈生成可执行的临床笔记评估清单

From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes

  • 从真实医生反馈中提炼结构化评估清单
  • 在2.1万例病历上验证,优于基线方法
  • 适合医疗AI质量评估与合规检查

AI生成的临床笔记在医疗中日益普及,但其质量评估因主观性强、专家评审难扩展而面临挑战。现有自动化指标常与医生实际偏好不符。为此,我们提出一个系统化流程,将真实用户反馈转化为可解释、可执行的结构化检查清单。该清单基于超过21,000例已去标识化临床会话数据(符合HIPAA安全港标准)构建,来自部署的AI医疗记录系统。离线评估显示,该反馈衍生清单在覆盖率、多样性及对人类评分的预测能力方面均优于基线方法。大量实验验证了清单对质量退化扰动的鲁棒性、与医生偏好的高度一致性,以及作为评估方法的实际价值。在离线研究场景中,该清单可有效识别未达标质量的笔记。

原文摘要 · Abstract (English)

AI-generated clinical notes are increasingly used in healthcare, but evaluating their quality remains a challenge due to high subjectivity and limited scalability of expert review. Existing automated metrics often fail to align with real-world physician preferences. To address this, we propose a pipeline that systematically distills real user feedback into structured checklists for note evaluation. These checklists are designed to be interpretable, grounded in human feedback, and enforceable by LLM-based evaluators. Using deidentified data from over 21,000 clinical encounters (prepared in accordance with the HIPAA safe harbor standard) from a deployed AI medical scribe system, we show that our feedback-derived checklist outperforms a baseline approach in our offline evaluations in coverage, diversity, and predictive power for human ratings. Extensive experiments confirm the checklist's robustness to quality-degrading perturbations, significant alignment with clinician preferences, and practical value as an evaluation methodology. In offline research settings, our checklist offers a practical tool for flagging notes that may fall short of our defined quality standards.

医疗AI评估清单临床笔记人机反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。