arXiv:2605.11533cs.CLcs.CV2026-05被引 1

将体检报告转化为患者可执行的行动卡片,助力精准医疗沟通。

Checkup2Action: A Multimodal Clinical Check-up Report Dataset for Patient-Oriented Action Card Generation

论文配图:Checkup2Action: A Multimodal Clinical Check-up Report Dataset for Patient-Oriented Action Card Generation
图 1 · 摘自论文原文
  • 构建多模态体检报告数据集,支持跨页跨模态信息整合。
  • 2000份真实报告生成行动卡,84%被专家评为合理有效。
  • 强调安全约束下平衡覆盖度与行动指导,适合医疗AI研究者。

常规体检报告包含检验数据、生理评估、影像结果和结构化信息,但很少告诉患者下一步该做什么。将报告转化为后续行动需模型跨页面、跨表格、跨模态关联证据,识别临床关键问题,并以无过度诊断或治疗承诺的方式传达建议。然而,这种报告到行动的能力尚未得到充分评估。我们提出C2A数据集与基准,用于从多模态体检报告生成结构化‘行动卡’,并设计了受约束的Checkup2Action工作流。C2A包含2000份去标识的真实世界报告,涵盖体格检查、实验室检测、心血管评估和影像证据。每张行动卡明确标注一个问题、优先级、推荐科室、随访时间窗、面向患者的解释及供医生提问的问题。评估涵盖问题覆盖率与精确率、优先级一致性、科室与时间准确性、行动复杂度、实用性、可读性与安全性。实验显示,通用与医学大模型在各维度表现各异。临床专家判断84%输出完全合理且有效。移除安全约束后问题召回率从0.527升至0.825,但出现更多诊断过度陈述。C2A因此支持系统性研究覆盖度、可操作指引与安全患者沟通之间的权衡。

原文摘要 · Abstract (English)

Routine clinical check-up reports combine laboratory measurements, physiological assessments, imaging findings and visually structured information, but rarely tell patients what to do next. Translating them into follow-up actions requires models to connect evidence across pages, tables and modalities, identify clinically relevant issues and communicate next steps without unsupported diagnostic or treatment claims. Yet this report-to-action capability remains poorly benchmarked. We introduce C2A, a dataset and benchmark for generating structured \textit{Action Cards} from multimodal check-up reports, together with Checkup2Action, a constrained workflow for the task. C2A contains 2,000 de-identified real-world reports covering physical examinations, laboratory tests, cardiovascular assessments and imaging evidence. Each card specifies one issue, its priority, recommended department, follow-up window, patient-facing explanation and questions for clinicians. Our evaluation measures issue coverage and precision, priority consistency, department and timing accuracy, action complexity, usefulness, readability and safety. Experiments across general-purpose and medical large language models show that no model performs best on every dimension. Clinical experts judged 84% of evaluated outputs fully reasonable and effective. Removing safety constraints increased problem recall from 0.527 to 0.825, but produced more diagnostic overstatement. C2A therefore supports systematic study of the balance between coverage, actionable guidance and safe patient communication.

医疗AI多模态行动卡数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。