让医生把关AI生成的慢病随访报告,既省时又保信任。
Human-in-the-Loop Interactive Report Generation for Chronic Disease Adherence
- AI负责组织内容并生成初稿,医生通过图表核对逐条确认。
- 医生仅需修改8.3%内容,报告质量接近人工水平(4.86/5分)。
- 适合需要高可信度的临床场景,如慢病管理与医患沟通。
慢性病管理需定期反馈依从性以避免可预防的住院,但临床医生缺乏时间制作个性化患者沟通材料。手动撰写虽准确但难以扩展;AI生成虽可扩展却可能削弱患者信任。我们提出一种医生在环的交互界面,通过基于识别的审查机制约束AI,保持医生全程监督。单页编辑器将AI生成的内容段落与时间对齐的可视化图表配对,支持带证据的逐条修改。该分工模式(AI组织,医生决策)兼顾效率与责任。在三位医生审核24个病例的试点中,AI生成的个性化初稿与医生手工写作表现相当(平均4.86/10对比5.0/10),平均仅需8.3%内容修改,无安全关键问题,证明内容生成环节可有效自动化。然而,审查时间仍与手动操作相当,揭示出问责悖论:在高风险临床场景中,专业责任要求完全验证,无论AI多准确。我们贡献三种临床AI协作模式:基于图表-文本配对的受限生成与识别式审查、自动紧急标记(分析生命体征趋势与依从性模式,关键任务遗漏时触发安全升级)、渐进式披露控制(降低认知负荷同时维持监督)。这些模式表明,临床AI效率不仅依赖模型精度,还需支持选择性验证的机制以保障问责。
原文摘要 · Abstract (English)
Chronic disease management requires regular adherence feedback to prevent avoidable hospitalizations, yet clinicians lack time to produce personalized patient communications. Manual authoring preserves clinical accuracy but does not scale; AI generation scales but can undermine trust in patient-facing contexts. We present a clinician-in-the-loop interface that constrains AI to data organization and preserves physician oversight through recognition-based review. A single-page editor pairs AI-generated section drafts with time-aligned visualizations, enabling inline editing with visual evidence for each claim. This division of labor (AI organizes, clinician decides) targets both efficiency and accountability. In a pilot with three physicians reviewing 24 cases, AI successfully generated clinically personalized drafts matching physicians' manual authoring practice (overall mean 4.86/10 vs. 5.0/10 baseline), requiring minimal physician editing (mean 8.3\% content modification) with zero safety-critical issues, demonstrating effective automation of content generation. However, review time remained comparable to manual practice, revealing an accountability paradox: in high-stakes clinical contexts, professional responsibility requires complete verification regardless of AI accuracy. We contribute three interaction patterns for clinical AI collaboration: bounded generation with recognition-based review via chart-text pairing, automated urgency flagging that analyzes vital trends and adherence patterns with fail-safe escalation for missed critical monitoring tasks, and progressive disclosure controls that reduce cognitive load while maintaining oversight. These patterns indicate that clinical AI efficiency requires not only accurate models, but also mechanisms for selective verification that preserve accountability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。