arXiv:2608.29481cs.CL2026-08中稿 · EMNLP

首个专为儿科重症沟通训练设计的模拟框架与评估基准

SIC-Agents: Benchmarking and Building an Adaptive Simulator for Pediatric Serious Illness Communication Training

论文配图:SIC-Agents: Benchmarking and Building an Adaptive Simulator for Pediatric Serious Illness Communication Training
图 1 · 摘自论文原文
  • 构建自适应模拟器,通过可编辑技能文档动态调整家长角色行为
  • 提出两个评测基准,分别评估对话轮次与全程表现,提升训练真实性
  • 适合医学教育研究者及临床沟通培训系统开发者使用

儿科重症沟通(SIC)至关重要,但面向临床医生的可扩展沟通培训仍有限。相较于其他对话模拟场景,儿科SIC面临多方互动、应对家长情绪及强反馈依赖等挑战。现有基于大模型的模拟器侧重通用对话质量,而非满足有效SIC培训所需的课程相关行为。本文联合教育专家与儿科临床医生,首次推出针对儿科SIC训练的基准套件与仿真框架。所提出的PitfallBench和DialogueBench基准,分别从回合级与全程对话角度评估模拟器性能。进一步提出SIC-Agents框架,具备自我优化能力,生成临床医生可编辑的技能文档以引导模拟行为。实验表明,SIC-Agents优于静态专家提示。为支持后续研究,项目已开源儿科SIC中家长角色模拟的基准数据集,地址见https://github.com/Beikewzh/sic-benchmarks。

原文摘要 · Abstract (English)

Pediatric serious illness communication (SIC) is critically important, yet scalable communication training for clinicians remains limited. Compared with other dialogue simulation settings, pediatric SIC poses additional challenges, including multi-party interactions, response to parental distress and strong dependence on feedback dynamics. Existing LLM-based simulators optimize generic dialogue quality rather than curriculum-contingent behavior required for effective SIC training. In collaboration with educators and pediatric clinicians, we introduce the first benchmark suite and simulation framework tailored to pediatric SIC training. Our benchmarks, PitfallBench and DialogueBench, evaluate simulators both at the turn-level and across full dialogues. We further propose SIC-Agents, a self-improving framework that generates a clinician-editable skill document to guide simulator behavior. Our experiments show that SIC-Agents outperforms static expert prompting. To support future research, we release our benchmarks for parent simulation in pediatric SIC at https://github.com/Beikewzh/sic-benchmarks

医疗AI对话模拟临床培训LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。