arXiv:2608.29995cs.CLcs.AI2026-08中稿 · Findings of the As…

用认知模型生成更真实临床案例,确保结构和因果关系准确。

Generating Clinical Vignettes that Preserve Cognitive Formulations

论文配图:Generating Clinical Vignettes that Preserve Cognitive Formulations
图 1 · 摘自论文原文
  • 将认知模型转为加权图,按个体配置生成案例
  • 完整条件生成的案例可恢复认知图(MCC=0.41)
  • 生成案例被专家评为更真实,85%像真人撰写

大语言模型能生成流畅的临床案例,但流畅不等于结构准确。本文提出FORMA框架,将创伤后应激障碍的埃勒斯-克拉克认知模型转化为有向加权图,采样500个个性化配置,结合11种生成模型与3种消融条件,生成16,500个案例。评估包含外部边恢复探测、两位临床专家、一个扩大的LLM评判器及100名持证临床医师用户研究。完整条件生成的案例可恢复认知图(MCC=+0.41,AUC=0.70),而零样本生成无法恢复(MCC=+0.01,AUC=0.50)。专家对完整案例评分更高,临床医生认为其85%像真人撰写(零样本仅22%)。该方法还使感知质量的性别/种族差异缩小1.5至7倍。结果表明认知构型可作为可审计的合成临床文本生成规范。代码与数据仓库已公开:https://github.com/Amit-Oren/FORMA。

原文摘要 · Abstract (English)

Large language models can generate fluent clinical case vignettes, but fluency alone does not ensure fidelity to a specifiable clinical structure. We introduce FORMA, a theory-grounded framework that compiles a cognitive model of a disorder into a directed weighted graph, samples a person-specific configuration of that graph, and validates whether the generated vignette preserves the specified components and causal links. We instantiate FORMA on Posttraumatic Stress Disorder using the Ehlers and Clark cognitive model, generating 16,500 vignettes across 500 personas, 11 generation models, and three ablation conditions. Evaluation combines an external edge-recovery probe, two clinical experts, a scaled LLM judge, and a clinician user study with 100 licensed practitioners. The cognitive graph is recoverable from full-condition vignettes (MCC = +0.41, AUC = 0.70) but not from zero-shot generation (MCC = +0.01, AUC = 0.50). Experts rate full vignettes substantially higher than zero-shot alternatives, and clinicians perceive them to be human-written 85% of the time, compared with 22% for zero-shot. FORMA also reduces demographic disparity in perceived quality by 1.5-7x. These results show that cognitive formulation can serve as an auditable specification for scalable synthetic clinical text generation. A repository with the data and code is available online: https://github.com/Amit-Oren/FORMA.

临床生成认知模型可审计性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。