对比模板与大模型在低资源环境下自动生成认知康复报告的效果。
Automated Clinical Report Generation for Remote Cognitive Remediation: Comparing Knowledge-Engineered Templates and LLMs in Low-Resource Settings

- 用规则模板和零样本大模型生成报告,输入变量一致。
- 模板系统更可靠,大模型输出更简洁流畅。
- 为远程康复报告系统提供可复现的设计指南。
认知康复治疗需求增长,但言语治疗师资源有限,推动了远程康复工具的应用。这些系统产生大量交互数据,临床医生难以高效审查。本文研究在无参考报告的低资源环境下,基于虚拟人引导的家庭认知康复会话的自动化临床报告生成。比较两种方法:(1) 基于规则的模板系统,嵌入言语治疗领域知识的显式决策规则与验证模板,保障临床可靠性与可追溯性;(2) 零样本大模型方法(GPT-4),追求更流畅简洁的输出。两者使用相同预提取、专家验证的结构化变量,实现受控的事实性对比。由八名言语治疗师和高年级学生按九项标准评估。结果表明临床可靠性与语言质量存在明显权衡:模板系统在流畅性、连贯性和结果呈现上得分更高,而GPT-4输出更简洁。各维度差异方向一致,但经校正后均未达统计显著性,反映专家评估规模限制。根据评估反馈,提炼出八条临床报告系统设计建议。本工作贡献了一种结合专家访谈、分类驱动生成与多维人工评估的可复现方法论,适用于低资源环境下的临床自然语言生成,并展示了受控比较如何促进生成式AI在医疗中的负责任应用。
原文摘要 · Abstract (English)
The growing demand for cognitive remediation therapy, combined with limited speech therapist availability, has accelerated the adoption of remote rehabilitation tools. These systems generate large volumes of interaction data that are difficult for clinicians to review efficiently. This paper investigates automated clinical report generation for avatar-guided, home-based cognitive remediation sessions in a low-resource setting with no reference reports. We present and compare two approaches: (1) a rule-based template system encoding speech therapy domain knowledge as explicit decision rules and validated templates, ensuring clinical reliability and traceability; and (2) a zero-shot LLM-based approach (GPT-4) aimed at more fluent and concise output. Both systems use identical pre-extracted, expert-validated structured variables, enabling a controlled factual comparison. Outputs were evaluated by eight speech therapists and final-year students using a nine-criterion questionnaire. Results reveal a clear trade-off between clinical reliability and linguistic quality. The template-based system scored higher on fluidity, coherence, and results presentation, while GPT-4 produced more concise output. Directional differences are consistent across evaluation dimensions, though no comparison reached statistical significance after correction, reflecting the scale constraints of expert clinical evaluation. Based on evaluator feedback, we derive eight design recommendations for clinical reporting systems in remote rehabilitation settings. More broadly, this work contributes a replicable methodology combining expert elicitation, taxonomy-driven generation, and multi-dimensional human evaluation for clinical NLG in low-resource settings, and illustrates how controlled comparisons can inform the responsible adoption of generative AI in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。