arXiv:2511.08600cs.CLcs.HC2025-11

用检索增强生成技术,自动生成符合专业标准的儿童语言病理学案例。

Retrieval-Augmented Generation of Pediatric Speech-Language Pathology vignettes: A Proof-of-Concept Study

  • 结合领域知识库与提示工程,用RAG提升大模型生成准确性。
  • 开源模型表现达标,可支持机构隐私保护部署。
  • 适合教育、临床训练及辅助决策系统开发人员使用。

临床案例是言语语言病理学(SLP)教育的核心工具,但人工编写耗时费力。通用大语言模型虽能生成文本,却缺乏专业领域知识,常产生幻觉且需专家大量修正。本研究提出一个概念验证系统,将检索增强生成(RAG)与精心整理的知识库结合,生成儿童SLP病例材料。系统集成多个商用(GPT-4o、Claude 3.5 Sonnet、Gemini 2.5 Pro)和开源模型(Llama 3.2、Qwen 2.5-7B),设计七种涵盖不同障碍类型与年级水平的测试场景。通过多维度评分体系评估生成案例的结构完整性、内部一致性、临床适宜性及个别化教育计划(IEP)目标与会话记录质量。结果表明,该技术在概念上可行;商用模型略优,但开源模型表现已达到可用水平,具备隐私保护部署潜力。知识库集成使生成内容符合专业指南。未来需经专家评审、学生试用与心理测量验证后方可用于教学或研究。潜在应用包括临床决策支持、自动化IEP目标生成与临床反思训练。

原文摘要 · Abstract (English)

Clinical vignettes are essential educational tools in speech-language pathology (SLP), but manual creation is time-intensive. While general-purpose large language models (LLMs) can generate text, they lack domain-specific knowledge, leading to hallucinations and requiring extensive expert revision. This study presents a proof-of-concept system integrating retrieval-augmented generation (RAG) with curated knowledge bases to generate pediatric SLP case materials. A multi-model RAG-based system was prototyped integrating curated domain knowledge with engineered prompt templates, supporting five commercial (GPT-4o, Claude 3.5 Sonnet, Gemini 2.5 Pro) and open-source (Llama 3.2, Qwen 2.5-7B) LLMs. Seven test scenarios spanning diverse disorder types and grade levels were systematically designed. Generated cases underwent automated quality assessment using a multi-dimensional rubric evaluating structural completeness, internal consistency, clinical appropriateness, and IEP goal/session note quality. This proof-of-concept demonstrates technical feasibility for RAG-augmented generation of pediatric SLP vignettes. Commercial models showed marginal quality advantages, but open-source alternatives achieved acceptable performance, suggesting potential for privacy-preserving institutional deployment. Integration of curated knowledge bases enabled content generation aligned with professional guidelines. Extensive validation through expert review, student pilot testing, and psychometric evaluation is required before educational or research implementation. Future applications may extend to clinical decision support, automated IEP goal generation, and clinical reflection training.

生成模型医疗AI教育工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。