arXiv:2412.16291cs.AIcs.CL2024-12

用小模型替代大模型,实现隐私保护的放疗患者报告摘要。

Benchmarking LLMs and SLMs for patient reported outcomes

  • 对比大模型与小模型在放疗患者问卷摘要中的表现。
  • 小模型在关键指标上接近大模型,但生成稳定性略逊。
  • 适合关注医疗数据隐私和本地部署的临床AI研发者。

大语言模型(LLMs)已广泛应用于医疗任务,尤其在将患者报告结果(PROs)浓缩为简洁自然语言报告方面备受关注,有助于医生聚焦关键问题并开展更有意义的对话。尽管GPT-4等大模型表现优异,但小语言模型(SLMs)因其可本地部署的优势,更有利于保障患者数据隐私并符合医疗监管要求。本研究在放疗场景下,对多种SLMs与LLMs进行基准测试,采用多种评估指标衡量其摘要的准确性和可靠性。结果显示,尽管SLMs在高风险医疗任务中展现出潜力,但仍存在生成一致性不足等局限性,推动更高效且隐私安全的AI医疗应用发展。

原文摘要 · Abstract (English)

LLMs have transformed the execution of numerous tasks, including those in the medical domain. Among these, summarizing patient-reported outcomes (PROs) into concise natural language reports is of particular interest to clinicians, as it enables them to focus on critical patient concerns and spend more time in meaningful discussions. While existing work with LLMs like GPT-4 has shown impressive results, real breakthroughs could arise from leveraging SLMs as they offer the advantage of being deployable locally, ensuring patient data privacy and compliance with healthcare regulations. This study benchmarks several SLMs against LLMs for summarizing patient-reported Q\&A forms in the context of radiotherapy. Using various metrics, we evaluate their precision and reliability. The findings highlight both the promise and limitations of SLMs for high-stakes medical tasks, fostering more efficient and privacy-preserving AI-driven healthcare solutions.

医疗AI小模型隐私保护文本摘要

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。