用大模型生成急救分诊案例,解决真实数据少难题。
Syn-STARTS: Synthesized START Triage Scenario Generation Framework for Scalable LLM Evaluation
- 用大模型自动构造急救分诊场景,替代人工收集。
- 生成案例与真实数据质量相当,分类准确率稳定。
- 适合医疗AI评估、灾难应急研究者使用。
在大规模伤亡事件(MCIs)中,分诊是决定幸存率的关键决策过程。尽管人工智能在资源与时间受限下做出最优决策的潜力受到关注,但其开发与评估仍需大量高质量基准数据。然而,由于真实事件罕见,现场数据难以积累,难以获取大规模真实数据用于研究。为此,我们提出Syn-STARTS框架,利用大语言模型生成分诊案例,并验证其有效性。结果表明,Syn-STARTS生成的案例在定性上与由培训材料手动构建的TRIAGE公开数据集无差别。此外,在使用每类(绿、黄、红、黑)数百个案例评估大模型准确性时,结果高度稳定,充分证明合成数据在构建严重医疗情境下高性能AI模型中的可行性。
原文摘要 · Abstract (English)
Triage is a critically important decision-making process in mass casualty incidents (MCIs) to maximize victim survival rates. While the role of AI in such situations is gaining attention for making optimal decisions within limited resources and time, its development and performance evaluation require benchmark datasets of sufficient quantity and quality. However, MCIs occur infrequently, and sufficient records are difficult to accumulate at the scene, making it challenging to collect large-scale realworld data for research use. Therefore, we developed Syn-STARTS, a framework that uses LLMs to generate triage cases, and verified its effectiveness. The results showed that the triage cases generated by Syn-STARTS were qualitatively indistinguishable from the TRIAGE open dataset generated by manual curation from training materials. Furthermore, when evaluating the LLM accuracy using hundreds of cases each from the green, yellow, red, and black categories defined by the standard triage method START, the results were found to be highly stable. This strongly indicates the possibility of synthetic data in developing high-performance AI models for severe and critical medical situations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。