对比三款兽医专用大模型,发现性能差异显著且评估方法可靠。
Context Matters: Comparison of commercial large language tools in veterinary medicine
- 用评分框架对比三款兽医大模型摘要能力,聚焦事实准确等五维度。
- 产品1得分最高(中位数4.61),在事实准确和时间顺序上表现完美。
- 评估结果高度可复现,证明该方法适合规模化评测兽医NLP系统。
大型语言模型(LLMs)在临床中的应用日益广泛,但在兽医医学领域的表现仍缺乏研究。本研究在标准化的兽医肿瘤学病历数据集上,评估了三款商用兽医专注型大模型摘要工具(产品1 [Hachiko] 及产品2、3)的表现。采用基于评分标准的LLM-as-a-judge框架,从事实准确性、完整性、时间顺序、临床相关性和组织性五个维度进行评分。产品1整体表现最佳,中位平均分达4.61(四分位距: 0.73),显著高于产品2(2.55,IQR: 0.78)和产品3(2.45,IQR: 0.92)。其在事实准确性和时间顺序上均获得满分中位数。为验证评分框架内部一致性,重复三次独立评估,结果显示评分器具有高可复现性:各产品平均得分标准差分别为0.015(产品1)、0.088(产品2)和0.034(产品3)。研究凸显了兽医专用商业大模型的重要性,并证实'大模型作为裁判'的评估方式在兽医临床NLP摘要任务中具备可扩展性和可重复性。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in clinical settings, yet their performance in veterinary medicine remains underexplored. We evaluated three commercially available veterinary-focused LLM summarization tools (Product 1 [Hachiko] and Products 2 and 3) on a standardized dataset of veterinary oncology records. Using a rubric-guided LLM-as-a-judge framework, summaries were scored across five domains: Factual Accuracy, Completeness, Chronological Order, Clinical Relevance, and Organization. Product 1 achieved the highest overall performance, with a median average score of 4.61 (IQR: 0.73), compared to 2.55 (IQR: 0.78) for Product 2 and 2.45 (IQR: 0.92) for Product 3. It also received perfect median scores in Factual Accuracy and Chronological Order. To assess the internal consistency of the grading framework itself, we repeated the evaluation across three independent runs. The LLM grader demonstrated high reproducibility, with Average Score standard deviations of 0.015 (Product 1), 0.088 (Product 2), and 0.034 (Product 3). These findings highlight the importance of veterinary-specific commercial LLM tools and demonstrate that LLM-as-a-judge evaluation is a scalable and reproducible method for assessing clinical NLP summarization in veterinary medicine.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。