构建首个韩语兽医临床推理评测基准,评估大模型在真实问诊场景中的表现。
PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning

- 基于真实宠物问诊数据构建多模态问答集,覆盖犬猫常见病
- 18个模型在零样本、RAG、微调三种设置下均表现不足,尤其在推理环节
- 提供五种语言翻译版本,适合兽医AI研究者与临床辅助系统开发者
我们提出PetQA,一个用于评估大型语言模型(LLMs)和大型视觉-语言模型(LVLMs)在兽医知识与临床推理能力上的韩语长文本问答基准。PetQA包含从真实犬猫问诊中提取的10,076个纯文本和8,751个多模态问答对,并配有专家兽医给出的答案。其测试集PetQA-Bench进一步标注了问题类型和临床状况。我们在零样本推理、检索增强生成(RAG)和监督微调(SFT)三种设置下,使用ROUGE、BERTScore及LLM-as-a-judge指标评估18个模型在事实性与实用性方面的表现。结果揭示当前模型在处理兽医临床问题时存在明显局限,凸显开发更有效适应方法以构建可信赖兽医AI系统的必要性。为促进广泛应用,我们还提供了PetQA-Bench的五种语言翻译版本。
原文摘要 · Abstract (English)
We introduce PetQA, a Korean long-form question-answering (QA) benchmark for evaluating veterinary knowledge and clinical reasoning in large language models (LLMs) and large vision-language models (LVLMs). PetQA contains 10,076 text-only and 8,751 multimodal QA pairs derived from real-world questions about dogs and cats, paired with answers from expert veterinarians. Its test split, PetQA-Bench, further includes annotations for question types and clinical conditions. We evaluate eighteen models using ROUGE, BERTScore, and LLM-as-a-judge metrics for factuality and helpfulness under three settings: zero-shot inference, retrieval-augmented generation (RAG), and supervised fine-tuning (SFT). The benchmarking results provide an overview of the strengths and limitations of current models in addressing veterinary clinical queries and highlight the need for more effective adaptation methods to develop clinically reliable AI systems for veterinary care. To facilitate broader use, we additionally provide translated versions of PetQA-Bench in five languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。