arXiv:2505.05070cs.CL2025-05被引 11

评测9个大模型在孟加拉语医疗查询摘要上的表现,发现无需训练也能达高质结果。

Performance Evaluation of Large Language Models in Bangla Consumer Health Query Summarization

  • 零样本测试9个大模型对孟加拉语医疗问诊的摘要能力。
  • Mixtral-8x22b-Instruct在ROUGE-1和ROUGE-L上最优,接近微调模型表现。
  • 为低资源语言医疗问答提供可扩展的智能摘要方案,适合医疗AI研究者。

孟加拉语消费者健康查询(CHQs)属于低资源语言,常含冗余信息,影响高效医疗回应。本研究评估了九种先进大语言模型(LLMs):GPT-3.5-Turbo、GPT-4、Claude-3.5-Sonnet、Llama3-70b-Instruct、Mixtral-8x22b-Instruct、Gemini-1.5-Pro、Qwen2-72b-Instruct、Gemma-2-27b和Athene-70B,在零样本场景下的孟加拉语CHQ摘要性能。基于包含2,350个标注查询-摘要对的BanglaCHQ-Summ数据集,使用ROUGE指标与经过微调的SOTA模型Bangla T5进行对比。结果显示,Mixtral-8x22b-Instruct在ROUGE-1和ROUGE-L上表现最佳,而Bangla T5在ROUGE-2上领先。实验表明,零样本大模型可在无需特定任务训练的情况下生成高质量摘要,展现了其在低资源语言医疗信息处理中的巨大潜力,为构建可扩展的医疗问答摘要系统提供了有力支持。

原文摘要 · Abstract (English)

Consumer Health Queries (CHQs) in Bengali (Bangla), a low-resource language, often contain extraneous details, complicating efficient medical responses. This study investigates the zero-shot performance of nine advanced large language models (LLMs): GPT-3.5-Turbo, GPT-4, Claude-3.5-Sonnet, Llama3-70b-Instruct, Mixtral-8x22b-Instruct, Gemini-1.5-Pro, Qwen2-72b-Instruct, Gemma-2-27b, and Athene-70B, in summarizing Bangla CHQs. Using the BanglaCHQ-Summ dataset comprising 2,350 annotated query-summary pairs, we benchmarked these LLMs using ROUGE metrics against Bangla T5, a fine-tuned state-of-the-art model. Mixtral-8x22b-Instruct emerged as the top performing model in ROUGE-1 and ROUGE-L, while Bangla T5 excelled in ROUGE-2. The results demonstrate that zero-shot LLMs can rival fine-tuned models, achieving high-quality summaries even without task-specific training. This work underscores the potential of LLMs in addressing challenges in low-resource languages, providing scalable solutions for healthcare query summarization.

大模型医疗摘要低资源语言孟加拉语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。