arXiv:2603.18898cs.IR2026-03

对比三大模型在泰卢固语孕产问答表现,发现选模型和用语言都关键。

Comparative Analysis of Large Language Models in Generating Telugu Responses for Maternal Health Queries

  • 用双语数据集测试三模型对泰卢固语孕产问题的回答能力。
  • Gemini在准确性与连贯性上最优,Perplexity在泰卢固语提示下表现好。
  • 研究强调改进区域语言医疗大模型的重要性,适合健康AI开发者参考。

大型语言模型(LLMs)在多个研究领域展现出潜力,但在低资源语言如泰卢固语、印地语、泰米尔语、乌尔都语等的急性孕产健康领域的表现仍缺乏研究。本研究评估了ChatGPT-4o、GeminiAI和Perplexity AI对不同语言提出的孕产相关问题的响应。采用双语数据集,结合BERT Score语义相似度指标与妇产科专家评估,从准确性、流畅性、相关性、连贯性和完整性等多维度评分。结果显示,Gemini在泰卢固语孕产相关回答中表现最佳,兼具准确与连贯;Perplexity在泰卢固语提示下亦表现良好;ChatGPT-4o仍有提升空间。研究强调,选择合适的模型与使用目标语言提示对获取有效医疗信息至关重要,呼吁提升大模型在区域语言医疗应用中的支持能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have been progressively exhibiting there capabilities in various areas of research. The performance of the LLMs in acute maternal healthcare area, predominantly in low resource languages like Telugu, Hindi, Tamil, Urdu etc are still unstudied. This study presents how ChatGPT-4o, GeminiAI, and Perplexity AI respond to pregnancy related questions asked in different languages. A bilingual dataset is used to obtain results by applying the semantic similarity metrics (BERT Score) and expert assessments from expertise gynecologists. Multiple parameters like accuracy, fluency, relevance, coherence and completeness are taken into consideration by the gynecologists to rate the responses generated by the LLMs. Gemini excels in other LLMs in terms of producing accurate and coherent pregnancy relevant responses in Telugu, while Perplexity demonstrated well when the prompts were in Telugu. ChatGPT's performance can be improved. The results states that both selecting an LLM and prompting language plays a crucial role in retrieving the information. Altogether, we emphasize for the improvement of LLMs assistance in regional languages for healthcare purposes.

大模型医疗问答泰卢固语孕产健康

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。