评估大模型在孕产咨询中的可靠性,发现Perplexity语义最准,ChatGPT-4o表达更清晰。
Trust, Safety, and Accuracy: Assessing LLMs for Routine Maternity Advice
- 用17个孕产问题测试3个大模型,对比专家回答的语义与可读性。
- Perplexity AI在语义相似度上最接近专业医生,ChatGPT-4o文本更易懂且术语准确。
- 适合关注乡村孕产健康、需兼顾准确与易懂的AI医疗应用开发者。
由于医疗资源和基础设施匮乏,印度农村地区获取可靠孕产健康信息面临重大挑战。尽管当地有超8.3亿互联网用户,近半数农村女性已上网,数字工具为健康教育带来新机遇。本研究评估了ChatGPT-4o、Perplexity AI和GeminiAI等大语言模型在提供孕产相关资讯方面的表现。针对17个孕产主题提问,将模型回复与孕产专家答案对比,使用语义相似度、名词重叠率及可读性指标衡量内容质量。结果显示,Perplexity AI在语义匹配度上最接近专家水平,而ChatGPT-4o生成文本更清晰,医学术语使用更准确。随着农村地区互联网普及,此类大模型有望成为可扩展的孕产健康教育辅助工具。研究强调,需发展兼顾准确性与表达清晰度的AI医疗工具,以改善弱势地区的医疗沟通效果。
原文摘要 · Abstract (English)
Access to reliable maternal healthcare information is a major challenge in rural India due to limited medical resources and infrastructure. With over 830 million internet users and nearly half of rural women online, digital tools offer new opportunities for health education. This study evaluates large language models (LLMs) like ChatGPT-4o, Perplexity AI, and GeminiAI to provide reliable and understandable pregnancy-related information. Seventeen pregnancy-focused questions were posed to each model and compared with responses from maternal health professionals. Evaluations used semantic similarity, noun overlap, and readability metrics to measure content quality. Results show Perplexity closely matched expert semantics, while ChatGPT-4o produced clearer, more understandable text with better medical terminology. As internet access grows in rural areas, LLMs could serve as scalable aids for maternal health education. The study highlights the need for AI tools that balance accuracy and clarity to improve healthcare communication in underserved regions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。