arXiv:2603.22291cs.CL2026-03

评估大模型对尼泊尔语生殖健康问答的准确性与安全性

Evaluating Large Language Models' Responses to Sexual and Reproductive Health Queries in Nepali

  • 构建多维度评估框架,涵盖准确性和文化适配性
  • 仅35.1%的回答在准确、充分、安全上达标
  • 揭示不同ChatGPT版本在可用性与安全性上的差异

随着大语言模型(LLMs)融入日常生活,人们越来越多地用其咨询性与生殖健康(SRH)问题,实现匿名交流且无心理负担。然而现有评估方法多聚焦高资源语言中客观问题的准确性,缺乏对低资源语言及敏感领域如SRH的可用性与安全性评估标准。本文提出LLM评估框架(LEAF),从准确性、语言、可用性缺口(相关性、充分性、文化适切性)和安全缺口(安全性、敏感性、保密性)四方面进行评估。基于该框架,我们对超过9000名用户提出的14,000条尼泊尔语SRH查询进行了人工标注,由生殖健康专家依据标准评分。结果显示,仅有35.1%的回答被评定为‘恰当’,即在准确、充分性以及无重大可用性或安全缺陷方面表现良好。研究还发现不同ChatGPT版本间虽准确性相似,但在可用性与安全性方面存在差异。该评估揭示了当前大模型在处理敏感议题上的显著局限,强调改进必要性。LEAF框架具有跨领域、跨语言适应性,尤其适用于对可用性与安全性要求高的场景。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) become integrated into daily life, they are increasingly used for personal queries, including Sexual and Reproductive Health (SRH), allowing users to chat anonymously without fear of judgment. However, current evaluation methods primarily focus on accuracy, often for objective queries in high-resource languages, and lack criteria to assess usability and safety, especially for low-resource languages and culturally sensitive domains like SRH. This paper introduces LLM Evaluation Framework (LEAF), that conducts assessments across multiple criteria: accuracy, language, usability gaps (including relevance, adequacy, and cultural appropriateness), and safety gaps (safety, sensitivity, and confidentiality). Using the LEAF framework, we assessed 14K SRH queries in Nepali from over 9K users. Responses were manually annotated by SRH experts according to the framework. Results revealed that only 35.1% of the responses were "proper", meaning they were accurate, adequate and had no major usability or safety related gaps. Insights include differences in performance between ChatGPT versions, such as similar accuracy but varying usability and safety aspects. This evaluation highlights significant limitations of current LLMs and underscores the need for improvement. The LEAF Framework is adaptable across domains and languages, particularly where usability and safety are critical, offering a pathway to better address sensitive topics.

大模型评估生殖健康低资源语言安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。