现有评测体系偏西方文化,无法真实反映AI健康助手在印度等地区的实际表现。
Beyond the Rubric: Cultural Misalignment in LLM Benchmarks for Sexual and Reproductive Health
- 用印度社区的性与生殖健康问题测试,发现主流评测打分偏低但实际内容准确
- 330次对话中多数回答在文化适配和医学准确性上达标,但被西方标准误判
- 适合关注AI公平性、全球健康技术本土化的研究者与实践者
大型语言模型有望提升全球南方地区医疗信息可及性,但其评估仍依赖基于西方规范的基准。我们对一个面向印度弱势群体的性与生殖健康聊天机器人进行了初步基准测试,使用OpenAI的HealthBench基准。从数据集中提取637个性与生殖健康问题,评估330次单轮对话。自动化评分系统依据评分表(rubric)给出低分,但经训练标注员和公共卫生专家的定性分析发现,许多回答在文化和医学层面均准确。我们指出常见问题,包括法律框架与社会规范(如公开哺乳)、饮食假设(如孕期食用鱼类安全)以及成本结构(如保险模式)的西方偏见。研究揭示当前基准难以捕捉为不同文化与医疗背景设计系统的有效性,主张开发兼顾质量标准与多元需求的文化适应型评估框架。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have been positioned as having the potential to expand access to health information in the Global South, yet their evaluation remains heavily dependent on benchmarks designed around Western norms. We present insights from a preliminary benchmarking exercise with a chatbot for sexual and reproductive health (SRH) for an underserved community in India. We evaluated using HealthBench, a benchmark for conversational health models by OpenAI. We extracted 637 SRH queries from the dataset and evaluated on the 330 single-turn conversations. Responses were evaluated using HealthBench's rubric-based automated grader, which rated responses consistently low. However, qualitative analysis by trained annotators and public health experts revealed that many responses were actually culturally appropriate and medically accurate. We highlight recurring issues, particularly a Western bias, such as for legal framing and norms (e.g., breastfeeding in public), diet assumptions (e.g., fish safe to eat during pregnancy), and costs (e.g., insurance models). Our findings demonstrate the limitations of current benchmarks in capturing the effectiveness of systems built for different cultural and healthcare contexts. We argue for the development of culturally adaptive evaluation frameworks that meet quality standards while recognizing needs of diverse populations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。