用社区真实需求构建医疗大模型评估基准,让评测更贴近现实。
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
- 联合民间组织与居民共创评估流程,确保问题源于真实场景。
- 在印度医疗场景中验证多语言大模型对复杂健康问题的理解能力。
- 提供可扩展的本土化评估方案,适合关注公平性与文化适配的研究者。
大型语言模型通常通过通用或领域特定的基准进行评估,但这些测试往往脱离终端用户的实际生活情境。医疗等关键领域需要超越人为或模拟任务的评估,反映社区的日常需求、文化实践和细微语境。本文提出 Samiksha,一个由公民社会组织(CSOs)和社区成员共同设计的社区驱动型评估流程。该方法通过文化敏感、社区参与的自动化管道,使社区反馈决定评估内容、基准构建方式及结果评分标准。我们在印度医疗场景中展示了该方法的应用。分析表明,当前多语言大模型能够回应复杂的社区健康咨询,同时为实现情境化、包容性的大模型评估提供了可扩展路径。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are typically evaluated through general or domain-specific benchmarks testing capabilities that often lack grounding in the lived realities of end users. Critical domains such as healthcare require evaluations that extend beyond artificial or simulated tasks to reflect the everyday needs, cultural practices, and nuanced contexts of communities. We propose Samiksha, a community-driven evaluation pipeline co-created with civil-society organizations (CSOs) and community members. Our approach enables scalable, automated benchmarking through a culturally aware, community-driven pipeline in which community feedback informs what to evaluate, how the benchmark is built, and how outputs are scored. We demonstrate this approach in the health domain in India. Our analysis highlights how current multilingual LLMs address nuanced community health queries, while also offering a scalable pathway for contextually grounded and inclusive LLM evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。