arXiv:2602.00491cs.CL2026-02

构建多语言公共卫生推理数据集,提升大模型在真实场景下的安全决策能力。

From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlas

  • 基于检索与去重的LLM辅助构建流程,确保数据质量与一致性。
  • 覆盖15个领域、17种语言,共28万条高质量推理样本。
  • 提出六维评估体系,适用于医疗安全关键任务的模型评测。

公共卫生推理需基于科学证据、专家共识和安全约束进行群体层面推断,但其作为结构化机器学习问题仍缺乏系统研究,且监督信号与基准有限。本文提出GlobalHealthAtlas,一个大规模多语言数据集,包含280,210条实例,覆盖15个公共健康领域及17种语言。通过引入大语言模型辅助的数据构建与质量控制流程(含检索、去重、证据对齐检查与标签验证),实现大规模数据的一致性保障。进一步设计了一个基于高置信度多LLM判断的领域对齐评估器,从准确性、推理能力、完整性、共识一致性、术语规范性与洞察力六个维度评估输出。上述贡献共同支持面向安全关键任务的公共卫生推理模型可复现训练与评估,超越传统问答基准。项目代码、评估器与模型已开源:https://github.com/Jan8217/GlobalHealthAtlas, https://huggingface.co/aerovane0/GlobalHealthAtlas_Public_Evaluator, https://huggingface.co/aerovane0/GlobalHealthAtlas_Public_Model。

原文摘要 · Abstract (English)

Public health reasoning requires population level inference grounded in scientific evidence, expert consensus, and safety constraints. However, it remains underexplored as a structured machine learning problem with limited supervised signals and benchmarks. We introduce GlobalHealthAtlas, a large scale multilingual dataset of 280,210 instances spanning 15 public health domains and 17 languages. We further propose a large language model (LLM) assisted construction and quality control pipeline with retrieval, deduplication, evidence grounding checks, and label validation to improve consistency at scale. Finally, we present a domain aligned evaluator distilled from high confidence judgments of diverse LLMs to assess outputs along six dimensions: Accuracy, Reasoning, Completeness, Consensus Alignment, Terminology Norms, and Insightfulness. Together, these contributions enable reproducible training and evaluation of LLMs for safety critical public health reasoning beyond conventional QA benchmarks. We publicly release project codebase, evaluator, and model at:: https://github.com/Jan8217/GlobalHealthAtlas, https://huggingface.co/aerovane0/GlobalHealthAtlas_Public_Evaluator and https://huggingface.co/aerovane0/GlobalHealthAtlas_Public_Model

公共卫生大模型评估多语言推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。