arXiv:2509.16360cs.CL2025-09被引 2

测试大模型在公共卫生问答中的可读性,发现多数模型表达不清。

RephQA: Evaluating Readability of Large Language Models in Public Health Question Answering

  • 构建包含533组问答的RephQA基准,评估模型表达清晰度。
  • 25个模型中多数未达可读性标准,专业术语过多影响理解。
  • 改进策略中,自适应调参的GRPO方法效果最佳,适合公众使用。

大型语言模型(LLMs)在解决复杂医疗问题方面具有潜力。然而,以往研究多关注准确性与推理能力,忽视了模型回答在公共卫生问答中的可读性——即能否以普通人能理解的方式清晰传达信息。为此,本文提出RephQA基准,涵盖来自13个主题、27个来源的533组专家评审问答对,并设计代理多项选择任务评估信息量,结合Flesch-Kincaid年级水平与专业评分两个可读性指标。对25个LLM的评估显示,多数模型未能达到可读性标准,暴露出推理能力与有效沟通之间的差距。为提升可读性,本文探索四种策略:标准提示、思维链提示、组相对策略优化(GRPO)及其基于词元的变体。其中,词元自适应的GRPO表现最优,推动更实用、用户友好的公共健康问答系统发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) hold promise in addressing complex medical problems. However, while most prior studies focus on improving accuracy and reasoning abilities, a significant bottleneck in developing effective healthcare agents lies in the readability of LLM-generated responses, specifically, their ability to answer public health problems clearly and simply to people without medical backgrounds. In this work, we introduce RephQA, a benchmark for evaluating the readability of LLMs in public health question answering (QA). It contains 533 expert-reviewed QA pairs from 27 sources across 13 topics, and includes a proxy multiple-choice task to assess informativeness, along with two readability metrics: Flesch-Kincaid grade level and professional score. Evaluation of 25 LLMs reveals that most fail to meet readability standards, highlighting a gap between reasoning and effective communication. To address this, we explore four readability-enhancing strategies-standard prompting, chain-of-thought prompting, Group Relative Policy Optimization (GRPO), and a token-adapted variant. Token-adapted GRPO achieves the best results, advancing the development of more practical and user-friendly public health agents. These results represent a step toward building more practical agents for public health.

可读性公共健康大模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。