arXiv:2507.21815cs.CLcs.CY2025-07被引 1

评测大模型在药物减害信息提供中的准确性和安全风险

HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs

  • 构建包含2160组问答证据的数据集,覆盖安全边界、定量信息与多药风险推断
  • 顶尖大模型在减害信息上仍存在错误,部分回答可能引发严重健康风险
  • 适合关注AI医疗安全、公共健康与伦理的开发者和研究者

数百万使用者的健康受物质滥用带来的危害影响。减害作为公共卫生策略,旨在改善其健康结果并降低安全风险。一些大型语言模型(LLMs)已展现出不错的医学知识水平,有望满足药物使用者(PWUD)的信息需求。然而,它们在相关任务中的表现仍缺乏系统评估。我们提出HRIPBench,一个用于评估大模型在减害信息提供中准确性与安全风险的基准。数据集HRIP-Basic包含2,160个问题-答案-证据对,涵盖三项任务:检查安全边界、提供定量数值、推断多药使用风险。我们设计了指令和RAG两种方案,评估模型基于自身知识与领域知识融合的表现。结果显示,当前顶尖大模型在提供减害信息时仍存在显著错误,甚至可能对使用者造成严重安全风险。因此,在减害场景中使用大模型应谨慎约束,以防引发负面健康后果。警告:本文包含可能诱发危害的非法内容。

原文摘要 · Abstract (English)

Millions of individuals' well-being are challenged by the harms of substance use. Harm reduction as a public health strategy is designed to improve their health outcomes and reduce safety risks. Some large language models (LLMs) have demonstrated a decent level of medical knowledge, promising to address the information needs of people who use drugs (PWUD). However, their performance in relevant tasks remains largely unexplored. We introduce HRIPBench, a benchmark designed to evaluate LLM's accuracy and safety risks in harm reduction information provision. The benchmark dataset HRIP-Basic has 2,160 question-answer-evidence pairs. The scope covers three tasks: checking safety boundaries, providing quantitative values, and inferring polysubstance use risks. We build the Instruction and RAG schemes to evaluate model behaviours based on their inherent knowledge and the integration of domain knowledge. Our results indicate that state-of-the-art LLMs still struggle to provide accurate harm reduction information, and sometimes, carry out severe safety risks to PWUD. The use of LLMs in harm reduction contexts should be cautiously constrained to avoid inducing negative health outcomes. WARNING: This paper contains illicit content that potentially induces harms.

大模型评估减害医疗安全AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。