为临床聊天机器人设计细粒度安全分类体系,提升响应安全性。
Taxonomy of Comprehensive Safety for Clinical Agents
- 将安全过滤与工具选择整合到用户意图分类中,构建21类细粒度安全标签
- 在自建数据集上验证,发现基础模型对临床安全知识存在分布偏差
- 适合医疗AI研发者、安全评估人员使用,助力临床应用落地
临床聊天机器人中的安全性至关重要,错误或有害回应可能引发严重后果。现有方法如防护机制和工具调用,难以满足临床场景的复杂需求。本文提出TACOS(TAxonomy of COmprehensive Safety for Clinical Agents),一个涵盖21个类别的细粒度安全分类体系,将安全过滤与工具选择统一纳入用户意图分类步骤。该分类体系覆盖广泛临床与非临床查询,显式建模不同安全阈值与外部工具依赖关系。为验证其有效性,我们构建了标注的TACOS数据集并开展大量实验。结果表明,专用于临床代理场景的新分类体系具有显著价值,并揭示了训练数据分布及基线模型预训练知识的重要洞见。
原文摘要 · Abstract (English)
Safety is a paramount concern in clinical chatbot applications, where inaccurate or harmful responses can lead to serious consequences. Existing methods--such as guardrails and tool calling--often fall short in addressing the nuanced demands of the clinical domain. In this paper, we introduce TACOS (TAxonomy of COmprehensive Safety for Clinical Agents), a fine-grained, 21-class taxonomy that integrates safety filtering and tool selection into a single user intent classification step. TACOS is a taxonomy that can cover a wide spectrum of clinical and non-clinical queries, explicitly modeling varying safety thresholds and external tool dependencies. To validate our taxonomy, we curate a TACOS-annotated dataset and perform extensive experiments. Our results demonstrate the value of a new taxonomy specialized for clinical agent settings, and reveal useful insights about train data distribution and pretrained knowledge of base models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。