arXiv:2601.09717cs.CLcs.AI2026-01被引 1

用大模型自动分类线上医疗对话的隐私风险,符合国家标准。

SALP-CG: Standard-Aligned LLM Pipeline for Classifying and Grading Large Volumes of Online Conversational Health Data

  • 结合提示引导与结构化输出,确保分类合规性。
  • 在中文医疗对话数据集上最高达0.900微F1值,敏感度分级准确。
  • 适合需要合规处理健康数据的机构或研究者使用。

在线医疗咨询产生大量包含个人健康信息的对话数据,需依据政策标准进行类别分类与风险等级划分。现有方法缺乏统一规范与可靠自动化手段。本文提出基于大模型的SALP-CG提取管道,依据国标GB/T 39725-2020制定健康数据分类与分级规则。通过少样本提示、JSON Schema约束解码及确定性高风险规则,该后端无关管道在多种大模型上实现强类别合规与可靠敏感度评估。在MedDialog-CN基准上,模型表现稳定,实体识别准确,模式符合率高,最高模型在最高等级预测中达到微F1=0.900。敏感度分层显示,二级至三级条目占主导,组合后易导致身份重识别;四级至五级条目虽少但危害极大。SALP-CG可跨模型可靠完成在线医疗对话的分类与敏感度分级,为健康数据治理提供实用方案。代码已开源。

原文摘要 · Abstract (English)

Online medical consultations generate large volumes of conversational health data that often embed protected health information, requiring robust methods to classify data categories and assign risk levels in line with policies and practice. However, existing approaches lack unified standards and reliable automated methods to fulfill sensitivity classification for such conversational health data. This study presents a large language model-based extraction pipeline, SALP-CG, for classifying and grading privacy risks in online conversational health data. We concluded health-data classification and grading rules in accordance with GB/T 39725-2020. Combining few-shot guidance, JSON Schema constrained decoding, and deterministic high-risk rules, the backend-agnostic extraction pipeline achieves strong category compliance and reliable sensitivity across diverse LLMs. On the MedDialog-CN benchmark, models yields robust entity counts, high schema compliance, and accurate sensitivity grading, while the strongest model attains micro-F1=0.900 for maximum-level prediction. The category landscape stratified by sensitivity shows that Level 2-3 items dominate, enabling re-identification when combined; Level 4-5 items are less frequent but carry outsize harm. SALP-CG reliably helps classify categories and grading sensitivity in online conversational health data across LLMs, offering a practical method for health data governance. Code is available at https://github.com/dommii1218/SALP-CG.

大模型隐私保护医疗数据分类分级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。