arXiv:2601.19191cs.CLcs.LG2026-01

为临床语言模型打造可审计的透明化发布工具包

Transparency-First Medical Language Models: Datasheets, Model Cards, and End-to-End Data Provenance for Clinical NLP

  • 构建统一的模型卡片、数据清单和溯源文件,支持机器检查
  • 在49.8万条合成病历上验证模型,识别774万处隐私信息
  • 适合医疗AI开发者与监管者用于模型可信性评估

我们提出TeMLM,一套以透明性为核心的临床语言模型发布工具包。TeMLM将数据溯源、数据透明度、建模透明度与治理规范整合为可机器校验的发布包。定义了TeMLM-Card、TeMLM-Datasheet、TeMLM-Provenance三类文件及轻量级合规检查清单,实现可重复审计。在包含49.8万条病历、774万处隐私信息标注(10类)及ICD-9-CM诊断标签的合成数据集Technetium-I上实例化,并报告了约1亿参数的ProtactiniumBERT在隐私信息去标识化(词级别分类)与前50个ICD-9编码提取(多标签分类)上的基准结果。强调合成基准对工具链验证有价值,但模型部署前必须在真实临床数据上验证。

原文摘要 · Abstract (English)

We introduce TeMLM, a set of transparency-first release artifacts for clinical language models. TeMLM unifies provenance, data transparency, modeling transparency, and governance into a single, machine-checkable release bundle. We define an artifact suite (TeMLM-Card, TeMLM-Datasheet, TeMLM-Provenance) and a lightweight conformance checklist for repeatable auditing. We instantiate the artifacts on Technetium-I, a large-scale synthetic clinical NLP dataset with 498,000 notes, 7.74M PHI entity annotations across 10 types, and ICD-9-CM diagnosis labels, and report reference results for ProtactiniumBERT (about 100 million parameters) on PHI de-identification (token classification) and top-50 ICD-9 code extraction (multi-label classification). We emphasize that synthetic benchmarks are valuable for tooling and process validation, but models should be validated on real clinical data prior to deployment.

医疗AI模型透明数据溯源隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。