arXiv:2601.09756cs.CRcs.AI2026-01

用合成数据提升兽医病历去标识化安全,但不能完全替代真实数据。

Synthetic Data for Veterinary EHR De-identification: Benefits, Limits, and Safety Trade-offs Under Fixed Compute

  • 用大模型生成无标识的合成病历,测试其在不同训练策略下的表现。
  • 合成数据适度混合可提升去标识效果,但高比例使用会增加信息泄露风险。
  • 合成数据适合扩大训练量,但无法独立保障隐私安全,适合有算力限制的研究者。

兽医电子健康记录(vEHRs)包含敏感身份信息,限制了二次使用。尽管PetEVAL提供了该领域的基准,但数据资源仍匮乏。本研究评估大型语言模型(LLM)生成的合成文本在不同训练策略下对去标识化安全的影响,重点考察(i)合成数据增强和(ii)固定预算替换。基于PetEVAL衍生语料库(3,750条测试/1,249条训练),生成10,382条合成病历,采用“仅模板”隐私保护策略(标识符在提示前移除)。使用三种Transformer模型(PetBERT、VetBERT、Bio_ClinicalBERT)在不同混合比例下训练。以文档级泄露率(至少遗漏一个标识符的文档占比)为主要安全指标。结果显示,在固定样本替换下,用合成数据替代真实数据导致泄露率单调上升,表明合成数据不可安全替代真实监督;在计算量匹配条件下,适度混合合成数据可达到与纯真实数据相当的性能,但高比例合成会降低效果。相反,按训练轮次扩展增强则带来性能提升:PetBERT的重叠分数从0.831升至0.850±0.014,泄露率从6.32%降至4.02%±0.19%。但这些改进主要源于训练暴露增加,而非合成数据本身质量。语料诊断发现合成与真实数据在病历长度和标签分布上存在系统性差异,与持续泄露现象一致。结论:合成数据增强有助于扩大训练暴露,但仅能互补,不能替代真实数据用于关键隐私场景。

原文摘要 · Abstract (English)

Veterinary electronic health records (vEHRs) contain privacy-sensitive identifiers that limit secondary use. While PetEVAL provides a benchmark for veterinary de-identification, the domain remains low-resource. This study evaluates whether large language model (LLM)-generated synthetic narratives improve de-identification safety under distinct training regimes, emphasizing (i) synthetic augmentation and (ii) fixed-budget substitution. We conducted a controlled simulation using a PetEVAL-derived corpus (3,750 holdout/1,249 train). We generated 10,382 synthetic notes using a privacy-preserving "template-only" regime where identifiers were removed prior to LLM prompting. Three transformer backbones (PetBERT, VetBERT, Bio_ClinicalBERT) were trained under varying mixtures. Evaluation prioritized document-level leakage rate (the fraction of documents with at least one missed identifier) as the primary safety outcome. Results show that under fixed-sample substitution, replacing real notes with synthetic ones monotonically increased leakage, indicating synthetic data cannot safely replace real supervision. Under compute-matched training, moderate synthetic mixing matched real-only performance, but high synthetic dominance degraded utility. Conversely, epoch-scaled augmentation improved performance: PetBERT span-overlap F1 increased from 0.831 to 0.850 +/- 0.014, and leakage decreased from 6.32% to 4.02% +/- 0.19%. However, these gains largely reflect increased training exposure rather than intrinsic synthetic data quality. Corpus diagnostics revealed systematic synthetic-real mismatches in note length and label distribution that align with persistent leakage. We conclude that synthetic augmentation is effective for expanding exposure but is complementary, not substitutive, for safety-critical veterinary de-identification.

合成数据去标识化兽医AI隐私安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。