arXiv:2507.02628cs.LG2025-07

用大模型自动生成医学数据检测用例,自动发现病历数据与医学常识的矛盾。

A Generative Approach for Semantic Auditing of Electronic Health Records

  • 借鉴软件测试思想,用大模型生成上下文感知的检测用例
  • 在三个数据集上每组生成数十条测试用例,发现数据分布异常
  • 适合医疗AI可信性验证、数据质量审计的研究者与开发者

临床人工智能的可靠性依赖高质量数据,但电子健康记录常与现有科学知识不一致。现有质量评估方法受限:或仅关注语法,或依赖耗时的手动规则来捕捉语义细节。为突破可扩展性瓶颈,我们提出医学数据啄食法(Medical Data Pecking),借鉴软件单元测试原理,引入语义数据覆盖率,利用大语言模型生成上下文感知的测试用例,以“啄食”方式探测观测数据与流行病学证据之间的不一致。我们构建了一个基于检索增强生成(Retrieval-Augmented Generation)架构的参考工具,将医学文献转化为可执行代码。应用于三个数据集时,该工具每队列生成数十个测试用例,识别出观测分布与流行病学先验之间的差异,这些差异既包含真实的数据不一致,也包含预期的队列选择效应。本工作为可扩展的语义审计提供了初步框架,推动可信AI所需的质量保障从手动规则转向生成式与上下文敏感的验证。

原文摘要 · Abstract (English)

The reliability of clinical artificial intelligence (AI) depends on high-quality data, yet Electronic Health Records are often inconsistent with existing scientific knowledge. Current quality assessments are limited: they either focus on syntax or rely on labor-intensive manual rules to capture semantic nuances. To overcome these scalability barriers, we propose Medical Data Pecking, a methodology that adopts software unit testing principles for medical data validation. It introduces Semantic Data Coverage, employing Large Language Models to generate context-aware tests that "peck" for inconsistencies between observed data and epidemiological evidence. To demonstrate this methodology, we implemented a reference tool using a Retrieval-Augmented Generation architecture that synthesizes medical literature into executable code. When applied to three datasets, this implementation generated dozens of tests per cohort, identifying discrepancies between observed distributions and epidemiological priors. These discrepancies encompass both genuine data inconsistencies and expected cohort-selection effects. This work provides an initial framework for scalable semantic auditing, shifting assurance from manual rules to the generative and context-sensitive verification required for trustworthy AI.

医疗AI数据审计大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。