arXiv:2608.29965cs.AI2026-08综述

让AI生成的医疗数据必须经源文件验证才可入档,防止虚假信息自我认证。

Review Before Trust: Source-Grounded Integrity Gates for AI-Assisted Personal Health Records

  • 用确定性监控器检查生成数据是否来自源文档中的唯一引用句
  • 97个数值结果中仅72个通过严格校验,25个留待人工复核
  • 适合关注医疗记录可信度的开发者与临床系统设计者

大语言模型可将医疗文档转为结构化数据,但生成内容可能缺乏源文件支持。若此类未经验证的数据存入长期健康记录(随时间积累患者信息),将带来数据完整性风险:未核实的信息可能影响后续摘要、趋势分析或预防性护理计算。本文提出一种证据约束的信任提升模型,确保生成数据在源文档验证前保持暂定状态。只有当源文档包含唯一支持引文、相关字段位于同一实验室行内且溯源信息完整时,候选数据才被允许用于下游用途。生成器无法自审,证据缺失或模糊则拒绝,被拒条目仍保留供人工审查而非无声丢弃。我们在个人健康记录应用Medical DataCloud中实现该模型,并通过自动化测试与历史提取结果回放进行评估。所有22项一致性与变异测试均通过。回放覆盖9份含102个手动标注行的历史实验室PDF报告,生成97个数值候选:模式验证通过全部97个,早期包级证据检查通过94个,而强化后的引文与行级策略仅接纳72个,保留25个待人工审查。研究评估系统完整性而非临床正确性或安全性。结果证明了防止生成声明自我授权在长期记录中重复使用的可执行边界的技术可行性。

原文摘要 · Abstract (English)

Large language models can convert medical documents into structured data, but plausible output may still be unsupported by the source. Persisting such output in a longitudinal health record, a record that accumulates patient information over time, therefore creates an integrity risk: unverified data may influence later summaries, trends, or preventive-care computations. We introduce an evidence-gated trust-promotion model that keeps generated data provisional until a deterministic monitor verifies it against the source document. The monitor admits a candidate for a specified downstream use only when the source contains a unique supporting quotation, the relevant fields occur within the same laboratory row, and the required provenance is preserved. The generator cannot approve its own output, missing or ambiguous evidence causes refusal, and refused candidates remain available for human review rather than being silently discarded. We implement the model in Medical DataCloud, a personal health-record application, and evaluate it through automated tests and a replay of saved extraction outputs. All 22 conformance and mutation tests pass. The replay covers nine historical laboratory PDF reports containing 102 manually labelled rows. The reports produce 97 numeric candidates: schema validation accepts all 97, an earlier packet-level evidence check accepts 94, and the hardened quotation- and row-level policy admits 72 while retaining 25 for review. The study evaluates system integrity rather than clinical correctness or clinical safety. The results demonstrate the technical feasibility of an enforceable boundary that prevents generated claims from authorizing their own reuse in a longitudinal health record.

医疗AI可信生成数据完整性健康记录

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。