arXiv:2608.22085cs.AI2026-08

用符号规则验证生成的癌症数据,防止临床错误污染模型。

Dissecting Neuro-Symbolic Quality Assurance for Synthetic Oncology Data Generation

论文配图:Dissecting Neuro-Symbolic Quality Assurance for Synthetic Oncology Data Generation
图 1 · 摘自论文原文
  • 引入符号门控机制,确保数据符合医学编码和分期逻辑
  • 无验证时20.1%数据存在错误分期,门控后减少至极低水平
  • 检索增强效果因模型而异,不能作为通用优化手段

大型语言模型生成合成临床数据可缓解癌症分期研究的数据稀缺问题,但肿瘤学幻觉具有根本危害:一次临床不可能的分期会污染所有下游模型。神经符号管道在生成过程中进行验证,但各质量保障组件的作用尚不明确。我们通过三组受控实验,分别考察门控必要性、约束归属和检索条件性,保持生成协议、多样性阈值及微调超参数不变。符号门控强制执行结构完整性、基于医学系统命名法的本体覆盖,以及根据美国癌症联合委员会第八版规则的分期逻辑一致性。未启用门控时,29.9%记录存在结构缺陷,20.1%含临床无效分期。结构验证是核心过滤器:在完整门控语料中,它排除了512条中的148条,本体对齐进一步排除24条,而分期逻辑验证未排除任何记录——唯一产生逻辑错误的生成器已在结构层面被剔除,说明临床逻辑验证仅是生成器依赖的次级防护。检索增强效果高度依赖模型:对一个生成器提升门控合规率12.5个百分点,对第二个无影响,对第三个甚至导致输出崩溃。在所有门控配置下,本体密度基本不变,表明符号验证提升的是临床有效性而非词汇丰富度。因此,符号门控虽提升语料临床有效性,但未能带来真实肺癌病历的语义相似度提升;检索应针对每类模型评估,且本体密度不可作为语料质量的代理指标。

原文摘要 · Abstract (English)

Synthetic clinical data generation with large language models addresses the scarcity that limits cancer staging research, but oncology hallucinations are categorically harmful: one clinically impossible staging assignment contaminates every downstream model trained on it. Neuro-symbolic pipelines validate during generation, yet the contribution of individual quality-assurance components remains unclear. We report three controlled studies isolating gate necessity, constraint attribution, and retrieval conditionality, holding generation protocol, diversity thresholds, and fine-tuning hyperparameters constant across adapter conditions. The symbolic gate enforces schema completeness, ontology coverage against the Systematized Nomenclature of Medicine, and staging-logic consistency under American Joint Committee on Cancer eighth-edition rules. Ungated, 29.9% of records contain schema failures and 20.1% contain clinically invalid staging. Schema validation is the load-bearing filter: within the fully gated corpus it rejects 148 of 512 records, ontology grounding a further 24, and staging-logic validation none---the only generator producing logic violations is already excluded on schema, making clinical-logic validation a generator-conditional safeguard rather than the dominant filter. Retrieval augmentation is strongly model-dependent: it improves gate compliance for one generator by 12.5 percentage points, has no measurable effect for a second, and collapses output in a third. Across gated configurations ontology density is largely unchanged, indicating that symbolic validation improves clinical validity rather than vocabulary richness. Symbolic gating therefore buys corpus validity but no commensurate gain on real lung-cancer notes in this study; retrieval should be evaluated per model, and ontology density should not be reported as a proxy for corpus quality.

医疗生成符号验证数据质量癌症研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。