arXiv:2608.05195q-bio.GNcs.LG2026-08

让癌症基因组学问答更准:自动识别并澄清语言歧义

CLARA: Clarification of Language Ambiguity through Result Analysis for Natural-Language Cancer Genomics Queries

论文配图:CLARA: Clarification of Language Ambiguity through Result Analysis for Natural-Language Cancer Genomics Queries
图 1 · 摘自论文原文
  • 将自然语言问题转为可执行查询,通过多解释对比判断是否需澄清
  • 在330个对比任务中准确识别115个结果敏感的模糊问题
  • 适合临床研究者与生物信息学工具开发者参考

自然语言接口可降低癌症基因组数据库使用门槛,但即使问题语法正确,科学含义仍可能模糊。本文提出CLARA框架,将问题表示为带类型的科学查询规范,考虑多种可能解释,执行后若结果差异显著则请求澄清。评估在8个TCGA PanCancer Atlas队列和30基因面板上进行,包含330个唯一可执行对比,其中115个为结果敏感(相对差异>0.10或绝对差异>5个百分点),215个为结果稳定。独立实现的pandas执行引擎与SQLite完全一致,复现全部660个结果。在120个由LLM生成、人工审核的语言压力测试中,CLARA识别出所有60个结果敏感对比,对60个稳定对比中误澄清13个(准确率89.2%,灵敏度100%,特异性78.3%)。纯机器学习方法总体准确率更高(97.5%),但漏检一个关键对比。结果表明,下游执行能区分重要与无关歧义,并揭示安全与负担间的明确权衡。

原文摘要 · Abstract (English)

A natural language interface can be used to make cancer genomics databases easier to use, but even if a question is perfectly fluent, its scientific meaning can be ambiguous. We propose CLARA, a framework that represents a question as a typed scientific query specification, considers a few possible interpretations, executes them, and asks for clarification when the estimates diverge. CLARA was assessed on mutation-prevalence contrasts among eight TCGA PanCancer Atlas cohorts and a 30-gene panel. This benchmark consisted of 330 unique executable contrasts varying in mutation scope, assay denominator, and sample context; 115 contrasts were result-sensitive and 215 were result-stable, per the preregistered definition of relative divergence greater than 0.10 or absolute divergence greater than 5 percentage points. An independently implemented pandas execution engine perfectly replicated all 660 results from the SQLite engine. In a separate 120-question LLM-generated, manually vetted language stress test, CLARA recognized all 60 result-sensitive contrasts and needlessly clarified 13 of 60 stable contrasts (accuracy 89.2%, sensitivity/recall 100%, specificity 78.3%). Standalone machine learning had superior overall accuracy (97.5%) but missed one critical contrast. This demonstrates that downstream execution can distinguish consequential from inconsequential ambiguity and reveal an explicit trade-off between safety and burden.

自然语言查询癌症基因组歧义检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。