arXiv:2509.02401cs.AI2025-09被引 9

让大模型知道何时不懂,用不确定性提升医学数据推理可靠性

Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning

  • 用检索与摘要双重不确定性信号指导推理决策
  • 总结正确且有用的内容数量提升近3倍,生存预测性能翻倍
  • 适合需要高可信度的生物医学数据分析场景

大型语言模型代理在结构化生物医学数据环境中日益普及,但面对复杂多表数据时常生成流畅却过度自信的输出。本文提出一种基于不确定性的代理,用于查询条件下的多表摘要任务,利用两个互补信号:(i) 检索不确定性——通过多次表选择滚动计算的熵;(ii) 摘要不确定性——结合自一致性与困惑度。摘要不确定性被融入分组相对策略优化(GRPO)的强化学习中,而两者共同指导推理时的过滤,并支持构建更高质量的合成数据集。在多组学基准测试中,该方法显著提升事实性和校准度,每篇摘要中正确且有用的说法从3.0增至8.4(内部数据集),3.6增至9.9(癌症多组学),下游生存预测性能也大幅提升(C-index由0.32升至0.63)。结果表明,不确定性可作为控制信号,使代理学会自我克制、传达置信度,成为复杂结构化数据环境中的更可靠工具。

原文摘要 · Abstract (English)

Large language model (LLM) agents are increasingly deployed in structured biomedical data environments, yet they often produce fluent but overconfident outputs when reasoning over complex multi-table data. We introduce an uncertainty-aware agent for query-conditioned multi-table summarization that leverages two complementary signals: (i) retrieval uncertainty--entropy over multiple table-selection rollouts--and (ii) summary uncertainty--combining self-consistency and perplexity. Summary uncertainty is incorporated into reinforcement learning (RL) with Group Relative Policy Optimization (GRPO), while both retrieval and summary uncertainty guide inference-time filtering and support the construction of higher-quality synthetic datasets. On multi-omics benchmarks, our approach improves factuality and calibration, nearly tripling correct and useful claims per summary (3.0\(\rightarrow\)8.4 internal; 3.6\(\rightarrow\)9.9 cancer multi-omics) and substantially improving downstream survival prediction (C-index 0.32\(\rightarrow\)0.63). These results demonstrate that uncertainty can serve as a control signal--enabling agents to abstain, communicate confidence, and become more reliable tools for complex structured-data environments.

大模型推理不确定性量化生物医学分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。