让大模型学会自我评估,提升医疗推理效率与准确率。
MedCoG: Maximizing LLM Inference Density in Medical Reasoning via Meta-Cognitive Regulation
- 通过自我评估任务复杂度、熟悉度等动态调节知识使用。
- 在5个医学基准上实现6.2倍推理密度提升。
- 适合追求高效精准医疗AI的开发者和研究者。
大型语言模型在复杂医疗推理中展现潜力,但其推理扩展性遵循收益递减规律。现有研究虽引入多种知识类型,却未明确额外成本如何转化为准确率提升。本文提出MedCoG——基于知识图谱的医学元认知代理,利用模型对自身认知状态的自我评估(如任务复杂度、熟悉度、知识密度)来动态调控程序性、情景性和事实性知识的使用。该以大模型为中心的按需推理机制,通过避免盲目扩展和过滤干扰知识,缓解了扩展定律下的收益递减问题。我们实证刻画了扩展曲线,并引入推理密度指标量化推理效率。实验表明,MedCoG在五个高难度医学基准上实现6.2倍推理密度提升。此外,基于理想假设的模拟研究(Oracle study)进一步验证了元认知调控的巨大潜力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown strong potential in complex medical reasoning yet face diminishing gains under inference scaling laws. While existing studies augment LLMs with various knowledge types, it remains unclear how effectively the additional costs translate into accuracy. In this paper, we explore how meta-cognition of LLMs, i.e., their self-assessment of their own cognitive states, can regulate the reasoning process. Specifically, we propose MedCoG, a Medical Meta-Cognition Agent with Knowledge Graph, where the meta-cognitive assessments of task complexity, familiarity, and knowledge density dynamically regulate utilization of procedural, episodic, and factual knowledge. The LLM-centric on-demand reasoning aims to mitigate the diminishing returns under scaling law by (1) reducing costs via avoiding indiscriminate scaling, (2) improving accuracy via filtering out distractive knowledge. To validate this, we empirically characterize the scaling curve and introduce inference density to quantify inference efficiency. Experiments demonstrate the effectiveness and efficiency of MedCoG on five hard sets of medical benchmarks, yielding 6.2x inference density. Furthermore, the Oracle study highlights the significant potential of meta-cognitive regulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。