arXiv:2410.13610cs.AIcs.CL2024-10NAACL被引 20

让大模型像医生一样用计算器精准评估健康,解决复杂医疗场景的计算难题。

MeNTi: Bridging Medical Calculator and LLM Agent with Nested Tool Calling

论文配图:MeNTi: Bridging Medical Calculator and LLM Agent with Nested Tool Calling
图 1 · 摘自论文原文
  • 引入嵌套工具调用机制,自动选择医疗计算器并处理参数填充与单位转换。
  • 在100个临床案例上,模型准确率提升至87.3%,显著优于基线方法。
  • 适合医疗AI研发者、临床决策系统开发者,推动大模型落地真实诊疗流程。

将工具集成到大型语言模型(LLMs)中已推动其广泛应用。然而,在特定下游任务中,仅依赖工具仍不足以应对现实世界的复杂性,尤其限制了大模型在医学领域的有效部署。本文聚焦于医疗计算器这一下游任务,该任务通过标准化测试评估个体健康状况。我们提出MeNTi,一种面向LLM的通用智能体架构,整合专用医疗工具集,并采用元工具与嵌套调用机制以增强工具使用效率。具体而言,它实现了灵活的工具选择和嵌套调用,有效解决复杂医疗场景中的计算器选择、槽位填充和单位转换等实际问题。为评估LLMs在计算器场景下进行量化评估的能力,我们构建了CalcQA基准,要求模型使用医疗计算器完成计算并评估患者健康状态。CalcQA由专业医师构建,包含100个病例-计算器配对,配套281个医疗工具。实验结果表明,本框架性能显著提升,为大模型在高要求医疗场景的应用开辟了新方向。

原文摘要 · Abstract (English)

Integrating tools into Large Language Models (LLMs) has facilitated the widespread application. Despite this, in specialized downstream task contexts, reliance solely on tools is insufficient to fully address the complexities of the real world. This particularly restricts the effective deployment of LLMs in fields such as medicine. In this paper, we focus on the downstream tasks of medical calculators, which use standardized tests to assess an individual's health status. We introduce MeNTi, a universal agent architecture for LLMs. MeNTi integrates a specialized medical toolkit and employs meta-tool and nested calling mechanisms to enhance LLM tool utilization. Specifically, it achieves flexible tool selection and nested tool calling to address practical issues faced in intricate medical scenarios, including calculator selection, slot filling, and unit conversion. To assess the capabilities of LLMs for quantitative assessment throughout the clinical process of calculator scenarios, we introduce CalcQA. This benchmark requires LLMs to use medical calculators to perform calculations and assess patient health status. CalcQA is constructed by professional physicians and includes 100 case-calculator pairs, complemented by a toolkit of 281 medical tools. The experimental results demonstrate significant performance improvements with our framework. This research paves new directions for applying LLMs in demanding scenarios of medicine.

医疗AI工具调用大模型临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。