arXiv:2508.17202cs.CL2025-08EMNLP被引 3

用百美元预算高效获取专家知识,提升专业领域大模型表现

Active Domain Knowledge Acquisition with 100-Dollar Budget: Enhancing LLMs via Cost-Efficient, Expert-Involved Interaction in Sensitive Domains

  • 主动筛选最合适的专家,兼顾能力、可用性与咨询成本
  • 在严格预算下使大模型在药物研发等场景性能显著提升
  • 适合需要专家协作但预算有限的医疗、科研类应用

大型语言模型虽具备广泛知识,但在药物研发、罕见病研究等高专业性且成本敏感的领域仍表现不佳,主要因缺乏专家知识。本文提出一种新型框架PU-ADKA,通过在固定预算内主动调用领域专家,显著提升领域专用大模型的能力。该框架能智能选择最适专家,综合考量其可用性、知识边界与咨询成本。我们基于PubMed数据进行模拟训练,并通过受控专家交互和真实药企团队部署验证了方法有效性。实验表明,在严格预算约束下,该方法可显著增强大模型在专业领域的表现。此外,本文还构建了一个新基准数据集CKAD,用于推动低成本大模型领域知识获取的研究。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated an impressive level of general knowledge. However, they often struggle in highly specialized and cost-sensitive domains such as drug discovery and rare disease research due to the lack of expert knowledge. In this paper, we propose a novel framework (PU-ADKA) designed to efficiently enhance domain-specific LLMs by actively engaging domain experts within a fixed budget. Unlike traditional fine-tuning approaches, PU-ADKA selectively identifies and queries the most appropriate expert from a team, taking into account each expert's availability, knowledge boundaries, and consultation costs. We train PU-ADKA using simulations on PubMed data and validate it through both controlled expert interactions and real-world deployment with a drug development team, demonstrating its effectiveness in enhancing LLM performance in specialized domains under strict budget constraints. In addition to outlining our methodological innovations and experimental results, we introduce a new benchmark dataset, CKAD, for cost-effective LLM domain knowledge acquisition to foster further research in this challenging area.

大模型专家系统药物研发成本优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。