arXiv:2510.00482cs.CL2025-10中稿 · AIxB 2025被引 1

通过自蒸馏微调大模型,在特定技术领域提升自主决策能力。

Agent Fine-tuning through Distillation for Domain-specific LLMs in Microdomains

  • 用领域手册和自生成推理轨迹训练模型,内化专业逻辑。
  • 在JP1认证题上比基线模型准确率提升14%。
  • 适合需要高精度专业推理的IT运维场景。

智能体大语言模型(LLM)在自主与外部环境交互及多步推理任务中表现突出。现有方法多依赖少样本提示进行上下文学习,但易导致输入过长、计算成本高。代理微调则通过在相关数据和示范轨迹上训练,使模型内化过程性推理与领域知识。尽管已有研究集中于通用领域,其在专业化技术微域中的效果尚不明确。本文探索了在日立JP1中间件这一特定IT运维微域中进行代理微调的方法。采用来自领域手册的JP1专属数据集,并通过大模型自身生成的推理轨迹进行蒸馏,提升了决策准确性和搜索效率。推理阶段结合检索增强生成与上下文-答案提取器,增强信息相关性。在JP1认证考试题上,该方法相较基线模型性能提升14%,证明了代理微调在复杂微域中实现领域专用推理的潜力。

原文摘要 · Abstract (English)

Agentic large language models (LLMs) have become prominent for autonomously interacting with external environments and performing multi-step reasoning tasks. Most approaches leverage these capabilities via in-context learning with few-shot prompts, but this often results in lengthy inputs and higher computational costs. Agent fine-tuning offers an alternative by enabling LLMs to internalize procedural reasoning and domain-specific knowledge through training on relevant data and demonstration trajectories. While prior studies have focused on general domains, their effectiveness in specialized technical microdomains remains unclear. This paper explores agent fine-tuning for domain adaptation within Hitachi's JP1 middleware, a microdomain for specialized IT operations. We fine-tuned LLMs using JP1-specific datasets derived from domain manuals and distilled reasoning trajectories generated by LLMs themselves, enhancing decision making accuracy and search efficiency. During inference, we used an agentic prompt with retrieval-augmented generation and introduced a context-answer extractor to improve information relevance. On JP1 certification exam questions, our method achieved a 14% performance improvement over the base model, demonstrating the potential of agent fine-tuning for domain-specific reasoning in complex microdomains.

大模型微调智能体系统领域适应知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。