arXiv:2601.07309cs.AIcs.LG2026-01被引 2

无需训练,通过角色引导的神经元移植,让大模型代理跨环境通用

ARM: Role-Conditioned Neuron Transplantation for Training-Free Generalist LLM Agent Merging

  • 基于角色激活分析选择最优神经元,实现精准移植
  • 在多个交互任务中超越专家模型和现有融合方法
  • 适合需要快速部署、跨场景通用的大模型应用

交互式大语言模型代理发展迅速,但多数仅适配单一环境,难以在其他环境中稳健表现。模型融合提供了一种无训练替代方案,将多个专家模型整合为单一模型。本文提出代理角色融合(ARM),一种基于激活引导、角色条件的神经元移植方法,用于大模型代理的模型融合。ARM将现有融合方法从静态自然语言任务拓展至多轮交互场景,显著提升跨多种交互环境的泛化能力。该方法采用三步框架:1)构建融合主干网络;2)基于角色条件激活分析进行选择;3)通过神经元移植实现细粒度优化。无需梯度优化,ARM在跨基准测试中实现更强泛化性与高效率。在多样化领域中,经ARM融合的模型优于先前模型融合方法及特定领域专家模型,并展现出强大的跨领域泛化能力。

原文摘要 · Abstract (English)

Interactive large language model agents have advanced rapidly, but most remain specialized to a single environment and fail to adapt robustly to other environments. Model merging offers a training-free alternative by integrating multiple experts into a single model. In this paper, we propose Agent-Role Merging (ARM), an activation-guided, role-conditioned neuron transplantation method for model merging in LLM agents. ARM improves existing merging methods from static natural language tasks to multi-turn agent scenarios, and over the generalization ability across various interactive environments. This is achieved with a well designed 3-step framework: 1) constructing merged backbones, 2) selection based on its role-conditioned activation analysis, and 3) neuron transplantation for fine-grained refinements. Without gradient-based optimization, ARM improves cross-benchmark generalization while enjoying efficiency. Across diverse domains, the model obtained via ARM merging outperforms prior model merging methods and domain-specific expert models, while demonstrating strong out-of-domain generalization.

大模型融合角色感知零训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。