arXiv:2512.20629cs.LGcs.AI2025-12

不微调模型参数,让语言智能体自主演化策略

Learning Evolving Latent Strategies for Multi-Agent Language Systems without Model Fine-Tuning

  • 用可更新的隐向量替代固定语义表示,实现策略持续进化
  • 多轮交互中隐空间收敛且关键节点有结构化变化
  • 适合需要低成本动态适应的复杂协作场景

本研究提出一种无需微调语言模型参数的多智能体语言框架,实现持续策略演化。核心思想是将抽象概念的隐向量从传统静态语义表征中解放,通过环境互动与强化反馈持续更新。构建双环架构:行为环根据环境奖励调整动作偏好,语言环通过生成文本的语义嵌入反思来更新外部隐向量。二者协同使智能体在长时多轮交互中发展出稳定且解耦的战略风格。实验显示,隐空间在反射驱动更新下呈现清晰收敛轨迹,关键时刻出现结构化迁移。系统还展现出隐式推断并持续适应情绪化智能体的能力,即使无共享奖励。结果表明,仅通过外部隐空间即可为语言智能体提供低成本、可扩展且可解释的抽象战略表示。

原文摘要 · Abstract (English)

This study proposes a multi-agent language framework that enables continual strategy evolution without fine-tuning the language model's parameters. The core idea is to liberate the latent vectors of abstract concepts from traditional static semantic representations, allowing them to be continuously updated through environmental interaction and reinforcement feedback. We construct a dual-loop architecture: the behavior loop adjusts action preferences based on environmental rewards, while the language loop updates the external latent vectors by reflecting on the semantic embeddings of generated text. Together, these mechanisms allow agents to develop stable and disentangled strategic styles over long-horizon multi-round interactions. Experiments show that agents' latent spaces exhibit clear convergence trajectories under reflection-driven updates, along with structured shifts at critical moments. Moreover, the system demonstrates an emergent ability to implicitly infer and continually adapt to emotional agents, even without shared rewards. These results indicate that, without modifying model parameters, an external latent space can provide language agents with a low-cost, scalable, and interpretable form of abstract strategic representation.

多智能体策略演化隐空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。