主代理协作学习中,用决策理论动态调整参数融合系数以达成最优共识。
A decision-theoretic model for a principal-agent collaborative learning problem
- 主代理基于测试集表现动态设定非负且和为1的聚合系数。
- 各智能体通过带平均场交互的Langevin动力学更新参数并达成共识。
- 无需了解数据分布或对方质量,仍具稳定性和泛化优势。
本文研究一种主代理协作学习框架,其中主方在每一步根据一组K个智能体基于独立测试集的表现,确定适当的聚合系数(非负且和为1),用于指导智能体的参数更新。智能体以离散时间Langevin动力学形式协同更新参数,其相互作用项具有平均场特征,但各自使用不同的训练数据集。本文提出一种决策理论框架,显式描述主方如何逐步确定聚合系数,最终使智能体达成一致的最优参数估计。由于智能体间固有的反馈与合作行为,该框架在稳定性与泛化性能上表现出优势,即便主方与智能体均无需了解样本分布或彼此数据质量。
原文摘要 · Abstract (English)
In this technical note, we consider a collaborative learning framework with principal-agent setting, in which the principal at each time-step determines a set of appropriate aggregation coefficients based on how the current parameter estimates from a group of $K$ agents effectively performed in connection with a separate test dataset, which is not part of the agents' training model datasets. Whereas, the agents, who act together as a team, then update their parameter estimates using a discrete-time version of Langevin dynamics with mean-field-like interaction term, but guided by their respective different training model datasets. Here, we propose a decision-theoretic framework that explicitly describes how the principal progressively determines a set of nonnegative and sum to one aggregation coefficients used by the agents in their mean-field-like interaction term, that eventually leading them to reach a consensus optimal parameter estimate. Interestingly, due to the inherent feedbacks and cooperative behavior among the agents, the proposed framework offers some advantages in terms of stability and generalization, despite that both the principal and the agents do not necessarily need to have any knowledge of the sample distributions or the quality of each others' datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。