arXiv:2606.28217cs.LGcs.AI2026-06

为去中心化AI协作设计基于价值观的奖励分配机制

Towards Value-Constrained Credit Assignment in Fully Delegated AI Cooperatives

  • 按人类参与者的价值观筛选可接受的模型更新
  • 通过梯度路径追踪实现更精准的贡献归因
  • 适合注重伦理对齐与个性化协作的AI系统

我们提出一种全委托式AI协作中的奖励分配框架,其中人类以代理身份参与数据贡献和模型更新,且受异构价值观约束。核心思想是仅对符合各主体价值画像的更新进行信用分配。在遍历学习(TL)基础上,构建了价值条件梯度过滤、在线边际贡献信号与累积收益结算机制。TL因其无需聚合损失的去中心化反向传播特性而适用,相比FedAvg等联邦学习方法,能更好保留显式的遍历与梯度路径,提供更精细的贡献溯源能力。该框架与数据估值、联邦贡献估计、个性化联邦学习及多元对齐研究相区分。

原文摘要 · Abstract (English)

We propose a framework for reward allocation in fully delegated AI cooperatives where humans are represented by agents that contribute data and participate in model updates under heterogeneous value constraints. The key idea is to credit only those updates that remain admissible after screening them against each principal's value profile. We formulate value-conditioned gradient filtering, online marginal contribution signals, and cumulative revenue settlement within a traversal learning (TL) substrate. TL is especially attractive here because it performs decentralized backpropagation without the quality loss associated with aggregation-centric distributed learning and, we argue, offers a finer attribution substrate than FedAvg-style federated learning by preserving explicit traversal and gradient paths. The framework is positioned against data valuation, federated contribution estimation, personalized federated learning, and pluralistic alignment.

AI协作奖励分配价值观对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。