用时间折扣机制解决长期公平分配中的记忆膨胀问题。
Past-Discounting is Key for Learning Markovian Fairness with Long Horizons
- 引入时间折扣策略,用有限状态跟踪历史分配,避免状态空间无限增长。
- 理论证明时间折扣可保持状态空间有界,而完美记忆方法无法做到。
- 实验显示时间折扣在长时程分配中表现优于完美记忆,适合大规模系统设计。
公平性是多智能体系统中动态资源分配的重要考量。现有方法多将公平性视为一次性问题,忽略时间动态,因而忽视了不平等的累积效应。近期方法通过追踪历史分配实现长期公平,但依赖对所有过往效用的完美记忆,导致状态空间随时间无界增长,影响学习算法的可扩展性和收敛性。受人类公平判断倾向于淡化远期事件的启发,本文提出一种带时间折扣的时序公平框架,实现即时公平与完美记忆之间的合理权衡。核心贡献在于构建了带时间折扣的记忆追踪机制,并理论证明该方法能保证状态空间有界且与时间无关,而完美记忆方法不具备此性质。这一结果使得在任意长时域内可高效学习公平策略。通过形式化框架、实验验证及路径分析,证实了时间折扣的必要性,为构建可扩展的公平资源分配系统提供了清晰方向。
原文摘要 · Abstract (English)
Fairness is an important consideration for dynamic resource allocation in multi-agent systems. Many existing methods treat fairness as a one-shot problem without considering temporal dynamics, which misses the nuances of accumulating inequalities over time. Recent approaches overcome this limitation by tracking allocations over time, assuming perfect recall of all past utilities. While the former neglects long-term equity, the latter introduces a critical challenge: the augmented state space required to track cumulative utilities grows unboundedly with time, hindering the scalability and convergence of learning algorithms. Motivated by behavioral insights that human fairness judgments discount distant events, we introduce a framework for temporal fairness that incorporates past-discounting into the learning problem. This approach offers a principled interpolation between instantaneous and perfect-recall fairness. Our central contribution is a past-discounted framework for memory tracking and a theoretical analysis of fairness memories, showing past-discounting guarantees a bounded, horizon-independent state space, a property that we prove perfect-recall methods lack. This result unlocks the ability to learn fair policies tractably over arbitrarily long horizons. We formalize this framework, demonstrate its necessity with experiments showing that perfect recall fails where past-discounting succeeds, and provide a clear path toward building scalable and equitable resource allocation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。