研究智能体通过微调特征获取更好标签的在线学习机制。
Online Learning with Improving Agents: Multiclass, Budgeted Agents and Bandit Learners
- 提出组合维度刻画该模型下的在线可学习性
- 扩展至多分类与带奖赏反馈场景
- 支持成本建模,适合策略优化研究者
我们研究了近期提出的改进型学习模型,其中智能体可通过对其特征值进行小幅调整以获得更优标签。本文系统拓展了已有成果:给出了刻画该模型在线可学习性的组合维度;分析了多分类情形;研究了带奖赏反馈设置下的可学习性;并引入对智能体改进成本的建模。结果为复杂决策环境中的动态学习提供了理论支撑。
原文摘要 · Abstract (English)
We investigate the recently introduced model of learning with improvements, where agents are allowed to make small changes to their feature values to be warranted a more desirable label. We extensively extend previously published results by providing combinatorial dimensions that characterize online learnability in this model, by analyzing the multiclass setup, learnability in a bandit feedback setup, modeling agents' cost for making improvements and more.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。