从行为数据中选出最优学习模型,适用于人类分类任务的上下文老虎机场景。
Model selection for behavioral learning data and applications to contextual bandits
- 基于行为数据设计两种模型选择方法,适应非平稳相关数据。
- 理论证明误差界接近独立同分布情况,具有可靠性保障。
- 在人工与真实人类分类实验中验证,适合心理学与机器学习交叉研究者。
动物或人类的学习过程使行为更适应环境,且高度依赖个体特征,通常仅通过个体行为观测得到。本文提出两种模型选择方法:通用留出法与类似AIC的准则,均适配非平稳相关数据。提供了理论误差界,其性能接近标准独立同分布情形。通过应用于上下文老虎机模型,在合成数据和人类分类实验数据上进行了对比验证,展示了方法的有效性与实用性。
原文摘要 · Abstract (English)
Learning for animals or humans is the process that leads to behaviors better adapted to the environment. This process highly depends on the individual that learns and is usually observed only through the individual's actions. This article presents ways to use this individual behavioral data to find the model that best explains how the individual learns. We propose two model selection methods: a general hold-out procedure and an AIC-type criterion, both adapted to non-stationary dependent data. We provide theoretical error bounds for these methods that are close to those of the standard i.i.d. case. To compare these approaches, we apply them to contextual bandit models and illustrate their use on both synthetic and experimental learning data in a human categorization task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。