无需环境奖励,仅靠感知信号推断正负反馈,实现在线学习。
Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards

- 用感知包的后续变化推断奖惩,不依赖外部评分。
- 在异或任务中准确判断价值符号,正确率达95.2%。
- 适合无标签、无奖励的强化学习场景,如医疗或安全系统。
我们研究在环境中无标量奖励或评价标签的情况下进行在线奖惩学习。每一步代理仅接收固定通道的感知数据包,疼痛、能量、接触、损伤或认知错误等被视为需通过转移后果推断其正负性的感知维度。OHIRL将角色分离为:M_psi负责预测下一数据包,D_omega建模残差动态,C_eta作为固定内部后转移轨迹评估器,B_xi则学习利用所得价值证据进行策略更新与动作评分。C_eta采用恢复正向、持续/增长负向的残差调节方向;系数原点审计显示,等单位、原始等值和随机单调变体保留超过92%的最优动作排名,而符号反转保留0%。奖励自由协议暴露观察转移,但隐藏环境奖励、延迟外部评估、成功标签和动作优劣标签。条件误差分解将B_xi的价值估计误差与残差策略优化误差分开。在2x2-XOR数据包任务中,药物与辣椒在视觉异或上下文中获得相反价值,相同疼痛或辣度可因后果结构而正可负;B_xi达到0.952的平衡奖惩符号准确率。全在线交错审计中,M_psi holdout R2=0.907,B_xi符号准确率0.940,策略最优动作准确率达0.979;而即时包评分、预测误差奖励、打乱目标、零奖励及误差减少控制均失效。隐含奖励的CartPole与Taxi控制、公开上下文无泄漏审计及模块角色消融进一步测试信息边界与组件必要性。
原文摘要 · Abstract (English)
We study online reward-punishment learning when the environment provides no scalar reward or evaluative label. At each step the agent receives only a fixed-channel perceptual packet, and quantities such as pain, energy, contact, damage, or cognitive error are treated as perceptual dimensions whose valence must be inferred from transition consequences. OHIRL separates four roles: M_psi learns next-packet prediction, D_omega models residual dynamics, C_eta is a fixed internal post-transition trajectory evaluator, and B_xi learns to use the resulting value evidence for later policy updates and action scoring. C_eta uses a recovery-positive and persistence/growth-negative residual-regulation orientation; a coefficient-origin audit shows that equal-unit, raw-equal, and random monotone variants preserve more than 92% of the released top-action rankings, while sign inversion preserves 0%. The reward-free protocol exposes observation transitions while withholding environment rewards, delayed external evaluators, success labels, and action-goodness labels. A conditional error decomposition separates B_xi evidence-estimation error from residual policy-optimization error. In a 2x2-XOR packet task, medicine and chili acquire opposite value under visual XOR contexts, and the same pain or spice increase can be positive or negative depending on consequence structure; B_xi reaches 0.952 balanced reward-sign accuracy. In a full online-interleaved audit, M_psi reaches holdout R2=0.907, B_xi reaches 0.940 sign accuracy, and the policy reaches 0.979 optimal-action accuracy, while immediate packet scores, prediction-error rewards, shuffled targets, zero reward, and error-reduction controls collapse. Hidden-reward CartPole and Taxi controls, public-context no-leakage audits, and module-role ablations further test information boundaries and component necessity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。