arXiv:2511.01415cs.AI2025-11中稿 · NeurIPS

AI智能体在双任务中时间判断偏长,揭示了类人时间处理机制。

Modulation of temporal decision-making in a deep reinforcement learning agent under the dual-task paradigm

  • 用深度强化学习构建双任务模型,模拟人类时间判断行为。
  • 双任务下时间估计比单任务平均多出30%以上,四种时长均成立。
  • 未发现专用计时器,提示时间感知可能源于通用神经机制。

本研究从人工智能视角探讨双任务范式中的时间处理干扰问题。实验采用简化版的Overcooked环境,设置单任务(T)和双任务(T+N)两种条件,两者均包含嵌入式时间生成任务,而双任务额外加入数字比较任务。分别训练两个深度强化学习(DRL)智能体应对不同任务。结果显示,双任务智能体的时间估计显著高于单任务智能体,该现象在四个目标时长下均一致。对智能体LSTM层的初步神经动力学分析未发现明确的专用或内在计时机制。因此需进一步研究其时间保持机制,以理解行为模式成因。本研究为探索人工智能与生物系统间行为相似性迈出一步,有助于深化对两类系统的认知。

原文摘要 · Abstract (English)

This study explores the interference in temporal processing within a dual-task paradigm from an artificial intelligence (AI) perspective. In this context, the dual-task setup is implemented as a simplified version of the Overcooked environment with two variations, single task (T) and dual task (T+N). Both variations involve an embedded time production task, but the dual task (T+N) additionally involves a concurrent number comparison task. Two deep reinforcement learning (DRL) agents were separately trained for each of these tasks. These agents exhibited emergent behavior consistent with human timing research. Specifically, the dual task (T+N) agent exhibited significant overproduction of time relative to its single task (T) counterpart. This result was consistent across four target durations. Preliminary analysis of neural dynamics in the agents' LSTM layers did not reveal any clear evidence of a dedicated or intrinsic timer. Hence, further investigation is needed to better understand the underlying time-keeping mechanisms of the agents and to provide insights into the observed behavioral patterns. This study is a small step towards exploring parallels between emergent DRL behavior and behavior observed in biological systems in order to facilitate a better understanding of both.

强化学习时间判断类脑机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。