arXiv:2512.14726cs.LGcs.AI2025-12

用量子机制提升决策模型,让旧数据学得更好更准。

Quantum Decision Transformers (QDT): Synergistic Entanglement and Interference for Offline Reinforcement Learning

  • 引入量子启发注意力与多路径网络,捕捉复杂依赖关系
  • 性能提升超2000%,在不同数据质量下都表现优异
  • 双机制协同效果显著,适合研究高效决策架构的学者

离线强化学习可在不与环境交互的情况下从预收集数据中学习策略,但现有决策变压器(DT)架构在长时序信用分配和复杂状态-动作依赖方面表现不佳。我们提出量子决策变压器(QDT),一种融合量子启发计算机制的新架构。该方法包含两个核心组件:具备纠缠操作的量子启发注意力,用于捕捉非局部特征相关性;以及具有多路径处理和可学习干扰的量子前馈网络,实现自适应计算。在连续控制任务上的实验表明,相比标准DT,性能提升超过2000%,且在不同数据质量下均具优越泛化能力。消融实验揭示两大组件间存在强协同效应:单独使用均无法达到竞争力,而联合使用则带来远超各自贡献的显著提升。这说明有效量子启发架构需整体协同设计,而非简单模块堆叠。分析识别出三大计算优势:通过非局部相关性增强信用分配、通过并行处理实现隐式集成行为、通过可学习干扰实现自适应资源分配。这些发现确立了量子启发设计原则为推进序列决策中变压器架构的可行方向,其影响可延伸至更广泛的神经网络架构设计。

原文摘要 · Abstract (English)

Offline reinforcement learning enables policy learning from pre-collected datasets without environment interaction, but existing Decision Transformer (DT) architectures struggle with long-horizon credit assignment and complex state-action dependencies. We introduce the Quantum Decision Transformer (QDT), a novel architecture incorporating quantum-inspired computational mechanisms to address these challenges. Our approach integrates two core components: Quantum-Inspired Attention with entanglement operations that capture non-local feature correlations, and Quantum Feedforward Networks with multi-path processing and learnable interference for adaptive computation. Through comprehensive experiments on continuous control tasks, we demonstrate over 2,000\% performance improvement compared to standard DTs, with superior generalization across varying data qualities. Critically, our ablation studies reveal strong synergistic effects between quantum-inspired components: neither alone achieves competitive performance, yet their combination produces dramatic improvements far exceeding individual contributions. This synergy demonstrates that effective quantum-inspired architecture design requires holistic co-design of interdependent mechanisms rather than modular component adoption. Our analysis identifies three key computational advantages: enhanced credit assignment through non-local correlations, implicit ensemble behavior via parallel processing, and adaptive resource allocation through learnable interference. These findings establish quantum-inspired design principles as a promising direction for advancing transformer architectures in sequential decision-making, with implications extending beyond reinforcement learning to neural architecture design more broadly.

强化学习决策模型量子启发离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。