量子世界模型可完美对齐真实与虚拟世界的最优策略,经典模型则不行。
An Irreducible Quantum Advantage in Aligning World Models with Reality

- 用量子系统建模,仅需一个qutrit即可精确模拟复杂环境
- 经典模型在特定路径上会误判动作优劣或持续高估次优动作
- 适合研究量子优势、强化学习对齐问题的学者参考
世界模型为真实世界提供数字模拟,使智能体可在低成本虚拟环境中训练与测试。每个时间步接收动作后生成匹配真实世界统计特性的观测与奖励。在依赖远期历史的复杂环境中,需具备记忆能力。尽管真实世界本身为经典系统,我们证明:即使增加记忆容量,经典世界模型也无法始终准确对齐真实与虚拟世界中的最优策略。我们构造了若干真实世界,其中所有有限经典模型在相同轨迹上均失效:要么无法区分真实世界偏好的动作,要么反复将最高期望奖励赋予次优动作,且其期望奖励估计始终存在不可消除的平均误差。相比之下,每个此类真实世界均可由单个qutrit构成的量子世界模型精确复现,其奖励估计与偏好始终与真实世界一致,确保真实与虚拟世界的最优策略完全对齐。
原文摘要 · Abstract (English)
World models provide digital simulacra of the true world, allowing agents to be trained and tested before costly real-world deployment. At each time step, they receive an action and generate an observation and reward matching the statistics of the true world. In complex environments where present outcomes depend on events far in the past, this requires memory. One might expect that, by increasing memory, we can always build a model accurately enough to align the optimal agent policies of the real and virtual worlds. We show that this is false for classical world models, even when the true world itself is classical. We construct true worlds for which every finite classical model fails along the same possible trajectory: it either loses the ability to distinguish actions when the true world clearly prefers one, or repeatedly assigns the highest expected reward to suboptimal actions. Its expected-reward estimates also retain a nonvanishing average error. In contrast, each such true world admits a quantum world model using a single qutrit that reproduces it exactly: its reward estimates and preferred actions always match those of the true world, ensuring that the optimal policies of the real and virtual worlds remain perfectly aligned.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。