arXiv:2602.02005cs.ARcs.LG2026-02被引 1

让FPGA在微秒级内完成学习与推理,实现物理系统实时自适应。

Position: The Need for Ultrafast Training

  • 在FPGA芯片内直接执行训练与推理,延迟低于1微秒。
  • 支持量子纠错、等离子体控制等需实时响应的科学场景。
  • 为高速动态系统提供闭环学习能力,适合高频率科研应用。

领域专用FPGA已在科学与工业工作负载中实现极低延迟推理,但现有加速器大多假设模型为离线静态训练,将学习与适应任务交由较慢的CPU或GPU处理。这种分离限制了在非平稳、高频环境中的系统性能,而这些环境中模型更新必须与底层物理过程同频。本文主张从仅支持推理的加速器转向超快速片上学习,使推理与训练均能在FPGA硬件中以确定性、亚微秒级延迟执行。将学习纳入与推理同一条实时数据通路,可实现随物理过程同步自适应的闭环系统,适用于量子纠错、低温量子比特校准、等离子体与聚变控制、加速器调参及自主科学实验。实现此类系统需协同重构算法、架构与工具链,有望使FPGA从静态推理引擎转变为实时学习机器。

原文摘要 · Abstract (English)

Domain-specialized FPGAs have delivered unprecedented performance for low-latency inference across scientific and industrial workloads, yet nearly all existing accelerators assume static models trained offline, relegating learning and adaptation to slower CPUs or GPUs. This separation fundamentally limits systems that must operate in non-stationary, high-frequency environments, where model updates must occur at the timescale of the underlying physics. In this paper, I argue for a shift from inference-only accelerators to ultrafast on-chip learning, in which both inference and training execute directly within the FPGA fabric under deterministic, sub-microsecond latency constraints. Bringing learning into the same real-time datapath as inference would enable closed-loop systems that adapt as fast as the physical processes they control, with applications spanning quantum error correction, cryogenic qubit calibration, plasma and fusion control, accelerator tuning, and autonomous scientific experiments. Enabling such regimes requires rethinking algorithms, architectures, and toolflows jointly, but promises to transform FPGAs from static inference engines into real-time learning machines.

FPGA实时学习闭环控制超快训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。