arXiv:2607.19421cs.ARcs.AI2026-07中稿 · publication in the…

光子芯片上实现低功耗视觉变压器微调,抗噪声且效率极高。

Opto-ViT-v2: Noise-Resilient On-Chip Fine-Tuning for Photonic Near-Sensor Vision Transformer Accelerators

论文配图:Opto-ViT-v2: Noise-Resilient On-Chip Fine-Tuning for Photonic Near-Sensor Vision Transformer Accelerators
图 1 · 摘自论文原文
  • 用低秩分解分离光学权重与电子参数,减少存储和更新开销。
  • 在真实光子噪声下仍保持接近纯软件精度,达100K FPS/W能效。
  • 适合边缘智能设备做实时场景适应,尤其光照/温度变化大的环境。

硅基光子(SiPh)加速器通过微环谐振器(MRR)阵列实现高吞吐、低功耗的视觉变压器(ViT)推理。将该平台扩展至片上微调仍面临挑战:反向传播需大量激活存储、频繁权重写回MRR,并耐受器件级噪声。本文提出Opto-ViT-v2,首个面向近传感器SiPh ViT加速器的参数高效微调(PEFT)框架。采用张量化低秩分解,将预训练光学权重与少量可训练电子因子(如ViT-Base仅8K参数)分离,显著降低激活存储与权重更新开销,支持实际片上训练。进一步设计梯度累积稀疏分类器,通过一次性top-k梯度掩码冻结低重要性权重,使分类器训练成本降低约40%。首次建立系统级光子片上训练噪声模型,涵盖MRR串扰、热漂移与激光幅度噪声在前向与反向传播中的影响。基于200多个制备的MRR器件实测校准,结果显示低秩因子更新比全量微调及传统层间低秩适配在相同噪声条件下更具鲁棒性。在VTAB-1K(19任务)与FGVC少样本基准测试中,Opto-ViT-v2在实测光子噪声下恢复精度达清洁软件精度的0.3%~0.8%,同时实现超过100K FPS/W,为光子边缘视觉系统提供实用的片上领域自适应能力。

原文摘要 · Abstract (English)

Silicon-photonic (SiPh) accelerators have emerged as a promising platform for Vision Transformer (ViT) inference by performing matrix multiplications on microring-resonator (MRR) banks with high throughput and energy efficiency. Extending these platforms to support on-chip fine-tuning remains challenging because backpropagation requires large activation storage, frequent weight write-back to MRRs, and tolerance to device-level noise. We present Opto-ViT-v2, the first framework for parameter-efficient fine-tuning (PEFT) on a near-sensor SiPh ViT accelerator. Our tensorized low-rank decomposition separates pretrained optical weights from a small set of trainable electronic factors (as few as 8K parameters for ViT-Base), greatly reducing activation storage and weight updates while enabling practical on-chip training. We further introduce a gradient-accumulated sparse classifier that freezes low-importance weights through one-shot top-k gradient masking, reducing classifier training cost by about 40 percent. We also develop the first system-level noise model for photonic on-chip training, capturing the effects of MRR crosstalk, thermal drift, and laser amplitude noise during both forward and backward propagation. Calibrated using measurements from more than 200 fabricated MRR devices, the model shows that low-rank factor updates are more robust than full fine-tuning and conventional layer-wise low-rank adaptation under identical noise conditions. Experiments on VTAB-1K (19 tasks) and FGVC few-shot benchmarks demonstrate that Opto-ViT-v2 recovers within 0.3 to 0.8 percent of clean software accuracy under measured photonic noise while achieving more than 100 KFPS/W, enabling practical on-chip domain adaptation for photonic edge vision systems.

光子计算视觉变压器边缘智能低功耗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。