arXiv:2608.00084cs.CVphysics.optics2026-08

用视觉图生成可执行的光子元件参数程序,精度超90%。

From Pixels to PCells: A Neurosymbolic Approach to Photonic Component Creation

论文配图:From Pixels to PCells: A Neurosymbolic Approach to Photonic Component Creation
图 1 · 摘自论文原文
  • 多模态智能体通过神经符号系统将图像转为几何语言参数程序。
  • 平均交并比达0.9以上,最高0.974,满足设计约束条件。
  • 适用于光子芯片设计优化,适合研发人员快速迭代原型。

我们提出PixCell,一个神经符号系统,使多模态智能体能将视觉呈现的光子元件转化为小领域特定语言(DSL)中的参数化程序。系统支持确定性视觉验证,使评估成本远低于生成成本。采用多种子采样和迭代修正的模型平均最佳回合交并比(IoU)仅为0.416,而通过PixCell接口与验证器,多模态智能体平均IoU超过0.9,八类元件目标中最高达0.974和0.955,同时满足源合约。实验表明,前沿多模态智能体可可靠理解并生成可执行参数表示。利用这些实时参数,对干涉仪进行跨堆栈研究,重构出满足8.0 nm自由光谱范围目标及原始尺寸约束的程序,适用于220纳米SOI、400纳米SiN和400纳米TFLN堆栈。PixCell成功从论文图中重建分束器,并通过全波仿真实现对称传播与平衡输出。此外,同一可执行验证器提供训练奖励与数据集,用于在无监督演示下训练Qwen3.6-35B-A3B模型(采用LoRA与GRPO)。在八个未参与训练的论文图上,其平均冠军IoU从初始8次尝试的0.422提升至3轮验证引导修订后的0.491。结果建立了一个可控框架,用于度量、重定向与改进视觉到参数化的光子元件设计。

原文摘要 · Abstract (English)

We present PixCell, a neurosymbolic system in which multimodal agents convert a visually presented photonic component into a parametric program over a small domain-specific language (DSL) of geometric primitives. A system enabling deterministic visual verification renders evaluation asymmetrically cheaper than the generation attempt. While models using multi-seed sampling and iterative revision reach a mean best-turn IoU of only 0.416, multimodal agents through PixCell's interface and verifier consistently exceed 0.9 mean IoU, with scores reaching 0.974 and 0.955 across eight component targets while also satisfying source contracts. These results demonstrate that frontier multimodal agents can reliably understand and render executable parametric representations from visual targets. Using these live parameters, cross-stack studies on an interferometer reconstruct primitive programs that satisfy an 8.0 nm free spectral range target and the original footprint constraint on modeled 220-nm SOI, 400-nm SiN, and 400-nm TFLN stacks. PixCell further carries a paper-derived splitter from visual reconstruction through SOI full-wave simulation, producing symmetric propagation and balanced outputs. Finally, the same executable verifier supplies a training reward and dataset used to train a Qwen3.6-35B-A3B model with LoRA and GRPO without supervised demonstrations. On eight training-excluded paper figures, its mean champion IoU rises from 0.422 after eight initial attempts to 0.491 after three verifier-guided revision rounds. These results therefore establish a controlled framework for measuring, retargeting, and improving visual-to-parametric photonic component design.

光子设计神经符号参数生成多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。