用模拟光计算处理百万级房贷数据,首次验证其在真实任务中的可行性
Analog Optical Inference on Million-Record Mortgage Data

- 构建光计算数字孪生模型,处理584万条房贷数据
- 达94.6%准确率,与XGBoost差距仅3.3个百分点
- 定位三类性能瓶颈,指导未来优化方向
模拟光计算机有望大幅提升机器学习推理效率,但此前尚未在小规模图像基准之外的场景中实现演示。本文将模拟光计算机(AOC)数字孪生模型应用于美国584万条HMDA房贷记录的审批分类任务,并识别出三类准确率损失来源。在原始19个特征下,AOC实现94.6%的平衡准确率,参数量为5,126(其中1,024个为光学参数),相较XGBoost的97.9%低3.3个百分点;当光学核心通道数从16增至48时,性能仅提升0.5个百分点,表明瓶颈在于架构而非硬件。将所有模型限制在127位二进制编码后,各模型准确率均降至89.4%-89.6%,数字模型编码损失8个百分点,而AOC仅损失5个百分点。七项校准后的硬件非理想因素未造成可测量影响。这三类限制层(编码、架构、硬件保真度)明确了性能损耗位置及后续改进重点。
原文摘要 · Abstract (English)
Analog optical computers promise large efficiency gains for machine learning inference, yet no demonstration has moved beyond small-scale image benchmarks. We benchmark the analog optical computer (AOC) digital twin on mortgage approval classification from 5.84 million U.S. HMDA records and separate three sources of accuracy loss. On the original 19 features, the AOC reaches 94.6% balanced accuracy with 5,126 parameters (1,024 optical), compared with 97.9% for XGBoost; the 3.3 percentage-point gap narrows by only 0.5pp when the optical core is widened from 16 to 48 channels, suggesting an architectural rather than hardware limitation. Restricting all models to a shared 127-bit binary encoding drops every model to 89.4--89.6%, with an encoding cost of 8pp for digital models and 5pp for the AOC. Seven calibrated hardware non-idealities impose no measurable penalty. The three resulting layers of limitation (encoding, architecture, hardware fidelity) locate where accuracy is lost and what to improve next.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。