用可编程超表面实现无线端到端语义对齐,降低设备算力负担。
Over-the-Air Semantic Alignment with Stacked Intelligent Metasurfaces
- 通过可堆叠智能超表面在波域直接对齐语义特征空间
- 高信噪比下任务准确率达90%,低信噪比仍保持鲁棒性
- 适合资源受限的AI终端,如物联网设备
语义通信系统旨在使具备人工智能能力的设备间传输任务相关的信息,但当收发端模型异构导致潜在表征不一致时,性能会下降。现有对齐方法通常依赖收发端额外的数字处理,增加设备复杂度。本文首次提出基于堆叠智能超表面(SIM)的无线语义对齐框架,可在波域直接实现潜在空间对齐,显著降低设备侧计算开销。我们将SIM建模为可训练的线性算子,能够模拟监督式线性对齐器和零样本的Parseval帧等效器。为物理实现这些算子,我们设计了一种基于梯度的优化方法,定制超表面传输函数以实现期望的语义映射。在异构视觉变换器(ViT)编码器上的实验表明,SIM可准确复现监督与零样本语义等效器,在高信噪比(SNR)下任务准确率最高达90%,且在低SNR下仍具备强鲁棒性。
原文摘要 · Abstract (English)
Semantic communication systems aim to transmit task-relevant information between devices capable of artificial intelligence, but their performance can degrade when heterogeneous transmitter-receiver models produce misaligned latent representations. Existing semantic alignment methods typically rely on additional digital processing at the transmitter or receiver, increasing overall device complexity. In this work, we introduce the first over-the-air semantic alignment framework based on stacked intelligent metasurfaces (SIM), which enables latent-space alignment directly in the wave domain, reducing substantially the computational burden at the device level. We model SIMs as trainable linear operators capable of emulating both supervised linear aligners and zero-shot Parseval-frame-based equalizers. To realize these operators physically, we develop a gradient-based optimization procedure that tailors the metasurface transfer function to a desired semantic mapping. Experiments with heterogeneous vision transformer (ViT) encoders show that SIMs can accurately reproduce both supervised and zero-shot semantic equalizers, achieving up to 90% task accuracy in regimes with high signal-to-noise ratio (SNR), while maintaining strong robustness even at low SNR values.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。