提出DECKER框架,让键盘声攻击在跨设备、跨用户下仍能精准识别击键。
DECKER: Domain-invariant Embedding for Cross-Keyboard Extraction and Recognition

- 通过四阶段设计消除键盘特征干扰,实现跨设备泛化。
- 在37种键盘上识别准确率超基线12.7%,噪声环境下依然有效。
- 适合研究安全漏洞或对抗性攻击的开发者与安全工程师。
键盘声学侧信道攻击(ASCA)可从打字声音中推断输入内容,构成严重安全威胁。现有研究受限于小规模数据集,缺乏用户、键盘和环境多样性。本文构建HEAR数据集,涵盖53名参与者使用37款笔记本键盘,在三种真实场景下采集:外接麦克风、设备内置麦克风(无网络噪声)、基于VoIP的流媒体录音。该数据集支持对键盘泛化、噪声适应和用户偏差的系统评估。在HEAR上,我们建立了一个涵盖传统特征与预训练音频/频谱表示的ASCA基准,包含单模态与多模态设置。提出DECKER框架,分四步:(1) 键盘签名归一化,减少设备色差;(2) 域对抗解耦,抑制键盘身份信息;(3) 监督式跨键盘对比对齐,确保按键一致性;(4) 声学风格随机化,合成未见键盘响应。进一步引入基于LLM的后处理层,利用语言上下文优化击键序列。实验表明,DECKER在跨键盘和跨用户场景下显著优于强基线,准确率提升达12.7%以上,且语言模型修正带来额外增益。结果证明ASCA在多样化用户、设备与噪声环境中仍具高有效性,凸显其现实安全风险。
原文摘要 · Abstract (English)
Acoustic side-channel attacks (ASCA) on keyboards pose a significant security risk, as keystrokes can be inferred from typing acoustics, revealing sensitive information. Prior ASCA studies are limited by small-scale datasets with restricted diversity in users, keyboards, and environments, constraining analysis across devices, microphones, and noise conditions. We introduce HEAR, a dataset designed to study ASCA along three axes: keyboard generalization, noise adaptation, and user bias. HEAR contains recordings from 53 participants using 37 laptop keyboards, collected in three realistic settings: (1) external microphone capture, (2) device microphone capture without network noise, and (3) VoIP-based streaming capture. This enables controlled evaluation across users, keyboards, and environments. On HEAR, we establish an ASCA benchmark spanning conventional features and pre-trained representations from raw audio and spectrograms in unimodal and multimodal settings. We propose DECKER, a domain-invariant keystroke inference framework with four stages: (1) Keyboard Signature Normalization to reduce device coloration, (2) domain-adversarial disentanglement to suppress keyboard identity, (3) supervised cross-keyboard contrastive alignment to enforce key consistency, and (4) Acoustic Style Randomization to synthesize unseen keyboard responses. We further explore sentence-level inference using an LLM-based post-processing layer to refine keystroke sequences via linguistic context. Results on HEAR show DECKER improves keystroke identification over strong baselines, particularly in cross-keyboard and cross-user settings, with further gains from language-model rectification. These findings highlight that ASCA remains effective across diverse users, devices, and noisy environments, underscoring its practical security risk.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。