用逆向对抗攻击提升HMI场景下OCR模型的抗噪能力
Towards Calibration Enhanced Network by Inverse Adversarial Attack
- 设计针对HMI场景的逆向对抗攻击,定位OCR决策边界
- 对抗训练后模型在多种噪声下准确率超基线12.3%以上
- 适合HMI自动化测试、工业视觉验证等高鲁棒性需求场景
随着人机界面(HMI)软件设计与内容复杂度持续上升,测试自动化变得愈发重要。当前主流方法依赖光学字符识别(OCR)从HMI界面自动提取文本信息进行验证,但关键挑战在于如何处理OCR模型对噪声的敏感性。本文提出利用对抗训练技术增强HMI测试场景中的OCR模型性能。具体而言,我们设计了一种面向HMI测试上下文的新型对抗攻击目标,用于发现模型决策边界,并通过对抗训练优化该边界以实现更鲁棒、更准确的OCR模型。此外,我们基于真实需求构建了HMI屏幕数据集,并引入多种扰动类型对原始数据进行增强,以覆盖更多潜在测试场景。实验表明,采用对抗训练后的模型在各类噪声干扰下表现出更强的鲁棒性,同时保持高识别准确率;进一步实验还显示,该模型对其他类型扰动也具备一定抵抗能力。
原文摘要 · Abstract (English)
Test automation has become increasingly important as the complexity of both design and content in Human Machine Interface (HMI) software continues to grow. Current standard practice uses Optical Character Recognition (OCR) techniques to automatically extract textual information from HMI screens for validation. At present, one of the key challenges faced during the automation of HMI screen validation is the noise handling for the OCR models. In this paper, we propose to utilize adversarial training techniques to enhance OCR models in HMI testing scenarios. More specifically, we design a new adversarial attack objective for OCR models to discover the decision boundaries in the context of HMI testing. We then adopt adversarial training to optimize the decision boundaries towards a more robust and accurate OCR model. In addition, we also built an HMI screen dataset based on real-world requirements and applied multiple types of perturbation onto the clean HMI dataset to provide a more complete coverage for the potential scenarios. We conduct experiments to demonstrate how using adversarial training techniques yields more robust OCR models against various kinds of noises, while still maintaining high OCR model accuracy. Further experiments even demonstrate that the adversarial training models exhibit a certain degree of robustness against perturbations from other patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。