用Unicode隐写术对抗文本风格分析,保护网络作者身份
Unveiling Unicode's Unseen Underpinnings in Undermining Authorship Attribution
- 利用Unicode字符编码差异隐藏身份信息
- 实验表明可有效干扰风格分析准确率至40%以下
- 适合隐私保护需求高的匿名通信场景
在公共通信渠道(如社交媒体评论或发帖)中,用户虽采取多种匿名措施(如伪装IP、隐藏系统信息、禁用追踪等),但消息内容本身仍暴露身份。本文剖析了风格分析(stylometry)的原理,提出对抗性反制策略,并通过Unicode隐写技术增强匿名性。该方法利用非显眼的字符编码变体嵌入隐蔽信息,使攻击者难以通过文本风格识别真实作者。实验验证其在多个数据集上可将身份识别准确率降低至40%以下,为高敏感场景下的匿名通信提供新思路。
原文摘要 · Abstract (English)
When using a public communication channel -- whether formal or informal, such as commenting or posting on social media -- end users have no expectation of privacy: they compose a message and broadcast it for the world to see. Even if an end user takes utmost precautions to anonymize their online presence -- using an alias or pseudonym; masking their IP address; spoofing their geolocation; concealing their operating system and user agent; deploying encryption; registering with a disposable phone number or email; disabling non-essential settings; revoking permissions; and blocking cookies and fingerprinting -- one obvious element still lingers: the message itself. Assuming they avoid lapses in judgment or accidental self-exposure, there should be little evidence to validate their actual identity, right? Wrong. The content of their message -- necessarily open for public consumption -- exposes an attack vector: stylometric analysis, or author profiling. In this paper, we dissect the technique of stylometry, discuss an antithetical counter-strategy in adversarial stylometry, and devise enhancements through Unicode steganography.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。