arXiv:2605.00924cs.LGcs.AI2026-05中稿 · ance

提出首个连续可控文本风格迁移框架,可绕过AIGC检测同时保持内容语义。

StyleShield: Exposing the Fragility of AIGC Detectors through Continuous Controllable Style Transfer

论文配图:StyleShield: Exposing the Fragility of AIGC Detectors through Continuous Controllable Style Transfer
图 1 · 摘自论文原文
  • 基于连续嵌入空间的流匹配框架,用冻结的Qwen-7B表示引导风格转移
  • 在中文多领域数据集上实现94.6%逃逸率,对未见检测器仍超99%逃逸
  • 支持平滑控制逃避与语义保留的权衡,适合研究检测系统脆弱性者

AI生成内容(AIGC)检测器被广泛应用于学术诚信审查等高风险场景,但其可靠性存在根本矛盾:随着语言模型在人类写作语料上训练优化,AI与人类写作风格的统计边界将逐渐消失。商业激励进一步扭曲这一局面——检测服务与‘去AI化’工具常同属一个供应链,将内容质量评估变为来源判断。本文提出StyleShield,首个面向条件文本风格迁移的流匹配框架,直接在连续词嵌入空间中运行,采用DiT骨干网络和零初始化的跨注意力适配器,以冻结的Qwen-7B表示作为条件。推理时,借鉴图像生成中的SDEdit范式,仅通过单个参数gamma即可实现逃避与语义保留之间的平滑连续控制。在多领域中文基准上,StyleShield对训练检测器实现94.6%的逃逸率,对三个未见检测器均≥99%逃逸,同时保持0.928的语义相似度。我们进一步提出RateAudit文档级调度算法,证明检测率判别可设为任意值,直接质疑基于分数的评估可靠性。

原文摘要 · Abstract (English)

AI-generated content (AIGC) detectors are increasingly deployed in high-stakes settings such as academic integrity screening, yet their reliability rests on a fundamental paradox: as language models are trained on human-written corpora, the statistical boundary between AI and human writing will inevitably dissolve as models improve. Commercial incentives have further distorted this landscape -- detection services and "de-AIification" tools often operate within the same supply chain, replacing evaluation of content quality with judgment of content origin. We present StyleShield, the first flow matching framework for conditional text style transfer, operating directly in continuous token embedding space via a DiT backbone with zero-initialized cross-attention adapters conditioned on frozen Qwen-7B representations. At inference, we adapt the SDEdit paradigm from image synthesis to text embeddings, with a single parameter gamma providing smooth continuous control over the evasion-preservation trade-off. On a multi-domain Chinese benchmark, StyleShield achieves 94.6% evasion against the training detector and >=99% against three unseen detectors, maintaining 0.928 semantic similarity. We further introduce RateAudit, a document-level scheduling algorithm that demonstrates detection-rate verdicts can be set to arbitrary values, directly questioning the reliability of score-based evaluation.

AIGC检测风格迁移文本安全检测规避

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。