arXiv:2608.12346cs.AIcs.CY2026-08中稿 · ICML

AI对齐技术本为防有害输出,却可能被用于信息操控与审查。

Position: The Alignment Community is Unintentionally Building a Censor's Toolkit

  • 将对齐技术映射为潜在的审查工具,揭示其双用途风险。
  • 快速普及的AI使信息控制权更易被滥用,加剧社会不平等。
  • 呼吁社区正视恶意使用风险,制定防范策略。

本文指出,现代AI对齐方法原本旨在防止有害输出,但其本身具有双重用途,可能被恶意行为者用于审查与操纵。通过分析现有对齐技术在实际中的滥用可能性,我们发现追求“完全对齐”的过程中,无意间为攻击者提供了日益精进的信息控制工具。随着用户快速采纳AI作为信息来源、经济权力失衡以及政治环境趋向威权主义,这种风险被进一步放大。本文呼吁研究社区正视对齐机制被故意滥用的潜在威胁,并提出缓解策略以保护公共信息生态。

原文摘要 · Abstract (English)

This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techniques to the possibility and actual cases of misuse, we show that the quest for a "perfectly aligned" model inadvertently also provides malicious actors with an ever-improving tool for informational dominance. We need to discuss this dual-use potential now, as its risk is exacerbated by rapid user adoption of AI as information provider, economic power asymmetries, and a political landscape that increasingly shifts towards authoritarianism. We conclude by urging the community to consider the intentional misuse of AI alignment mechanisms and propose mitigation strategies to safeguard against this dual-use potential.

AI对齐信息审查双用途风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。