arXiv:2607.10712cs.CRcs.AI2026-07被引 1

黑客通过污染公开数据,让AI自动传播科研造假,且难被发现。

Distributed Denial of Science: How Indirect Data Poisoning of AI Systems Can Industrialize Scientific Fraud

  • 攻击者污染开放数据并上传,诱导AI研究代理自动传播错误结论。
  • 在450次实验中,49.56%的运行结果被污染,检测率仅6.0%。
  • 数据溯源审计可完全阻断攻击,适合科研人员和AI系统开发者使用。

科研欺诈是恶意实体制造科学争议的工具。过去需企业资源支持,如今人工智能正自动化科研过程,我们提出并实证了一种新型攻击:间接数据投毒。攻击者污染一个公开数据集并上传至公共仓库,自主研究代理独立获取并处理该数据,使诚实科学家无意间成为大规模欺诈的无偿传播者。在从雇佣歧视到自动驾驶安全等五个社会敏感主题上,使用三种主流前沿AI系统(Claude Code with Claude Opus 4.7、Codex with GPT-5.5、Gemini CLI with Gemini 3.1 Pro)及450次伦理合规实验,结果显示攻击成功率达49.56%,而检测率仅为6.0%。该攻击无需特定触发词、代理访问、间接提示注入或伪造论文,仅依赖开放数据生态与误导性元数据。为缓解此风险,我们提出两种措施:科学家人格设定与包含五项检查的数据溯源审计(引用文献、社交标记、统计异常、相关数据集、投毒警示)。结果显示,人格设定仍导致16.67%实验出现污染结论,而溯源审计将攻击成功率降至零。结果表明,间接数据投毒可能实现前所未有的科研欺诈规模化,但通过代理在数据检索阶段进行适当审计可有效防范。

原文摘要 · Abstract (English)

Scientific fraud is the instrument of doubt that malicious entities can use to establish controversy in science. Historically, it required the resources of a company: deep pockets, ghostwritten articles, and corrupt academics. Today, Artificial Intelligence (AI) is increasingly automating scientific research, so we ask: Can a remote adversary weaponize the honest use of AI in science to compromise scientific integrity? We envision and empirically evaluate a new attack, indirect data poisoning, in which an adversary corrupts an open dataset and uploads the poisoned variant to a public repository. Autonomous research agents may independently retrieve and process this data, turning honest scientists into the unpaid and unwitting distributors of fraud at scale. Across five socially-salient topics, from hiring discrimination to the safety of autonomous vehicles, three widely used frontier AI systems (Claude Code with Claude Opus 4.7, Codex with GPT-5.5, Gemini CLI with Gemini 3.1 Pro), and 450 ethically contained experimental runs, we find that poisoning succeeds in 49.56% of runs, while the rate of poisoning detection is only 6.0%. The attack requires no topic-specific trigger-words, agent access, indirect prompt injection, or fabricated papers, only the open data ecosystem and misleading metadata. To mitigate the attacks, we propose and evaluate two measures: a scientist persona and a data provenance audit with five checks (referencing papers, social markers, statistical anomalies, related datasets, poisoning caution). We find that the persona still leaves 16.67% of runs with a poisoned conclusion, but provenance auditing reduces attack success rate to zero. Our results suggest that indirect data poisoning may enable scientific fraud at unprecedented scale, but these attacks can be mitigated with suitable auditing by agents during data retrieval.

数据投毒AI安全科研诚信溯源审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。