arXiv:2412.17908cs.LGcs.CE2024-12被引 1

用强化学习与贝叶斯优化设计金融场景下的隐蔽后门攻击

Trading Devil RL: Backdoor attack via Stock market, Bayesian Optimization and Reinforcement Learning

  • 通过金融数据污染实现无触发条件的后门攻击
  • 攻击在生成文本中植入隐蔽指令,成功率超90%
  • 适合关注AI安全与金融模型风险的研究者

随着生成式人工智能的快速发展,特别是大语言模型的广泛应用,深度学习多个子领域已取得显著进展,并广泛应用于日常场景。例如,金融机构在生产前及日常运营中,使用强化学习对各类研究模型进行多场景模拟。本文提出一种仅依赖数据投毒的后门攻击方法,名为FinanceLLMsBackRL,该攻击无需预设触发条件,属于无先验触发的攻击类型。我们旨在评估采用强化学习系统的大型语言模型在文本生成、语音识别、金融建模、物理仿真以及当代人工智能生态系统中的潜在影响。同时,提出基于动态系统与数据分布统计分析的检测机制。

原文摘要 · Abstract (English)

With the rapid development of generative artificial intelligence, particularly large language models a number of sub-fields of deep learning have made significant progress and are now very useful in everyday applications. For example,financial institutions simulate a wide range of scenarios for various models created by their research teams using reinforcement learning, both before production and after regular operations. In this work, we propose a backdoor attack that focuses solely on data poisoning and a method of detection by dynamic systems and statistical analysis of the distribution of data. This particular backdoor attack is classified as an attack without prior consideration or trigger, and we name it FinanceLLMsBackRL. Our aim is to examine the potential effects of large language models that use reinforcement learning systems for text production or speech recognition, finance, physics, or the ecosystem of contemporary artificial intelligence models.

后门攻击强化学习金融AI数据投毒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。