arXiv:2510.00332cs.AIcs.CE2025-10中稿 · AAAI被引 5

测试AI在高风险金融骗局中的抗骗能力,发现模型易被误导且无法自主决策。

When Hallucination Costs Millions: Benchmarking AI Agents in High-Stakes Adversarial Financial Markets

  • 用加密市场真实事件设计178个任务,考验模型识破操纵与决策能力。
  • 无工具时模型准确率仅28%,即使有工具也难超67.4%,低于人类80%基准。
  • 模型偏好不可靠搜索结果,暴露其选择机制存在根本缺陷。

我们提出CAIA基准,揭示当前AI评估的盲点:先进模型在充满欺骗、后果不可逆的高风险对抗性环境中表现脆弱。以2024年遭攻击损失300亿美元的加密市场为测试场景,评估17个模型在178个时间锚定任务上的表现,要求其辨别真相与操纵、处理碎片化信息、在对抗压力下做出不可逆财务决策。结果显示,缺乏工具时,顶尖模型在初级分析师常规任务中准确率仅为28%;引入工具后性能提升至67.4%,仍远低于人类80%基准。更严重的是,模型系统性地选择不可靠网络搜索,受搜索引擎优化和社交媒体操纵影响,即便正确答案可通过专业工具直接获取。这表明问题非知识缺失,而是根本性决策机制缺陷。此外,Pass@k指标掩盖了自主部署中危险的试错行为。该研究对网络安全、内容审核等对抗性领域具有广泛启示。我们发布带污染控制与持续更新的CAIA基准,确立对抗鲁棒性是可信AI自主性的必要条件。

原文摘要 · Abstract (English)

We present CAIA, a benchmark exposing a critical blind spot in AI evaluation: the inability of state-of-the-art models to operate in adversarial, high-stakes environments where misinformation is weaponized and errors are irreversible. While existing benchmarks measure task completion in controlled settings, real-world deployment demands resilience against active deception. Using crypto markets as a testbed where $30 billion was lost to exploits in 2024, we evaluate 17 models on 178 time-anchored tasks requiring agents to distinguish truth from manipulation, navigate fragmented information landscapes, and make irreversible financial decisions under adversarial pressure. Our results reveal a fundamental capability gap: without tools, even frontier models achieve only 28% accuracy on tasks junior analysts routinely handle. Tool augmentation improves performance but plateaus at 67.4% versus 80% human baseline, despite unlimited access to professional resources. Most critically, we uncover a systematic tool selection catastrophe: models preferentially choose unreliable web search over authoritative data, falling for SEO-optimized misinformation and social media manipulation. This behavior persists even when correct answers are directly accessible through specialized tools, suggesting foundational limitations rather than knowledge gaps. We also find that Pass@k metrics mask dangerous trial-and-error behavior for autonomous deployment. The implications extend beyond crypto to any domain with active adversaries, e.g. cybersecurity, content moderation, etc. We release CAIA with contamination controls and continuous updates, establishing adversarial robustness as a necessary condition for trustworthy AI autonomy. The benchmark reveals that current models, despite impressive reasoning scores, remain fundamentally unprepared for environments where intelligence must survive active opposition.

AI安全金融诈骗对抗测试模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。