arXiv:2607.03233cs.CRcs.AI2026-07

用智能体+生成式AI提升开源情报效率,解决幻觉与评估难题

Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions

论文配图:Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions
图 1 · 摘自论文原文
  • 构建11类分类体系,区分智能体AI与普通大模型提示工程
  • 发现20余项研究指出现实中幻觉问题严重但仅1个系统实测验证
  • 适合安全分析、情报研判人员及AI评估研究者参考

公开数字信息的激增使传统人工开源情报(OSINT)分析难以应对现代情报、网络安全与网络调查需求。具备工具调用、多步推理和迭代生成能力的大语言模型(LLMs)与智能体AI系统成为有力解决方案,但评估框架未能跟上其能力进展。本综述系统分析74项研究,提出四项贡献:首先,将智能体AI确立为独立分析类别,构建涵盖LLM基础、智能体架构、检索增强生成(RAG)、知识图谱等11个维度的分类体系;其次,识别出‘幻觉-验证’鸿沟这一全局性问题:尽管超过20项研究指出幻觉是可靠性重大隐患,但仅有1个面向OSINT的RAG系统在非可复现条件下进行过端到端幻觉测量,相关推理与事实修正研究多聚焦通用领域问答而非OSINT场景;第三,映射现有研究至OSINT生命周期,显示对采集与分析支持较强,但在验证、报告、传播与决策支持环节覆盖有限;第四,提出包含评估、基准测试、幻觉量化、对抗鲁棒性、暗网覆盖、多模态智能与治理在内的十点研究议程。结论认为,人类-人工智能共驾模式——即大模型协助信息采集与筛选,分析人员保留验证与决策责任——是近中期最合理的部署架构。

原文摘要 · Abstract (English)

The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for modern intelligence, cybersecurity, and cyber investigation. Large language models (LLMs) and agentic AI systems, capable of tool use, multi-step reasoning, and iterative intelligence generation, have emerged as promising solutions, yet evaluation frameworks have not kept pace with reported capabilities. This survey systematically reviews 74 studies and makes four contributions. First, it establishes agentic AI as a distinct analytical category rather than an extension of LLM prompting, organising the literature through an 11-category taxonomy covering LLM foundations, agentic architectures, retrieval-augmented generation (RAG), knowledge graphs, prompt engineering, domain adaptation, evaluation benchmarks, and risk. Second, it identifies the hallucination-validation gap as a corpus-level finding: although hallucination is recognised as a major reliability concern in over twenty studies, end-to-end hallucination is empirically measured in only one OSINT-specific RAG-based system, non-reproducible conditions, while related reasoning and factual-correction studies evaluate general-domain question answering rather than OSINT. Third, it maps existing research to the OSINT lifecycle, showing strong support for collection and analysis but limited coverage of verification, reporting, dissemination, and decision support. Fourth, it derives a ten-point research agenda addressing evaluation, benchmarking, hallucination measurement, adversarial robustness, dark-web coverage, multimodal intelligence, and governance. It concludes that a human-AI co-pilot model, where LLMs assist collection and triage while analysts retain responsibility for verification and decision-making, represents the most defensible near-term deployment architecture.

智能体AI开源情报幻觉检测安全分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。