arXiv:2507.23611cs.CRcs.AI2025-07

用大模型分析截图,自动揪出木马感染路径。

LLM-Based Identification of Infostealer Infection Vectors from Screenshots: The Case of Aurora

  • 用GPT-4o-mini分析受感染截图,提取可疑链接和文件
  • 从1000张截图中发现337个可行动链接和246个相关文件
  • 识别出3个独立传播链,适合威胁情报与应急响应团队

信息窃取类恶意软件会从受感染系统中窃取凭证、会话Cookie及敏感数据。2024年报告的窃取日志超过2900万条,人工分析与大规模应对已几乎不可行。尽管多数研究聚焦于主动检测恶意软件,但对窃取日志及其关联产物(如截图)的被动分析仍存在显著空白。本文提出一种新方法,利用大语言模型(特别是gpt-4o-mini)分析感染时生成的截图,以提取潜在的入侵指标(IoCs)、映射感染向量并追踪攻击活动。针对‘Aurora’窃取木马,我们展示了模型如何从截图中识别恶意链接、安装文件及被利用的软件主题。在1000张截图中,共提取337个可行动链接和246个相关文件,揭示了主要分发方式与社会工程手法。通过关联文件名、链接与感染主题,识别出三个不同传播活动,证明了基于大模型的分析在还原攻击流程与提升威胁情报方面的潜力。该研究将分析范式从传统日志驱动转向反应式、产物驱动,为规模化识别感染路径与实现早期干预提供了可行方案。

原文摘要 · Abstract (English)

Infostealers exfiltrate credentials, session cookies, and sensitive data from infected systems. With over 29 million stealer logs reported in 2024, manual analysis and mitigation at scale are virtually unfeasible/unpractical. While most research focuses on proactive malware detection, a significant gap remains in leveraging reactive analysis of stealer logs and their associated artifacts. Specifically, infection artifacts such as screenshots, image captured at the point of compromise, are largely overlooked by the current literature. This paper introduces a novel approach leveraging Large Language Models (LLMs), more specifically gpt-4o-mini, to analyze infection screenshots to extract potential Indicators of Compromise (IoCs), map infection vectors, and track campaigns. Focusing on the Aurora infostealer, we demonstrate how LLMs can process screenshots to identify infection vectors, such as malicious URLs, installer files, and exploited software themes. Our method extracted 337 actionable URLs and 246 relevant files from 1000 screenshots, revealing key malware distribution methods and social engineering tactics. By correlating extracted filenames, URLs, and infection themes, we identified three distinct malware campaigns, demonstrating the potential of LLM-driven analysis for uncovering infection workflows and enhancing threat intelligence. By shifting malware analysis from traditional log-based detection methods to a reactive, artifact-driven approach that leverages infection screenshots, this research presents a scalable method for identifying infection vectors and enabling early intervention.

恶意软件分析大模型应用威胁情报截图识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。