arXiv:2510.03761cs.CRcs.AI2025-10被引 2

用大模型分析预印本源码,发现数千处隐私泄露

You Have Been LaTeXpOsEd: A Systematic Analysis of Information Leakage in Preprint Archives Using Large Language Models

  • 四阶段框架结合规则匹配与大模型检测隐藏信息
  • 10万篇论文中发现数以千计的个人隐私和密钥泄露
  • 适合关注科研安全与数据伦理的研究者参考

预印本平台如arXiv加速了科学传播,但也带来未被重视的安全风险。除PDF外,这些平台开放访问原始源码、辅助代码、图表及嵌入注释。若无清理,提交内容可能泄露敏感信息,供攻击者通过开源情报获取。本文首次开展大规模安全审计,分析超过1.2TB的arXiv源数据(来自10万篇论文),提出LaTeXpOsEd框架,融合模式匹配、逻辑过滤、传统抓取技术和大语言模型(LLMs),挖掘非引用文件与LaTeX注释中的隐秘披露。为评估LLM的漏洞检测能力,构建了LLMSec-DB基准,测试25个前沿模型。分析发现数千起个人身份信息(PII)泄露、带地理标签的EXIF文件、公开的Google Drive与Dropbox链接、可编辑的私有SharePoint链接、暴露的GitHub与谷歌凭证及云API密钥,还发现内部沟通记录、学术分歧和会议投稿凭证,对研究者与机构声誉构成严重威胁。我们呼吁科研界与平台立即行动修复隐患。为支持开放科学,发布全部脚本与方法,但保留可能被滥用的敏感发现,符合伦理规范。代码与资料见项目主页 https://github.com/LaTeXpOsEd

原文摘要 · Abstract (English)

The widespread use of preprint repositories such as arXiv has accelerated the communication of scientific results but also introduced overlooked security risks. Beyond PDFs, these platforms provide unrestricted access to original source materials, including LaTeX sources, auxiliary code, figures, and embedded comments. In the absence of sanitization, submissions may disclose sensitive information that adversaries can harvest using open-source intelligence. In this work, we present the first large-scale security audit of preprint archives, analyzing more than 1.2 TB of source data from 100,000 arXiv submissions. We introduce LaTeXpOsEd, a four-stage framework that integrates pattern matching, logical filtering, traditional harvesting techniques, and large language models (LLMs) to uncover hidden disclosures within non-referenced files and LaTeX comments. To evaluate LLMs' secret-detection capabilities, we introduce LLMSec-DB, a benchmark on which we tested 25 state-of-the-art models. Our analysis uncovered thousands of PII leaks, GPS-tagged EXIF files, publicly available Google Drive and Dropbox folders, editable private SharePoint links, exposed GitHub and Google credentials, and cloud API keys. We also uncovered confidential author communications, internal disagreements, and conference submission credentials, exposing information that poses serious reputational risks to both researchers and institutions. We urge the research community and repository operators to take immediate action to close these hidden security gaps. To support open science, we release all scripts and methods from this study but withhold sensitive findings that could be misused, in line with ethical principles. The source code and related material are available at the project website https://github.com/LaTeXpOsEd

安全审计大模型隐私泄露预印本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。