收集了万张真实社交平台上的AI生成图像,揭示其传播现状与可信度问题。
GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment

- 通过多语言文本+平台标识+模型名匹配,从推特抓取并验证10217张图像
- 82%图像含可识别文字,59.2%含人脸,共检测22583个面部特征
- 发现推特云端自动移除内容凭证,导致溯源不可行,适合研究者参考
OpenAI发布GPT-Image-2标志着人工智能生成图像进入新阶段,真实与合成内容的界限愈发模糊。本文推出首个公开的GPT-Image-2 Twitter数据集,基于2026年4月21日模型上线后六天内公开的推文,通过Twitter API v2采集27,662条记录,并采用多阶段清洗流程:涵盖英语、日语、中文的文本启发式筛选、浏览器自动化验证“Made with AI”徽章、以及模型名称变体匹配,最终确认10,217张图像。我们开展四项分析:基于CLIP的零样本主题分类、光学字符识别(文本可读性达82.0%)、人脸检测(59.2%图像含人脸,总计22,583个面部)及语义聚类(137个CLIP ViT-L/14聚类)。关键负面发现是:推特内容分发网络在上传时系统性删除C2PA内容凭证,导致社交媒体来源的AI图像无法进行加密溯源验证。数据集与全部清洗代码已公开。
原文摘要 · Abstract (English)
The release of GPT-image-2 by OpenAI marks a watershed moment in AI-generated imagery: the boundary between photographic reality and synthetic content has never been more difficult to discern. We introduce the GPT-Image-2 Twitter Dataset, the first published dataset of GPT-image-2 generated images, sourced from publicly available Twitter/X posts in the immediate aftermath of the model's April 21, 2026 release. Leveraging the Twitter API v2 and a multi-stage curation pipeline spanning multilingual text heuristics (English, Japanese, and Chinese), browser-automated Twitter "Made with AI" badge verification, and model name variant matching, we curate 10,217 confirmed GPT-image-2 images from 27,662 collected records over a six-day window. We characterize the dataset across four analyses: CLIP-based zero-shot subject taxonomy, OCR text legibility (82.0% of images contain detectable text), face detection (59.2% of images, 22,583 total faces), and semantic clustering (137 CLIP ViT-L/14 clusters). A key negative result is that C2PA content credentials are systematically stripped by Twitter's CDN on upload, rendering cryptographic provenance verification infeasible for social-media-sourced AI images. The dataset and all curation code are released publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。