能通过用户名访问社交媒体并推断用户年龄性别等信息
Web-Browsing LLMs Can Access Social Media Profiles and Infer User Demographics
- 用用户名搜索社交平台,自动分析个人资料
- 在48个模拟账号上准确预测性别与年龄,1384人调查验证有效性
- 可能引发隐私泄露风险,建议限制公开使用
大型语言模型(LLMs)传统上依赖静态训练数据,知识停留在固定快照。近期进展使它们具备网页浏览能力,可实时检索并推理动态网络内容。尽管已有研究证明其能访问和分析网站,但对直接获取与分析社交媒体数据的能力尚无探索。本文通过48个合成的X(原Twitter)账号数据集及1,384名国际参与者的调查数据集,验证了具备网页浏览能力的LLMs能否仅凭用户名推断用户人口统计特征。结果显示,这些模型可成功访问社交媒体内容,并以合理准确度预测用户性别、年龄等属性。对合成数据集的分析表明,模型会解析和解读用户资料,但对活跃度低的账号可能产生性别与政治倾向偏差。该能力虽为后API时代的计算社会科学带来机遇,但也加剧信息操纵与精准广告滥用风险,亟需建立防护机制。建议模型提供商在公开应用中限制此功能,仅对经验证的研究用途开放可控访问。
原文摘要 · Abstract (English)
Large language models (LLMs) have traditionally relied on static training data, limiting their knowledge to fixed snapshots. Recent advancements, however, have equipped LLMs with web browsing capabilities, enabling real time information retrieval and multi step reasoning over live web content. While prior studies have demonstrated LLMs ability to access and analyze websites, their capacity to directly retrieve and analyze social media data remains unexplored. Here, we evaluate whether web browsing LLMs can infer demographic attributes of social media users given only their usernames. Using a synthetic dataset of 48 X (Twitter) accounts and a survey dataset of 1,384 international participants, we show that these models can access social media content and predict user demographics with reasonable accuracy. Analysis of the synthetic dataset further reveals how LLMs parse and interpret social media profiles, which may introduce gender and political biases against accounts with minimal activity. While this capability holds promise for computational social science in the post API era, it also raises risks of misuse particularly in information operations and targeted advertising underscoring the need for safeguards. We recommend that LLM providers restrict this capability in public facing applications, while preserving controlled access for verified research purposes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。