arXiv:2606.06755cs.CLcs.ET2026-06

研究发现用户输入的简短提示词可作为稳定的身份生物特征。

PromptPrint: Behavioral Biometrics Through Natural Language Prompting in LLMs

论文配图:PromptPrint: Behavioral Biometrics Through Natural Language Prompting in LLMs
图 1 · 摘自论文原文
  • 通过分析1034人2万条提示词,发现词汇选择比语义表达更易识别身份
  • 用户在不同场景下行为不一致,但整体上具有高度独特性
  • 轻微改写提示词不影响识别,但语义改写会显著削弱身份信号

传统作者归属研究关注长篇表达文本,但大语言模型(LLMs)交互通常为简短任务型提示。这引发核心问题:此类提示是否包含稳定、可识别且独特的身份信号?我们提出PromptPrint,系统研究基于提示的身份特征,假设用户的习惯用词、语法与话语模式构成可学习的行为生物特征。基于1,034名用户共20,680条真实提示数据,得出三项关键发现:第一,词汇表征显著优于语义编码,支持“词汇稳定性假说”——身份主要编码于表面用词而非抽象意图;第二,风格特征呈现“独特性-一致性悖论”:用户群体间高度独特,但跨场景行为不一致;第三,对抗分析显示存在明显脆弱性谱系:身份信号对轻微词汇扰动鲁棒,但在语义改写下显著退化。整体结果表明大规模识别性能强,确立提示身份作为可行行为生物特征。本研究为LLM交互中的用户建模提供新视角,对安全与隐私具重要意义。数据与代码将在论文被接受后公开。

原文摘要 · Abstract (English)

Authorship attribution research has traditionally focused on long-form, expressive texts; however, interactions with large language models (LLMs) are typically brief and task-driven prompts. This raises a fundamental question: do such prompts contain a stable, author-identifiable, and distinctive signal? We introduce PromptPrint, a systematic study of prompt-based identity, the hypothesis that a user's habitual vocabulary, syntax, and discourse patterns form a learnable behavioral biometric. Using 20,680 real prompts from 1,034 users, we establish three key findings. First, lexical representations significantly outperform semantic encoders, supporting the "lexical stability hypothesis": identity is primarily encoded in surface-level word choice rather than abstract intent. Second, stylometric features exhibit a "uniqueness-consistency paradox": users are highly distinctive across the population, yet behaviorally inconsistent across contexts. Third, adversarial analysis reveals a clear vulnerability spectrum: identity signals are robust to minor lexical perturbations but degrade substantially under semantic paraphrasing. Overall, our results demonstrate strong identification performance at scale, establishing prompt-based identity as a viable behavioral biometric. This work introduces a new perspective on user modeling in LLM interactions, with important implications for security and privacy. Data and code will be released upon the acceptance of our work.

身份识别提示工程行为生物特征隐私安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。