arXiv:2506.10095cs.CL2025-06被引 3

同一语义的提示词,模型回应却不同,研究发现模型行为漂移有规律。

When Meaning Stays the Same, but Models Drift: Evaluating Quality of Service under Token-Level Behavioral Instability in LLMs

  • 设计新框架PBSS,检测同义提示下的模型行为差异
  • 十项任务中模型响应出现可复现的稳定偏移
  • 揭示分词与解码机制影响模型服务稳定性

我们研究大语言模型对仅在词元层面不同但语义相同的提示的响应,这一现象称为提示变异性。提出基于提示的语义偏移(PBSS)诊断框架,用于衡量模型在语义等价提示重述下的行为漂移。应用于十项约束任务,结果揭示出一致且模型特有的响应偏移,表明其与分词和解码存在统计关联。这些发现突显了提示重述下模型评估稳定性被忽视的维度,并暗示分词策略与解码动态可能造成训练后服务质量的不稳定性。

原文摘要 · Abstract (English)

We investigate how large language models respond to prompts that differ only in their token-level realization but preserve the same semantic intent, a phenomenon we call prompt variance. We propose Prompt-Based Semantic Shift (PBSS), a diagnostic framework for measuring behavioral drift in LLMs under semantically equivalent prompt rewordings. Applied to ten constrained tasks, PBSS reveals consistent, model-specific response shifts, suggesting statistical regularities linked to tokenization and decoding. These results highlight an overlooked dimension of model evaluation stability under rephrasing and suggest that tokenization strategies and decoding dynamics may contribute to post-training quality of service instability.

大模型评估行为漂移提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。