arXiv:2510.01288cs.LGcs.AI2025-10被引 1

用微眼动启发的扰动法,无须训练即可发现大模型漏洞。

Microsaccade-Inspired Probing: Positional Encoding Perturbations Reveal LLM Misbehaviours

  • 通过位置编码轻量扰动,探测模型潜在错误
  • 无需微调即可检测事实性、安全、毒性等多类问题
  • 适合研究模型可靠性与安全性的人士

我们受微眼动(microsaccades)——人类感知中揭示隐含动态的微小非自主眼动——启发,提出一种类比的大型语言模型(LLM)探针方法。正如微眼动暴露视觉中的微妙信息变化,我们发现轻量级位置编码扰动能激发模型内部隐藏信号,揭示其错误行为。该方法无需微调或任务特定监督,却能在多种场景下检测到事实性、安全性、毒性及后门攻击等问题。在多个最先进的大模型上实验表明,这种基于扰动的探针能有效暴露模型缺陷,同时保持计算高效。结果表明,预训练大模型已内嵌自我识别失败所需的信息,而受微眼动启发的干预手段为检测和缓解不良行为提供了新路径。

原文摘要 · Abstract (English)

We draw inspiration from microsaccades, tiny involuntary eye movements that reveal hidden dynamics of human perception, to propose an analogous probing method for large language models (LLMs). Just as microsaccades expose subtle but informative shifts in vision, we show that lightweight position encoding perturbations elicit latent signals that indicate model misbehaviour. Our method requires no fine-tuning or task-specific supervision, yet detects failures across diverse settings including factuality, safety, toxicity, and backdoor attacks. Experiments on multiple state-of-the-art LLMs demonstrate that these perturbation-based probes surface misbehaviours while remaining computationally efficient. These findings suggest that pretrained LLMs already encode the internal evidence needed to flag their own failures, and that microsaccade-inspired interventions provide a pathway for detecting and mitigating undesirable behaviours.

大模型探针位置编码模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。