arXiv:2605.00113cs.CLcs.AI2026-05

测试大模型如何响应神经多样性提示,发现需明确指令才有效调整输出结构。

How Frontier LLMs Adapt to Neurodivergence Context: A Measurement Framework for Surface vs. Structural Change in System-Prompted Responses

  • 构建576条输出的NDBench基准,测试大模型在神经多样性情境下的响应变化。
  • 明确指令下模型输出更长更结构化,头部和步骤细节显著增加(p<10^-8)。
  • 仅通过角色声明无法抑制潜在危害,需显式指令才能实现36%-44%的减少。

我们研究前沿对话型大语言模型是否能根据系统提示中的神经多样性(ND)背景调整输出,并描述这种调整的性质。提出NDBench基准,包含两个前沿模型、三种系统提示类型(基础、ND角色声明、带调整指令的ND角色声明)、四种典型ND角色及24个跨四类的提示,其中一类采用对抗性掩码策略。四项趋势一致出现:第一,模型在ND情境下有显著适应性,明确指令条件下输出更长更结构化,表现为更高词元数、更多标题和更细粒度步骤(p < 10^-8,Holm校正);第二,此类适应主要为结构性调整:列表密度变化不大,但标题频率和每步细节明显上升;第三,仅靠角色声明无法抑制潜在有害倾向,掩码强化仅在明确指令下减少36%-44%,角色声明组无明显变化。此外,基于模型的危害评估可靠性分析显示,六维中仅有两项(掩码与强化、验证质量)达到预设评判者一致性标准(alpha ≥ 0.67),可作为主结果。NDBench已公开,含提示、输出、代码及其他资源,形成可复现的未来大模型神经多样性感知审计框架。

原文摘要 · Abstract (English)

We examine if frontier chat-based large language models (LLMs) adjust their outputs based on neurodivergence (ND) context in system prompts and describe the nature of these adjustments. Specifically, we propose NDBench, a 576-output benchmark involving two frontier models, three system prompt types (baseline, ND-profile assertion, and ND-profile assertion with explicit instructions for adjustments), four canonical ND profiles, and 24 prompts across four categories, one of which involves an adversarial masking strategy. Four trends emerge consistently from our findings. First, LLMs show significant adaptation under ND context, where fully instructed conditions yield lengthier and more structured outputs, characterized by higher token counts, more headings, and more granular steps (p < 10^-8, Holm-corrected). Second, such adaptation is largely structural in nature: although list density does not change much, there is a marked rise in the frequency of headings and per-step detail. Third, ND persona assertion alone fails to suppress potentially harmful tendencies, as masking-reinforcement decreases only in explicitly instructed cases (36-44% reduction); the reduction rate barely changes in persona assertion conditions. Moreover, reliability analysis of LLM-based harm assessment reveals that only two out of the six dimensions (masking and reinforcement, validation quality) exceed the pre-defined inter-judge agreement criterion (alpha >= 0.67) and thus can be considered primary results. NDBench is made publicly available along with its prompts, outputs, code, and other resources, forming a reproducible framework for auditing future LLMs' adaptation to ND awareness.

大模型评测神经多样性提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。