arXiv:2510.05864cs.CLcs.CY2025-10

研究大模型在长文本中识别隐含有害内容的能力,发现识别效果受位置和显隐性影响。

On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs

  • 通过控制输入长度、有害比例、危害类型和位置,系统测试模型对有害内容的敏感度。
  • 有害内容占比适中时识别率最高,越长输入识别能力越弱,开头内容更易被识别。
  • 显性有害内容比隐性内容更容易被识别,适用于模型安全评估与改进。

大型语言模型(LLMs)越来越多地处理长输入,但当有害句子稀疏嵌入其中时,其行为仍不明确。我们开展敏感性分析,探究LLMs如何从长输入中提取有害句子。通过组合中性与有害句子构建长输入,系统调节四个因素:输入长度(600–30,000词元)、有害句子比例(0.01–0.50)、危害表现形式(显性与隐性)及有害句位置(开头、中间、结尾),实现可控的压力测试。在有毒、冒犯性和仇恨内容上,针对LLaMA-3.1、Qwen-2.5和Mistral的实验揭示一致规律:敏感性随有害比例变化呈非单调,峰值出现在中等水平;输入越长,敏感性越低;有害句位于开头时更易被优先识别;显性危害比隐性危害更可靠地被识别。该研究为理解模型在长输入中对有害内容的优先级提供了系统视角,凸显了安全应用中的潜在优势与挑战。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly operate on long inputs, yet their behavior when harmful sentences are sparsely embedded within such inputs remains poorly understood. We present a sensitivity analysis that probes how LLMs extract harmful sentences embedded in long inputs. We construct long inputs by combining neutral and harmful sentences, and systematically vary four factors: input length (600--30,000 tokens), the proportion of harmful sentences (0.01--0.50), harm realization (explicit vs. implicit), and the position of harmful sentences within the input (beginning, middle, end), enabling a controlled stress-test evaluation. Experiments across toxic, offensive, and hate content, and across LLaMA-3.1, Qwen-2.5, and Mistral, reveal consistent patterns: sensitivity is non-monotonic with respect to harmful prevalence, peaking at moderate levels; sensitivity degrades as input length increases; harmful sentences placed earlier in the input are more strongly prioritized; and explicit harm is more reliably identified than implicit harm. These findings provide a systematic view of how LLMs prioritize harmful sentences in long input under controlled stress conditions, highlighting both emerging strengths and remaining challenges for safety-related use.

大模型安全有害内容检测长序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。