arXiv:2604.10834cs.SEcs.AI2026-04

测试发现大模型难以准确识别安全相关代码评论,无法替代人工标注。

LLMs for Qualitative Data Analysis Fail on Security-specificComments in Human Experiments

论文配图:LLMs for Qualitative Data Analysis Fail on Security-specificComments in Human Experiments
图 1 · 摘自论文原文
  • 用详细代码说明提升模型识别能力
  • 仅在部分代码上表现改善,整体仍不可靠
  • 适合研究人机协作或提示工程的学者

对人类实验中自由文本解释进行主题分析可提供重要定性洞察,但需多名领域专家标注,成本高昂。大语言模型(LLMs)看似可替代人工标注。然而,识别安全相关要素(如代码标识符、行号、安全关键词)需更深层上下文理解,可能超出情感分类能力。本文在LiveBench上测试四个顶尖LLM,针对九类安全相关代码,在人类用户对漏洞代码片段的自由文本评论中进行检测,结果与人工标注者使用Cohen's Kappa比较。采用不同提示策略,包括新兴代码、带示例的详细代码手册及冲突案例。结果显示,仅使用详细代码描述时有明显提升,但该提升不具普遍性,且不足以可靠替代人工标注。尚需更多模型与任务的后续研究。

原文摘要 · Abstract (English)

[Background:] Thematic analysis of free-text justifications in human experiments provides significant qualitative insights. Yet, it is costly because reliable annotations require multiple domain experts. Large language models (LLMs) seem ideal candidates to replace human annotators. [Problem:] Coding security-specific aspects (code identifiers mentioned, lines-of-code mentioned, security keywords mentioned) may require deeper contextual understanding than sentiment classification. [Objective:] Explore whether LLMs can act as automated annotators for technical security comments by human subjects. [Method:] We prompt four top-performing LLMs on LiveBench to detect nine security-relevant codes in free-text comments by human subjects analyzing vulnerable code snippets. Outputs are compared to human annotators using Cohen's Kappa (chance-corrected accuracy). We test different prompts mimicking annotation best practices, including emerging codes, detailed codebooks with examples, and conflicting examples. [Negative Results:] We observed marked improvements only when using detailed code descriptions; however, these improvements are not uniform across codes and are insufficient to reliably replace a human annotator. [Limitations:] Additional studies with more LLMs and annotation tasks are needed.

大模型评估安全分析文本标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。