研究大模型回答用户安全问题的表现与改进方向
Learned, Lagged, LLM-splained: LLM Responses to End User Security Questions
- 在900个系统收集的安全问题上评估3个主流大模型
- 模型普遍存在过时、错误及沟通不直接的问题
- 适合安全知识普及者和普通用户参考交互策略
回答终端用户安全问题具有挑战性。尽管GPT、LLAMA、Gemini等大型语言模型(LLMs)尚未完全可靠,但在安全领域外已展现出解答各类问题的潜力。本文通过定性评估3个流行大模型在900个系统收集的终端用户安全问题上的表现,发现它们虽具备广泛的通用安全知识,但仍存在答案过时、不准确,以及间接或不回应的沟通风格等问题,显著影响信息质量。基于这些模式,本文提出模型优化方向,并为用户提供与大模型互动以获取安全帮助的实用策略。
原文摘要 · Abstract (English)
Answering end user security questions is challenging. While large language models (LLMs) like GPT, LLAMA, and Gemini are far from error-free, they have shown promise in answering a variety of questions outside of security. We studied LLM performance in the area of end user security by qualitatively evaluating 3 popular LLMs on 900 systematically collected end user security questions. While LLMs demonstrate broad generalist ``knowledge'' of end user security information, there are patterns of errors and limitations across LLMs consisting of stale and inaccurate answers, and indirect or unresponsive communication styles, all of which impacts the quality of information received. Based on these patterns, we suggest directions for model improvement and recommend user strategies for interacting with LLMs when seeking assistance with security.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。