arXiv:2604.04852cs.CRcs.AI2026-04中稿 · the 12th Intellige…

用结构化提示增强大模型推理可靠性,提升网络安全分析可信度。

Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework

  • 设计四维度16因子框架,约束推理过程防幻觉与漂移。
  • 在小型模型上推理准确率提升40%,跨模型表现稳定。
  • 适合安全敏感场景,可解释性强,无需复杂调参或微调。

链式思维(CoT)提示被用于提升大语言模型的推理能力,但在安全敏感分析任务中其可靠性尚未充分验证,尤其缺乏结构化人类评估。相比之下,提示工程提供轻量、透明且可控的引导方式。本文提出一种结构化提示框架,通过16个因素分属四个核心维度——上下文与范围控制、证据锚定与可追溯性、推理结构与认知控制、安全特异性分析约束——显式干预推理过程,缓解幻觉与推理漂移,增强安全场景下的可解释性。以软件定义网络中的DDoS攻击检测为案例,对比结构化与非结构化提示下多个模型家族的表现。帕累托前沿分析与消融实验表明,小模型推理性能最高提升40%,跨规模保持稳定精度增益;人类评估显示评分者间一致性高(Cohen's k > 0.80),结果证实该框架在可靠且可解释的AI驱动网络安全分析中具有有效性与实用性。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) prompting has been used to enhance the reasoning capability of LLMs. However, its reliability in security-sensitive analytical tasks remains insufficiently examined, particularly under structured human evaluation. Alternative approaches, such as model scaling and fine-tuning can be used to help improve performance. These methods are also often costly, computationally intensive, or difficult to audit. In contrast, prompt engineering provides a lightweight, transparent, and controllable mechanism for guiding LLM reasoning. This study proposes a structured prompt engineering framework designed to strengthen CoT reasoning integrity while improving security threat and attack detection reliability in local LLM deployments. The framework includes 16 factors grouped into four core dimensions: (1) Context and Scope Control, (2) Evidence Grounding and Traceability, (3) Reasoning Structure and Cognitive Control, and (4) Security-Specific Analytical Constraints. Rather than optimizing the wording of the prompt heuristically, the framework introduces explicit reasoning controls to mitigate hallucination and prevent reasoning drift, as well as strengthening interpretability in security-sensitive contexts. Using DDoS attack detection in SDN traffic as a case study, multiple model families were evaluated under structured and unstructured prompting conditions. Pareto frontier analysis and ablation experiments demonstrate consistent reasoning improvements (up to 40% in smaller models) and stable accuracy gains across scales. Human evaluation with strong inter-rater agreement (Cohen's k > 0.80) confirms robustness. The results establish structured prompting as an effective and practical approach for reliable and explainable AI-driven cybersecurity analysis.

提示工程推理增强安全分析可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。