arXiv:2606.06871cs.LG2026-06

用可验证证据构建可靠诊断系统,解决无线抓包分析中模型胡编乱造的问题。

Evidence-Grounded Ensemble Diagnosis of 802.11 Packet Captures: A Multi-Stage Pipeline with Deterministic Reliability Scoring

  • 多阶段流水线将抓包转为可验证文本,逐帧追踪协议证据
  • 集成投票结合交叉验证,将关键故障检出率提升至96%
  • 自动生成确定性可信度评分,避免模型自我打分误导

802.11抓包诊断依赖专家知识,效率低且结果不一致。现有大模型方法易虚构不存在的协议事件(尤其在截断数据中),产生未校准的置信度,且因参考答案由模型自身生成导致评估偏差。本文提出PROBE(基于证据集成的协议推理),包含四部分:(i) 帧级可验证的抓包到文本转换;(ii) 多轮、多候选集成,支持跨模型复核与渐进式混淆;(iii) 以‘无故障证据’作为有效证据的判据框架;(iv) 完全确定性的综合可信度评分,基于证据有效性、运行稳定性与跨模型一致性,无需模型自评。在87个企业级Wi-Fi抓包(104对捕获-评审)上测试,单次大模型分析使加权证据F1从专家基准0.871提升至0.912,但35%案例遗漏关键帧。朴素集成投票降至0.842,因多数表决放大保守判断:50%已确认故障被误判为‘无问题’或‘证据不足’。加入证据锚定校正后,F1达0.957,自动接受率96%,最差情况仍高于0.70。模型自报置信度集中在0.95(71%精确为0.95),表明其完全无效。此外,引入基于字段断言匹配的模型无关评估框架,消除模型自产参考答案带来的循环偏差。

原文摘要 · Abstract (English)

Diagnosing 802.11 packet captures requires expert protocol knowledge, is slow, inconsistent across engineers, and unscalable. LLM-based approaches sound plausible but fabricate protocol events absent from captures (especially truncated traces), produce uncalibrated confidence scores, and suffer evaluation bias when golden references are co-produced by the model under test. We introduce PROBE (Protocol Reasoning Over evidence-Based Ensembles), a multi-stage pipeline addressing all three failures. It integrates (i) deterministic PCAP-to-text normalization with frame-level verifiability, (ii) multi-run, multi-candidate ensembles with optional cross-model second opinion and progressive obfuscation, (iii) a verdict-aware evidence framework treating absence of failure evidence as contributing evidence, and (iv) a fully deterministic composite reliability score from evidence validity, run-to-run stability, and cross-model agreement without LLM self-assessment. On 87 enterprise Wi-Fi captures (104 capture-reviewer pairs), single-pass LLM analysis raises weighted evidence F1 from 0.871 (expert baseline) to 0.912 but misses critical frames in 35% of cases. Naive ensemble voting drops below baseline (0.842) as majority voting amplifies conservative verdicts: 50% of confirmed failures are misclassified as 'no issue' or 'insufficient evidence.' Adding evidence-grounded reconciliation achieves 0.957 F1, a 96% auto-accept rate, and a worst-case floor above 0.70. LLM self-reported confidence clusters at 0.95 regardless of difficulty (71% report exactly 0.95), confirming it is uninformative. We also introduce a model-agnostic evaluation framework using per-field assertion matching, eliminating circular bias from model-co-produced golden references.

网络诊断大模型评估可靠性无线通信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。