arXiv:2608.13608cs.AIcs.CR2026-08中稿 · CAMLIS 2025

无需标签数据,用教师模型评估智能学习系统性能。

Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis

论文配图:Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis
图 1 · 摘自论文原文
  • 用更强的教师模型提供稀疏修正,评估学生模型的持续学习效果。
  • 教师相对提升与真实性能提升高度相关,可替代标签基准。
  • 适合在无标签或标签不可靠的安全场景中评估学习系统。

智能持续学习系统将大语言模型与检索或记忆机制结合,通过反馈改进而无需重新训练,在网络安全领域价值日益凸显。然而,其效果传统上依赖有标签基准评估,但实际安全场景中标签稀缺、过时且不具代表性,导致无法判断系统是否有效或比较两个系统优劣。传统的以LLM为裁判的方法信号弱,因裁判能力受限于被评估的代理;且在稀疏、偶然、有偏的标签下,蒸馏方法不可靠。本文提出一种基于缩放假说的端到端评估框架:由更强的教师模型对较小的学生模型提供稀疏采样的修正,通过学生模型随时间向教师收敛的程度来评分。在多种安全任务、模型族和系统设计下,我们证明相对于教师的提升与相对于保留黄金标准的提升高度相关,验证了教师相对提升可作为无标签场景下的真实性能代理。进一步实验显示,同级别模型间的LLM裁判无法提供可用信号。结果表明,当人类提供类似稀疏高精度修正时,教师级模型亦可通过相同学习系统得到改进。

原文摘要 · Abstract (English)

Agentic "Continual Learning Harnesses", systems that pair an LLM with retrieval or memory to improve from feedback without retraining, have shown growing value in cybersecurity. But their value is conventionally measured by gains against labeled benchmarks, an approach that often fails in operational security settings. Benchmark labels are scarce, stale, and unrepresentative, so a practitioner often cannot tell whether a given harness helps at all or which of two is better for their task. Traditional LLM-as-a-judge offers little signal because it is no stronger than the agent it evaluates, and distillation is unreliable on scarce, sporadic, and biased labels. We propose a framework for evaluating learning harnesses end-to-end without a labeled benchmark, grounded in the scaling hypothesis. A stronger teacher model provides sparsely sampled corrections to a smaller student with a continual learning harness. We score a harness by how much its student converges toward the teacher over time. Across security tasks, model families, and harness designs, we show that improvement relative to the teacher correlates with improvement relative to a held-out gold standard, validating teacher-relative lift as a proxy for true harness uplift when labels are absent. We further show that LLM-as-a-judge between similarly powered models yields no usable signal. These results suggest that a teacher-sized model can be improved through the same harness when humans provide the same kind of sparse, high-precision corrections.

持续学习无标签评估智能体缩放假说

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。