arXiv:2603.20508cs.MAcs.AI2026-03中稿 · ICML被引 2

提出衡量推理模型从弱到强可读性的新方法,解决大模型输出难被小模型理解的问题。

Measuring Weak-to-Strong Legibility of Reasoning Models

  • 设计新指标评估弱模型能否读懂强模型的推理过程
  • 发现现有简洁性度量忽略推理完整性,导致误判可读性
  • 适用于安全监控、模型蒸馏等需要跨层级协作的场景

推理语言模型(RLMs)及其生成的中间思维链在多智能体系统中日益关键,如模型间监控或蒸馏至小型模型。当不同能力层级的智能体需协作时,强模型必须生成弱模型可理解的决策路径。我们称此目标为“弱到强可读性”。大模型的可信度部分依赖于这一属性。尤其在安全监督中,使用弱监控器可能成为低成本可靠保障的标准。可读性要求这些决策路径的形式对弱监控器而言是可访问的。现有基于效率的可读性度量未能捕捉“彻底性”,仅关注简洁性。

原文摘要 · Abstract (English)

Reasoning language models (RLMs) and the intermediate chains of thought they emit play an increasingly central role in multi-agent setups such as inter-model monitoring or distillation into smaller models. When agents at different capability tiers must cooperate, strong models need to produce traces digestible by weaker ones. We refer to this goal as "weak-to-strong legibility". Trustworthiness of large models depends in part on this legibility property. For safety oversight in particular, adoption of weak monitors may become a standard for reliability scaffolds on a healthy budget. Legibility requires that the shape of these decision-making traces takes some form accessible to weaker monitors. Existing efficiency-based metrics for legibility fail to capture "thoroughness", instead focusing on conciseness.

推理模型可读性多智能体安全监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。