arXiv:2602.21231cs.LGcs.AI2026-02被引 1

基于自洽方差动态分配模型数量,提升多模型系统效率与可审计性。

ACAR: Adaptive Complexity Routing for Multi-Model Ensembles with Auditable Decision Traces

  • 用3个样本的自洽方差决定任务分配到单/双/三模型执行
  • 准确率达55.6%,比双模型基线高1.2个百分点,54.2%任务无需全量集成
  • 无学习组件、模型无关,适合需透明决策路径的研究与工业场景

我们提出ACAR(自适应复杂度与归因路由),一个在可审计条件下研究多模型协同的测量框架。ACAR通过N=3个探测样本计算自洽方差(sigma),将任务路由至单模型、双模型或三模型执行模式。系统基于TEAMLLM实现,该平台具备确定性执行、不可变产物和完整决策轨迹。我们在1,510个任务上评估,覆盖MathArena、Reasoning Gym、LiveCodeBench和SuperGPQA四个基准,使用Claude Sonnet 4、GPT-4o和Gemini 2.0 Flash,完成超过7,550次可审计运行。结果表明,基于sigma的路由达到55.6%准确率,优于双模型基线(54.4%),且在54.2%的任务中避免了全量集成。路由机制模型无关,无需学习组件。我们还记录了负面结果:第一,检索增强使准确率下降3.4个百分点,因中位检索相似度仅为0.167,缺乏语义对齐反而引入噪声;第二,当模型一致错误(sigma=0)时,任何下游集成均无法修复,此“一致但错误”模式使准确率上限约低于全集成8个百分点;第三,基于响应相似性或熵的代理归因信号与真实留一法归因相关性弱,表明实际归因需显式反事实计算。本工作揭示了实践中的失效假设,并为未来路由、检索与多模型归因研究提供可验证基线。

原文摘要 · Abstract (English)

We present ACAR (Adaptive Complexity and Attribution Routing), a measurement framework for studying multi-model orchestration under auditable conditions. ACAR uses self-consistency variance (sigma) computed from N=3 probe samples to route tasks across single-model, two-model, and three-model execution modes. The system is implemented on top of TEAMLLM, a deterministic execution substrate with immutable artifacts and complete decision traces. We evaluate ACAR on 1,510 tasks spanning four benchmarks: MathArena, Reasoning Gym, LiveCodeBench, and SuperGPQA, using Claude Sonnet 4, GPT-4o, and Gemini 2.0 Flash, producing more than 7,550 auditable runs. Results show that sigma-based routing achieves 55.6 percent accuracy, exceeding the two-model baseline of 54.4 percent while avoiding full ensembling on 54.2 percent of tasks. The routing mechanism is model-agnostic and requires no learned components. We also document negative results. First, retrieval augmentation reduced accuracy by 3.4 percentage points, as median retrieval similarity was only 0.167, demonstrating that experience injection without semantic alignment introduces noise rather than grounding. Second, when models agree on incorrect answers (sigma equals zero), no downstream ensemble can recover; this agreement-but-wrong failure mode is intrinsic to self-consistency and bounds achievable accuracy at approximately eight percentage points below full ensembling. Third, attribution estimates based on proxy signals such as response similarity and entropy showed weak correlation with ground-truth leave-one-out values, indicating that practical attribution requires explicit counterfactual computation. This work documents which assumptions fail in practice and provides falsifiable baselines for future research on routing, retrieval, and multi-model attribution.

多模型集成自洽性可审计性路由机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。