让智能体自主决定采集哪些数据,还能保证结果可信。
Guaranteed Adaptive Modality Acquisition: When the Policy Chooses Its Own Calibration Group
- 用策略自适应选择输入,基于最终采集模式做校准。
- 临床心电图任务中仅用48.8%成本,达成71.2%诊断覆盖率。
- 首次实现策略与校准组联动的可证明可靠性,适合高风险场景。
多模态系统在推理时可能只持有部分输入,其余需付费获取。通过自适应采集,策略决定最终观测输入,因此我们基于该终端输入模式给出保证。传统条件校准假设分组映射独立于校准样本,而策略诱导的分组不满足此条件。本文刻画了何时模式条件保证仍有效,并提出两种有限样本构造:无阈值路由,在终端模式上应用校准;以及同时认证完整策略-模式对,使校准数据可选择部署策略。反例表明,针对固定分组设计的保证无法迁移至策略决定分组的情形。提出的方法称为RouteCert。在具有分阶段、成本有序导联协议的临床心电图任务中,经认证的策略在48.8%预设序数成本下,对71.2%的患者作出响应,与心脏病专家诊断的分歧仅为7.4%,且三个采集阶段均拥有独立证书。在掩码多模态基准测试中,按终端模式逐点认证将最差模式选择性风险控制在0.034(对比全信息参考决策),而联合设计达到0.145,低于0.10上限;在预算匹配的同步比较中,应答率降至0.305。
原文摘要 · Abstract (English)
A multimodal system may begin inference holding only some of its inputs and may acquire the rest at a cost. With adaptive acquisition, the policy determines which inputs are ultimately observed, so we state the guarantee conditional on that terminal input pattern. Conditional calibration normally assumes the grouping map is fixed independently of the calibration sample, which policy-induced grouping does not satisfy. We characterize when pattern-conditional guarantees remain valid and give two finite-sample constructions: threshold-free routing with calibration applied at the terminal pattern, and simultaneous certification of complete policy-pattern pairs, which lets calibration data select the deployed policy. A counterexample shows that a guarantee proved for a calibration-independent grouping map need not transfer once the policy makes the terminal group calibration-dependent. We call the resulting method RouteCert. On a clinical electrocardiogram task with a staged, cost-ordered lead protocol, the certified policy answers 71.2% of held-out patients at an observed 7.4% disagreement with the cardiologist's diagnosis at 48.8% of the prespecified ordinal cost of acquiring every stage, and all three acquisition stages carry their own certificate. On masked multimodal benchmarks, certifying pointwise at each terminal pattern holds observed worst-pattern selective risk, measured against the full-information reference decision rather than the true label, at 0.034 where a pooled design reaches 0.145 against a 0.10 cap, at a comparable answered fraction (0.350 vs 0.342); under the budget-matched simultaneous comparison the answered fraction falls to 0.305.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。