arXiv:2510.01020cs.LGcs.AI2025-10中稿 · AISTATS 2026被引 1

在医疗筛查中,用最少检测次数实现低错误率的在线分类。

The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification

  • 通过保守阈值决定何时检测,联合估计风险参数和特征分布。
  • 错误率低于设定阈值,检测次数仅比理想情况多√T量级。
  • 适合需要安全、高效医疗筛查的场景,如疾病早期诊断。

我们研究二元疾病结果的序列测试问题,其中风险服从未知的逻辑回归模型。每轮决策者可选择支付费用进行检测以获得真实标签,或基于患者特征和历史数据进行预测。目标是在最小化检测成本的同时,确保误分类率低于α的概率不低于1−δ。本文提出一种方法,联合估计逻辑回归参数θ⋆和特征分布,采用保守的逻辑得分阈值决定是否检测。理论证明该方法在高概率下达到目标误差,且所需检测次数仅比已知完整信息的最优策略多˜O(√T)。这是首个针对误差约束下的逻辑回归测试的无遗憾保证,具有直接的医学筛查应用价值。模拟结果验证了理论结论,表明能安全分类患者,并以极少额外检测次数准确估计θ⋆。

原文摘要 · Abstract (English)

We study sequential testing for a binary disease outcome when risk follows an unknown logistic model. At each round, the decision maker may either pay for a test revealing the true label or predict the outcome based on patient features and past data. The goal is to minimize costly tests while ensuring the misclassification rate stays below $α$ with probability at least $1-δ$. We propose a method that jointly estimates the logistic parameter $θ^{\star}$ and the feature distribution, using a conservative threshold on the logistic score to decide when to test. We prove our procedure achieves the target error with high probability and requires only $\widetilde O(\sqrt{T})$ more tests than an oracle with full knowledge. This is the first no-regret guarantee for error-constrained logistic testing, with direct applications to medical screening. Simulations corroborate our theoretical results, showing safe classification of patients and efficient estimation of $θ^{\star}$ with few excess tests.

在线学习医疗筛查逻辑回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。