arXiv:2608.15095cs.AI2026-08

在观测受限时,智能选择高效低噪的特征表示。

Validation-Frontier Representation Selection under Constrained Observation

  • 提出验证前沿选择器,权衡准确率与特征成本、过拟合、稳定性。
  • 相比全特征,平均减少22.7个特征,前沿得分提升0.0258。
  • 适合资源受限或观测不稳定的机器学习部署场景。

在真实环境中,AI系统常依赖不完整、不稳定、高成本或失效的观测数据。本文研究受限观测下的表示选择问题:当原始准确率非唯一目标时如何选状态表示。提出一种验证前沿选择器,综合平衡准确率,并惩罚特征成本、过拟合差距与验证-测试不稳定性。在基于三个scikit-learn数据集的公开表格基准中,覆盖五种观测模式、45个任务组合、720个候选动作与405个表示行,该自适应选择器在保持平衡准确率相近(无显著差异)的前提下,使前沿得分比全特征基线提升0.025801,同时平均特征数减少22.733。更广泛的离线压力测试结果参差。因此结论有限:在匹配基准下,自适应表示选择可提升受限观测下的鲁棒性-效率前沿,但无法普遍超越全特征基线。

原文摘要 · Abstract (English)

AI systems deployed outside clean benchmark settings often rely on observations that are incomplete, unstable, costly, or degraded by monitoring failures. This paper studies representation selection under constrained observation: choosing a state representation when raw accuracy is not the only operational criterion. We propose a validation-frontier selector that combines balanced accuracy with penalties for feature cost, overfit gap, and validation-test instability. In a focused public-tabular benchmark using three scikit-learn datasets, five observation regimes, 45 matched task cells, 720 candidate actions, and 405 representation rows, the adaptive selector improves frontier score over full trace features by 0.025801 while reducing mean feature count by 22.733. Balanced-accuracy difference is small and not statistically significant. A broader offline stress test gives mixed results. The supported claim is therefore bounded: adaptive representation selection can improve a constrained-observation robustness-efficiency frontier in matched benchmark settings, but does not universally dominate trace baselines.

表示学习特征选择鲁棒性优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。