arXiv:2510.12624cs.LGcs.AI2025-10被引 3

让模型自适应选特征,跨任务通用且无需重训。

Learning-To-Measure: In-Context Active Feature Acquisition

  • 基于不确定性引导的贪心选特征策略,提升信息获取效率。
  • 在标签少、缺失多场景下,性能超越单任务方法。
  • 支持直接在历史数据上操作,适合多任务实际应用。

主动特征获取(AFA)是一种序列决策问题,目标是通过自适应选择要获取的特征来提升模型对测试实例的性能。现实中,AFA方法常依赖带有系统性缺失特征和有限任务标签的回顾性数据。现有工作大多针对单一预设任务,难以扩展。为此,我们提出元主动特征获取(meta-AFA)问题,旨在学习跨任务的获取策略。我们提出Learning-to-Measure(L2M),包含:i)对未见任务的可靠不确定性量化;ii)基于不确定性的贪心特征获取代理,最大化条件互信息。我们采用序列建模或自回归预训练方法,实现对任意缺失模式任务的可靠不确定性估计。L2M可直接在具有回顾性缺失的数据集上运行,以“上下文内”方式完成元AFA,无需针对每个任务重新训练。在合成与真实世界表格数据基准上,L2M在标签稀缺和高缺失率条件下表现优于或媲美任务特定基线。

原文摘要 · Abstract (English)

Active feature acquisition (AFA) is a sequential decision-making problem where the goal is to improve model performance for test instances by adaptively selecting which features to acquire. In practice, AFA methods often learn from retrospective data with systematic missingness in the features and limited task-specific labels. Most prior work addresses acquisition for a single predetermined task, limiting scalability. To address this limitation, we formalize the meta-AFA problem, where the goal is to learn acquisition policies across various tasks. We introduce Learning-to-Measure (L2M), which consists of i) reliable uncertainty quantification over unseen tasks, and ii) an uncertainty-guided greedy feature acquisition agent that maximizes conditional mutual information. We demonstrate a sequence-modeling or autoregressive pre-training approach that underpins reliable uncertainty quantification for tasks with arbitrary missingness. L2M operates directly on datasets with retrospective missingness and performs the meta-AFA task in-context, eliminating per-task retraining. Across synthetic and real-world tabular benchmarks, L2M matches or surpasses task-specific baselines, particularly under scarce labels and high missingness.

特征获取元学习缺失数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。