arXiv:2505.16791cs.LGcs.AI2025-05被引 3

针对多模态数据缺失,提出按队列级主动选样补数据的新方法。

Cohort-Based Active Modality Acquisition

  • 基于样本预测置信度估计补数据收益,指导选哪些样本补。
  • 在15模态数据集上,比随机、熵值等方法更有效提升性能。
  • 适用于大规模生物队列研究,如英国生物银行疾病预测场景。

现实中的多模态机器学习常面临模态缺失且获取成本高的问题,如何在预算限制下优先选择哪些样本进行额外模态采集仍具挑战。现有工作多聚焦单样本或训练阶段的采集策略,而对测试阶段、队列级别的采集研究较少。本文提出一种新的测试阶段队列级模态采集框架CAMA,引入基于插补的采集策略,通过估计补采缺失模态的预期效用,实现高效决策;同时提供上界启发式方法用于基准对比。在最多包含15个模态的数据集上实验表明,该方法相比仅依赖预采集信息、基于熵或随机选择的方法,能更有效地引导额外模态的获取。我们进一步在大型前瞻性队列——英国生物银行(UK Biobank)中验证了该方法在疾病预测中指导蛋白质组学数据采集的实际价值与可扩展性。本工作为资源受限环境下实现高效的模态采集优化提供了有效方案。

原文摘要 · Abstract (English)

Real-world multimodal machine learning often faces missing, costly-to-acquire modalities, raising the problem of which samples to prioritize for additional acquisition under a budget. Prior work mainly studies per-sample or training-time acquisition while test-time, cohort-level acquisition is less explored. We propose Cohort-based Active Modality Acquisition (CAMA), a novel test-time cohort-level modality acquisition setting, and introduce imputation-based acquisition strategies that estimate the expected utility of acquiring a missing modality, along with upper-bound heuristics for benchmarking. Experiments on datasets with up to 15 modalities demonstrate that our proposed imputation-based strategies can more effectively guide the acquisition of an additional modality for selected samples compared with methods relying solely on pre-acquisition information, entropy-based guidance, or random selection. We showcase the real-world relevance and scalability of our method by demonstrating its ability to guide the acquisition of proteomics data for disease prediction in a large prospective cohort, the UK Biobank (UKB). Our work provides an effective approach for optimizing modality acquisition at the cohort level, enabling more effective use of resources in constrained settings.

多模态主动学习生物信息队列研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。