arXiv:2505.14964cs.LGcs.AI2025-05被引 1

用智能数据筛选提升高风险模型性能,少标也能强。

The Achilles Heel of AI: Fundamentals of Risk-Aware Training Data for High-Consequence Models

  • 基于模型反馈和标签多样性,动态选最优数据。
  • 仅用20%~40%精选数据,罕见事件召回率不降反升。
  • 适合军事、应急等高风险场景的高效模型训练。

在国防、情报和灾害响应等高后果领域,AI系统需在资源受限下检测稀有高影响事件。传统以标注量为主的策略引入冗余与噪声,限制模型泛化能力。本文提出智能缩放(smart-sizing)数据策略,强调标签多样性、模型引导选择与边际效用停机。通过自适应标签优化(ALO)实现预标注筛选、标注者分歧分析与迭代反馈,优先选取显著提升模型性能的标签。实验表明,使用20%至40%的精选数据训练的模型,可达到或超过全数据基线表现,尤其在罕见类别召回率与边缘案例泛化上优势明显。研究还揭示训练与验证集中潜藏的标注错误会扭曲评估结果,凸显嵌入式审计工具与性能感知治理的重要性。智能缩放将标注重构为与任务目标对齐的反馈驱动过程,以更少标签构建更鲁棒模型,支持前沿模型与实战系统的高效开发。

原文摘要 · Abstract (English)

AI systems in high-consequence domains such as defense, intelligence, and disaster response must detect rare, high-impact events while operating under tight resource constraints. Traditional annotation strategies that prioritize label volume over informational value introduce redundancy and noise, limiting model generalization. This paper introduces smart-sizing, a training data strategy that emphasizes label diversity, model-guided selection, and marginal utility-based stopping. We implement this through Adaptive Label Optimization (ALO), combining pre-labeling triage, annotator disagreement analysis, and iterative feedback to prioritize labels that meaningfully improve model performance. Experiments show that models trained on 20 to 40 percent of curated data can match or exceed full-data baselines, particularly in rare-class recall and edge-case generalization. We also demonstrate how latent labeling errors embedded in training and validation sets can distort evaluation, underscoring the need for embedded audit tools and performance-aware governance. Smart-sizing reframes annotation as a feedback-driven process aligned with mission outcomes, enabling more robust models with fewer labels and supporting efficient AI development pipelines for frontier models and operational systems.

高风险AI数据筛选模型效率标注优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。