arXiv:2604.00904cs.LG2026-04被引 2

考虑人类疲劳的智能决策系统,能动态调整人机协作策略。

Fatigue-Aware Learning to Defer via Constrained Optimisation

  • 用疲劳曲线建模人类表现随工作量变化,优化人机协同决策。
  • 在多数据集上优于现有方法,覆盖度0到1之间时效果更优。
  • 可零样本适配不同疲劳模式的专家,适合实际人机协作场景。

学习性延迟(L2D)通过判断AI何时自主决策或交由人工处理,实现人机协作。现有方法假设人类表现恒定,与疲劳导致能力下降的心理学发现相悖。本文提出基于约束优化的疲劳感知学习性延迟(FALCON),利用心理学支持的疲劳曲线显式建模随工作量变化的人类性能。FALCON将L2D建模为包含任务特征与累积工作量的状态的约束马尔可夫决策过程(CMDP),并通过PPO-拉格朗日训练,在人机协作预算下优化准确率。我们进一步构建了FA-L2D基准,系统性地改变疲劳动态,从近似静态到快速退化。跨多个数据集的实验表明,FALCON在不同覆盖水平下持续优于前沿L2D方法,零样本泛化至具有不同疲劳模式的未见专家,并在覆盖度严格介于0与1之间时展现出自适应人机协作相对于纯AI或纯人工决策的优势。

原文摘要 · Abstract (English)

Learning to defer (L2D) enables human-AI cooperation by deciding when an AI system should act autonomously or defer to a human expert. Existing L2D methods, however, assume static human performance, contradicting well-established findings on fatigue-induced degradation. We propose Fatigue-Aware Learning to Defer via Constrained Optimisation (FALCON), which explicitly models workload-varying human performance using psychologically grounded fatigue curves. FALCON formulates L2D as a Constrained Markov Decision Process (CMDP) whose state includes both task features and cumulative human workload, and optimises accuracy under human-AI cooperation budgets via PPO-Lagrangian training. We further introduce FA-L2D, a benchmark that systematically varies fatigue dynamics from near-static to rapidly degrading regimes. Experiments across multiple datasets show that FALCON consistently outperforms state-of-the-art L2D methods across coverage levels, generalises zero-shot to unseen experts with different fatigue patterns, and demonstrates the advantage of adaptive human-AI collaboration over AI-only or human-only decision-making when coverage lies strictly between 0 and 1.

人机协作决策优化疲劳建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。