ML工程代理在公平性上表现不佳,需改进以保障敏感领域应用安全
Be Fair! Can Machine Learning Engineering Agents Adhere to Fairness Constraints?

- 提出以责任为中心的评估框架,关注公平性等关键约束
- 实验显示代理生成模型在肤色公平性上显著低于人工设计基准
- 适合关注ML自动化安全性的研究人员与合规从业者
机器学习工程(MLE)代理有望从原始数据和自然语言指令中自动构建端到端的机器学习流程,使非技术领域的专家也能使用。但在敏感和受监管的领域,这种抽象带来了责任空白:最终用户无法了解影响正确性、鲁棒性、公平性和合规性的设计决策。我们指出现有基准不足以评估MLE代理在这些场景中的安全性。本文提出了责任导向的评估框架要求,并在皮肤癌分类任务中开展探索性研究,重点关注不同肤色调下的公平性。评估两个近期的MLE代理发现,其生成的流水线存在高度变异性,且在预测性能和公平性上均持续逊于人工设计的基线,即使使用了强调公平性的提示。初步结果表明,亟需重新设计MLE代理,以支持人类引导搜索过程,并可靠评估生成流水线的合规性与质量。
原文摘要 · Abstract (English)
Machine learning engineering (MLE) agents promise to automate end-to-end ML pipeline development from raw data and natural language instructions, potentially making ML accessible to non-technical domain experts. However, in sensitive and regulated domains, this abstraction creates a responsibility gap: end-users may lack visibility into design choices that affect correctness, robustness, fairness, and regulatory compliance. We argue that existing benchmarks are insufficient to assess whether MLE agents can be safely applied in such settings. We propose desiderata for a responsibility-centered evaluation framework and conduct an exploratory study on melanoma classification, focusing on fairness across skin tones as a responsibility constraint. When evaluating two recent MLE agents, we find that agent-generated pipelines show high variance and consistently underperform manually designed baselines in both predictive quality and fairness, despite fairness-oriented prompts. These preliminary results suggest that further research is needed towards redesigning MLE agents to allow humans to guide the search process and reliably assess the compliance and quality of the generated ML pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。