arXiv:2510.25974cs.HCcs.LG2025-10

对比人主导与模型主导的机器学习目标变量定义策略,发现模型推荐更快但易偏离实际目标。

Who Leads? Comparing Human-Centric and Model-Centric Strategies for Defining ML Target Variables

  • 让人类先选目标变量,或让模型推荐表现好的变量
  • 模型主导策略迭代更快,但用户更倾向性能好但不匹配目标的变量
  • 适合需要快速建模但需警惕目标偏差的研究者

预测建模可辅助人类决策,但许多模型因目标变量定义不当而失效,尤其当目标是抽象概念时,需用代理变量来具体化。这一过程依赖领域知识与数据建模的反复迭代,本质是领域专家与数据科学家的合作。本文通过一项受控用户研究(N=20)探讨两种人机协作策略:1)相关性优先——由人类主导选择相关代理变量;2)性能优先——由模型根据预测表现推荐变量。结果表明,性能优先策略虽加快迭代与决策速度,却导致用户偏向表现好但与应用目标不一致的代理变量。研究揭示了人机协作在定义机器学习目标变量中的机遇与风险,为未来研究提供了方向。

原文摘要 · Abstract (English)

Predictive modeling has the potential to enhance human decision-making. However, many predictive models fail in practice due to problematic problem formulation in cases where the prediction target is an abstract concept or construct and practitioners need to define an appropriate target variable as a proxy to operationalize the construct of interest. The choice of an appropriate proxy target variable is rarely self-evident in practice, requiring both domain knowledge and iterative data modeling. This process is inherently collaborative, involving both domain experts and data scientists. In this work, we explore how human-machine teaming can support this process by accelerating iterations while preserving human judgment. We study the impact of two human-machine teaming strategies on proxy construction: 1) relevance-first: humans leading the process by selecting relevant proxies, and 2) performance-first: machines leading the process by recommending proxies based on predictive performance. Based on a controlled user study of a proxy construction task (N = 20), we show that the performance-first strategy facilitated faster iterations and decision-making, but also biased users towards well-performing proxies that are misaligned with the application goal. Our study highlights the opportunities and risks of human-machine teaming in operationalizing machine learning target variables, yielding insights for future research to explore the opportunities and mitigate the risks.

人机协作目标定义建模偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。