数据科学家如何将模糊概念转化为可预测变量
Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks
- 通过拼凑现有数据,创造性地构建目标变量
- 需满足有效性、可预测性等五个核心标准
- 适合关注模型设计过程的研究者与实践者
数据科学家常需处理诸如学生写作‘真实性’或患者‘医疗需求’等模糊概念的预测任务。然而,如何将这些模糊概念转化为具体可操作的目标变量,仍缺乏理解。本研究访谈了教育领域(N=8)和医疗领域(N=7)共15位数据科学家,探讨其目标变量构建过程。研究发现,数据科学家采用‘拼凑式’方法,基于有限数据进行创造性和务实性调整。他们需在有效性、简洁性、可预测性、可迁移性及资源要求五个维度上达成平衡。为此,他们会动态调整问题定义,例如替换不达标的变量,或合并多个结果形成综合目标变量以覆盖更全面的建模目标。研究为未来人机交互(HCI)、计算机支持协同工作(CSCW)及机器学习领域提供了支持目标变量构建的机遇。
原文摘要 · Abstract (English)
Data scientists often formulate predictive modeling tasks involving fuzzy, hard-to-define concepts, such as the "authenticity" of student writing or the "healthcare need" of a patient. Yet the process by which data scientists translate fuzzy concepts into a concrete, proxy target variable remains poorly understood. We interview fifteen data scientists in education (N=8) and healthcare (N=7) to understand how they construct target variables for predictive modeling tasks. Our findings suggest that data scientists construct target variables through a bricolage process, in which they use creative and pragmatic approaches to make do with the limited data at hand. Data scientists attempt to satisfy five major criteria for a target variable through bricolage: validity, simplicity, predictability, portability, and resource requirements. To achieve this, data scientists adaptively apply problem (re)formulation strategies, such as swapping out one candidate target variable for another when the first fails to meet certain criteria (e.g., predictability), or composing multiple outcomes into a single target variable to capture a more holistic set of modeling objectives. Based on our findings, we present opportunities for future HCI, CSCW, and ML research to better support the art and science of target variable construction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。