用大模型挖掘隐藏特征,提升数据稀缺场景的预测效果
Latent Feature Mining for Predictive Model Enhancement with Large Language Models
- 将隐含特征挖掘转化为文本推理任务,利用LLM生成潜在变量
- 在司法与医疗领域验证,新特征显著提升分类器性能
- 适合数据受限或伦理难采的高风险领域研究者使用
预测建模常因数据稀缺与质量差而受阻,尤其在特征与结果关联弱、新增采集受限于伦理或实际困难的领域。传统机器学习难以融入未观测但关键的因素。本文提出FLAME(Faithful Latent Feature Mining for Predictive Model Enhancement)框架,将隐含特征挖掘建模为文本到文本的命题逻辑推理,借助大语言模型(LLMs)为已有特征补充潜在特征,从而增强下游任务的预测能力。该框架具备跨领域泛化性,通过融入各领域的上下文信息,可有效迁移至面临类似数据挑战的其他场景。我们在两个案例中验证:(1)刑事司法系统,数据收集受限且具伦理争议;(2)医疗领域,患者隐私与医学数据复杂性制约全面特征获取。结果表明,推断出的隐含特征与真实标签高度一致,并显著提升下游分类器表现。
原文摘要 · Abstract (English)
Predictive modeling often faces challenges due to limited data availability and quality, especially in domains where collected features are weakly correlated with outcomes and where additional feature collection is constrained by ethical or practical difficulties. Traditional machine learning (ML) models struggle to incorporate unobserved yet critical factors. In this work, we introduce an effective approach to formulate latent feature mining as text-to-text propositional logical reasoning. We propose FLAME (Faithful Latent Feature Mining for Predictive Model Enhancement), a framework that leverages large language models (LLMs) to augment observed features with latent features and enhance the predictive power of ML models in downstream tasks. Our framework is generalizable across various domains with necessary domain-specific adaptation, as it is designed to incorporate contextual information unique to each area, ensuring effective transfer to different areas facing similar data availability challenges. We validate our framework with two case studies: (1) the criminal justice system, a domain characterized by limited and ethically challenging data collection; (2) the healthcare domain, where patient privacy concerns and the complexity of medical data limit comprehensive feature collection. Our results show that inferred latent features align well with ground truth labels and significantly enhance the downstream classifier.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。