用大模型推理实现工业场景下约束感知的特征选择
LLM-Driven Reasoning for Constraint-Aware Feature Selection in Industrial Systems
- 基于大模型进行分步推理,融合语义与量化信息做特征筛选
- 在三个真实工业场景中提升准确率并减少特征复杂度
- 适合需要可解释性与高效特征工程的生产系统
特征选择是大规模工业机器学习系统中的关键步骤,直接影响模型准确性、效率和可维护性。传统方法依赖标注数据和统计启发式规则,在生产环境中难以应用,因标注数据稀缺且需满足多重运行约束。为此,我们提出模型驱动框架 MoFA,通过结合语义与定量特征信息,进行序列化、基于推理的特征选择。MoFA 将特征定义、重要性得分、相关性及元数据(如特征组或类型)整合为结构化提示,实现可解释、约束感知的特征选择。我们在三个真实工业应用中评估:(1) 真实兴趣与时间价值预测,提升精度同时降低特征组复杂度;(2) 价值模型增强,发现高阶交互项,在线上实验中带来显著参与度提升;(3) 通知行为预测,选出紧凑且高价值的特征子集,兼顾模型精度与推理效率。结果表明,基于大模型推理的特征选择在真实生产系统中具有实用性和有效性。
原文摘要 · Abstract (English)
Feature selection is a crucial step in large-scale industrial machine learning systems, directly affecting model accuracy, efficiency, and maintainability. Traditional feature selection methods rely on labeled data and statistical heuristics, making them difficult to apply in production environments where labeled data are limited and multiple operational constraints must be satisfied. To address this, we propose Model Feature Agent (MoFA), a model-driven framework that performs sequential, reasoning-based feature selection using both semantic and quantitative feature information. MoFA incorporates feature definitions, importance scores, correlations, and metadata (e.g., feature groups or types) into structured prompts and selects features through interpretable, constraint-aware reasoning. We evaluate MoFA in three real-world industrial applications: (1) True Interest and Time-Worthiness Prediction, where it improves accuracy while reducing feature group complexity, (2) Value Model Enhancement, where it discovers high-order interaction terms that yield substantial engagement gains in online experiments, and (3) Notification Behavior Prediction, where it selects compact, high-value feature subsets that improve both model accuracy and inference efficiency. Together, these results demonstrate the practicality and effectiveness of LLM-based reasoning for feature selection in real production systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。