arXiv:2605.23102stat.MLcs.LG2026-05

用LLM权重提升变量选择鲁棒性,自动过滤无效信息。

LLM Sparsity Prior for Robust Feature Selection

论文配图:LLM Sparsity Prior for Robust Feature Selection
图 1 · 摘自论文原文
  • 将LLM权重融入斯皮克-平板模型先验,通过可解释超参数控制稀疏性和权重集中度。
  • 在低数据条件下预测准确率提升,识别出基线遗漏的临床相关特征。
  • 对提示词变化不敏感,适合医疗等小样本高维数据场景。

大型语言模型(LLMs)为高维变量选择提供了可扩展的领域先验信息获取机制。然而,现有方法如LLM-Lasso对权重质量敏感,当LLM生成的权重不准确时性能显著下降。为此,我们首先提出一个量化LLM生成权重质量的框架,支持在不同权重条件下对LLM引导方法进行严格评估。随后提出LLM稀疏先验(LSP),通过两个可解释的超参数(控制全局稀疏性和权重集中度),将LLM生成的权重整合到斯皮克-平板和斯皮克-平板Lasso模型的先验包含概率中。这些参数的层次化超先验使模型能动态削弱无信息或误导性权重的影响,提升鲁棒性,同时在权重准确时不损失收益。最后,我们设计了合理的提示工程策略,并在一项研究急性肾损伤的私有医学数据集上验证方法。LSP在预测准确率上优于基线,并识别出基线遗漏的临床相关特征,在提示词变化下表现稳定,尤其在低数据场景下效果突出。

原文摘要 · Abstract (English)

Large language models (LLMs) offer a scalable mechanism to elicit domain-informed prior information for high-dimensional variable selection. However, existing methods such as LLM-Lasso are sensitive to weight quality, with performance degrading substantially when LLM-generated weights are inaccurate. To address this challenge, we first introduce a framework for quantifying the quality of LLM-generated weights, enabling rigorous evaluation of LLM-informed methods across varying weight regimes. We then propose the LLM Sparsity Prior (LSP), which integrates LLM-generated weights into the prior inclusion probabilities of Spike-and-Slab and Spike-and-Slab Lasso models via two interpretable hyperparameters governing global sparsity and weight concentration. Hierarchical hyperpriors on these parameters allow the model to dynamically discount uninformative or misleading weights, improving robustness without sacrificing gains when weights are accurate. Finally, we develop principled prompt engineering strategies and validate the method on a private medical dataset studying Acute Kidney Injury. LSP improves prediction accuracy and identifies clinically relevant features missed by the baselines, with robustness to prompt variation and particular effectiveness in low-data regimes.

变量选择稀疏先验低数据医学分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。