arXiv:2412.06303cs.LGcs.AI2024-12

DSAI能无偏地提取数据真实特征,让大模型真正读懂数据本身。

DSAI: Unbiased and Interpretable Latent Feature Extraction for Data-Centric AI

  • 通过多阶段流程+可量化指标,从数据中挖掘客观特征
  • 在合成数据上识别出专家定义特征的召回率很高
  • 适合需要解释性分析的现实数据场景

大型语言模型(LLMs)常因依赖预训练知识而难以客观识别大数据集中的潜在特征。为解决这一数据根基不足的问题,我们提出数据科学家人工智能(DSAI)框架,通过多阶段流水线与可量化的显著性度量,实现无偏且可解释的特征提取。在具有已知真值特征的合成数据集上,DSAI展现出对专家定义特征的高召回率,并准确反映数据内在结构。在真实数据集上的应用表明,该框架能在极少专家干预下发现有意义的模式,支持可解释分类等实际场景。论文标题亦由DSAI根据生成标准选定。

原文摘要 · Abstract (English)

Large language models (LLMs) often struggle to objectively identify latent characteristics in large datasets due to their reliance on pre-trained knowledge rather than actual data patterns. To address this data grounding issue, we propose Data Scientist AI (DSAI), a framework that enables unbiased and interpretable feature extraction through a multi-stage pipeline with quantifiable prominence metrics for evaluating extracted features. On synthetic datasets with known ground-truth features, DSAI demonstrates high recall in identifying expert-defined features while faithfully reflecting the underlying data. Applications on real-world datasets illustrate the framework's practical utility in uncovering meaningful patterns with minimal expert oversight, supporting use cases such as interpretable classification. The title of our paper is chosen from multiple candidates based on DSAI-generated criteria.

数据驱动可解释性特征提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。