arXiv:2606.08800cs.AI2026-06

让AI自动提取符合专家标准的可解释特征,提升关键领域模型可信度。

Bridging Expert Knowledge and Automated Feature Engineering via Self-Evolution

论文配图:Bridging Expert Knowledge and Automated Feature Engineering via Self-Evolution
图 1 · 摘自论文原文
  • 用双流生成+树状迭代演化,从原始文本图像中自动提炼可审计特征
  • 在20个任务中17个领先,平均准确率比基线高4.2个百分点
  • 能将'保持专业语气'等定性要求转为可操作特征,适合需要人工审核的场景

在品牌合规、临床护理和内容审核等高风险场景中,机器学习不能作为黑箱使用:从业者需审查驱动模型决策的特征,且模型必须遵循领域专家文档。实际数据常为非结构化内容,所提取特征须具备可解释性、判别力,并与专家关注点一致。现有方法存在局限:仅针对表格输入、缺乏专家对齐验证,无法将如‘保持专业语气’等定性标准转化为精确特征。本文提出FEST(自演化树引导的特征工程),结合双流特征生成(语义与确定性)、语义去重及树状引导的迭代演化,从原始文本和图像中发现可审计特征。FEST在20个分类任务组合中于17个取得领先,五种分类器平均性能提升4.2个百分点。以LLM为裁判的评估显示,其在严格语义对齐阈值下覆盖了60-80%专家设计的品牌特征,人类专家评估也证实特征具有高度相关性、清晰性和可操作性。当引入专家指南时,FEST可将定性标准转化为可执行特征,使各品牌平均准确率提升6-12个百分点。为系统评估自动化特征工程中的专家对齐性,我们发布BrandGuide——首个将专家设计特征与超过100万资产、2,683个品牌配对的数据集。通过将特征工程扎根于专家知识,FEST为需要人工监督的领域提供了可解释机器学习的实用路径。

原文摘要 · Abstract (English)

In high-stakes settings such as brand compliance, clinical care, and content moderation, machine learning cannot be deployed as opaque oracles: practitioners inspect the features driving model decisions, and models must leverage the expert documentation governing these domains. In practice, the data arrives as unstructured content, and features extracted from it must be interpretable, discriminative, and aligned with what experts consider important. Existing methods fall short: they target tabular inputs, lack demonstrated expert alignment, and cannot operationalize qualitative criteria such as 'maintain professional tone' into precise features. We present FEST (Feature Engineering with Self-evolving Trees), combining dual-stream feature generation (semantic and deterministic), semantic deduplication, and tree-guided iterative evolution to discover auditable features from raw text and images. FEST leads in 17 of 20 classifier-task combinations across brand classification, content authenticity detection, and stress detection, with a mean gain of 4.2 pp over the strongest baseline across five classifiers. An LLM-as-judge evaluation shows FEST achieves 60-80% coverage of expert-designed brand features at strict semantic-alignment thresholds, corroborated by a human expert study rating features highly on relevance, clarity, and actionability. When seeded with expert guidelines, FEST refines qualitative criteria into operational features, improving accuracy by 6-12 pp on average across brands. To enable systematic evaluation of expert alignment in automated feature engineering, we release BrandGuide, the first dataset pairing expert-designed features with 1M+ assets across 2,683 brands. By grounding feature engineering in expert knowledge, FEST opens a practical pathway for interpretable ML in domains demanding human oversight.

可解释AI特征工程专家知识文本分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。