arXiv:2507.18993cs.IRcs.LG2025-07被引 2

用LLM智能挖掘文本中的多值特征,提升推荐系统效果

Agent0: Leveraging LLM Agents to Discover Multi-value Features from Text for Enhanced Recommendations

  • 多个LLM代理协作自动提取文本关键特征
  • 在10个真实数据集上显著提升推荐准确率
  • 适合需要自动化特征工程的推荐系统研发者

大型语言模型(LLMs)及其基于代理的框架在自动化信息抽取方面取得显著进展,而这一能力对现代推荐系统至关重要。尽管这些多任务框架广泛应用于代码生成,但在以数据为中心的研究中仍鲜有应用。本文提出Agent0,一种基于LLM的代理系统,用于从原始非结构化文本中自动完成信息抽取与特征构建。类别型特征对大规模推荐系统极为重要,但获取成本高昂。Agent0通过协调一组交互式LLM代理,自动识别对后续任务(如模型训练或AutoML流程)最有价值的文本特征。除了特征工程能力外,Agent0还提供一种基于动态反馈环的自动化提示工程调优方法,利用来自专家模型(oracle)的反馈。实验表明,这种闭环方法在自动化特征发现方面既实用又高效,而该环节正是当前推荐系统开发中最具挑战性的阶段之一。

原文摘要 · Abstract (English)

Large language models (LLMs) and their associated agent-based frameworks have significantly advanced automated information extraction, a critical component of modern recommender systems. While these multitask frameworks are widely used in code generation, their application in data-centric research is still largely untapped. This paper presents Agent0, an LLM-driven, agent-based system designed to automate information extraction and feature construction from raw, unstructured text. Categorical features are crucial for large-scale recommender systems but are often expensive to acquire. Agent0 coordinates a group of interacting LLM agents to automatically identify the most valuable text aspects for subsequent tasks (such as models or AutoML pipelines). Beyond its feature engineering capabilities, Agent0 also offers an automated prompt-engineering tuning method that utilizes dynamic feedback loops from an oracle. Our findings demonstrate that this closed-loop methodology is both practical and effective for automated feature discovery, which is recognized as one of the most challenging phases in current recommender system development.

推荐系统LLM代理特征工程自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。