arXiv:2510.08081cs.AIcs.CL2025-10EMNLP综述被引 4

用大模型自动发现可解释的评价质量特征,提升电商推荐效果。

AutoQual: An LLM Agent for Automated Discovery of Interpretable Features for Review Quality Assessment

  • 通过反思与工具调用,让大模型自主生成并验证评价特征
  • 在亿级用户平台测试,使人均浏览评论数提升0.79%
  • 适合需要可解释性、避免黑箱决策的推荐系统场景

按内在质量对在线评论排序是电商平台和信息服务的关键任务,直接影响用户体验与商业结果。然而质量具有领域依赖性和动态性,评估难度大。传统手工特征方法难以跨领域扩展且无法适应内容变化;现代深度学习方法常产生不可解释的黑箱模型,可能过度关注语义而非真实质量。为此,我们提出AutoQual,一种基于大模型的智能体框架,自动化发现可解释的特征。该框架模拟人类研究过程,通过迭代生成特征假设、自主实现工具化验证,并将经验存入持久记忆。我们在拥有十亿级用户的大型平台上部署此方法。大规模A/B测试表明,其有效提升了平均每位用户浏览评论数0.79%,并将评论读者的转化率提高0.27%。

原文摘要 · Abstract (English)

Ranking online reviews by their intrinsic quality is a critical task for e-commerce platforms and information services, impacting user experience and business outcomes. However, quality is a domain-dependent and dynamic concept, making its assessment a formidable challenge. Traditional methods relying on hand-crafted features are unscalable across domains and fail to adapt to evolving content patterns, while modern deep learning approaches often produce black-box models that lack interpretability and may prioritize semantics over quality. To address these challenges, we propose AutoQual, an LLM-based agent framework that automates the discovery of interpretable features. While demonstrated on review quality assessment, AutoQual is designed as a general framework for transforming tacit knowledge embedded in data into explicit, computable features. It mimics a human research process, iteratively generating feature hypotheses through reflection, operationalizing them via autonomous tool implementation, and accumulating experience in a persistent memory. We deploy our method on a large-scale online platform with a billion-level user base. Large-scale A/B testing confirms its effectiveness, increasing average reviews viewed per user by 0.79% and the conversion rate of review readers by 0.27%.

大模型应用可解释性推荐系统自动特征发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。