AI助手比人类更擅长客服,新评测工具揭示这一趋势。
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
- 用SelfScore评测系统对比AI与人类处理客服问题的表现。
- 含RAG的AI模型在专业咨询中表现优于无RAG模型和人类。
- 适合关注AI替代人力、评估自动化效果的研究者与企业。
本文提出并验证了SelfScore——一种用于评估大型语言模型(LLM)代理在客服与专业咨询任务中表现的新基准。随着AI在各行业尤其是客户服务领域的深入应用,SelfScore填补了自动化代理与人类工作者对比评估的空白。该基准从问题复杂度与回复帮助性两个维度进行评分,确保评分体系透明简洁。研究构建了自动化的LLM代理以测试SelfScore,并探索了检索增强生成(RAG)在特定领域任务中的优势,结果显示:结合RAG的自动化代理显著优于未使用RAG的代理,且整体表现超越人类对照组。基于此,研究对人工智能可能取代人类岗位表示担忧,尤其是在AI表现突出的领域。最终,SelfScore为理解AI在客服环境中的影响提供了基础工具,并呼吁在自动化转型过程中重视伦理考量。
原文摘要 · Abstract (English)
The current paper presents the development and validation of SelfScore, a novel benchmark designed to assess the performance of automated Large Language Model (LLM) agents on help desk and professional consultation tasks. Given the increasing integration of AI in industries, particularly within customer service, SelfScore fills a crucial gap by enabling the comparison of automated agents and human workers. The benchmark evaluates agents on problem complexity and response helpfulness, ensuring transparency and simplicity in its scoring system. The study also develops automated LLM agents to assess SelfScore and explores the benefits of Retrieval-Augmented Generation (RAG) for domain-specific tasks, demonstrating that automated LLM agents incorporating RAG outperform those without. All automated LLM agents were observed to perform better than the human control group. Given these results, the study raises concerns about the potential displacement of human workers, especially in areas where AI technologies excel. Ultimately, SelfScore provides a foundational tool for understanding the impact of AI in help desk environments while advocating for ethical considerations in the ongoing transition towards automation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。