arXiv:2508.14066cs.IRcs.AI2025-08中稿 · presentation at th…被引 15

调研13位从业者,揭示工业界RAG落地的真实场景与挑战

Retrieval-Augmented Generation in Industry: An Interview Study on Use Cases, Requirements, Challenges, and Evaluation

  • 通过访谈13名从业者,梳理RAG在真实业务中的应用方式
  • 发现当前多用于垂直领域问答,系统普遍处于原型阶段
  • 强调数据安全与质量是核心需求,评估仍依赖人工

检索增强生成(RAG)是近年来快速发展的AI技术,通过从外部知识源检索信息来提升大模型输出质量。尽管工业界开始采纳RAG,但其实际应用研究仍严重不足。为此,我们对13位行业实践者进行了半结构化访谈,探索RAG在真实场景中的应用现状。研究揭示了:(1)行业典型用例;(2)系统核心需求;(3)实践中遇到的关键挑战与经验教训;(4)当前的评估方法。主要发现包括:现有应用多集中于特定领域的问答任务,系统大多处于原型阶段;行业最关注数据保护、安全与质量,而伦理、偏见和可扩展性则较少被重视;数据预处理仍是主要难题,且系统评估以人工为主,缺乏自动化手段。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) is a well-established and rapidly evolving field within AI that enhances the outputs of large language models by integrating relevant information retrieved from external knowledge sources. While industry adoption of RAG is now beginning, there is a significant lack of research on its practical application in industrial contexts. To address this gap, we conducted a semistructured interview study with 13 industry practitioners to explore the current state of RAG adoption in real-world settings. Our study investigates how companies apply RAG in practice, providing (1) an overview of industry use cases, (2) a consolidated list of system requirements, (3) key challenges and lessons learned from practical experiences, and (4) an analysis of current industry evaluation methods. Our main findings show that current RAG applications are mostly limited to domain-specific QA tasks, with systems still in prototype stages; industry requirements focus primarily on data protection, security, and quality, while issues such as ethics, bias, and scalability receive less attention; data preprocessing remains a key challenge, and system evaluation is predominantly conducted by humans rather than automated methods.

RAG工业应用访谈研究评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。