研究航运公司如何为RAG系统定义检索需求,解决用户对AI完美输出的期待与实际准确性之间的矛盾。
Towards Requirements Engineering for RAG Systems
- 通过与用户迭代实验,识别特定场景下的检索需求以判断生成内容正确性
- 数据科学家需在真实业务场景中持续调整,应对系统能力限制
- 适合关注领域专用AI系统落地的工程师与产品经理参考
本文通过一家航运服务提供商的案例研究,探讨大型语言模型(LLM)在专家场景中开发与集成时的需求工程问题。聚焦于检索增强生成(RAG)系统的构建过程,揭示数据科学家面临的核心矛盾:用户对AI输出完美的期望与实际生成结果的准确性之间存在张力。研究发现,只有用户能判断输出正确性,因此数据科学家必须通过与用户反复协作,逐步识别出上下文相关的“检索需求”。本文提出一个实证性过程模型,描述了数据科学家如何在实践中获取这些需求并管理系统局限。该工作深化了软件工程对复杂领域专用RAG系统需求工程的理解。
原文摘要 · Abstract (English)
This short paper explores how a maritime company develops and integrates large-language models (LLM). Specifically by looking at the requirements engineering for Retrieval Augmented Generation (RAG) systems in expert settings. Through a case study at a maritime service provider, we demonstrate how data scientists face a fundamental tension between user expectations of AI perfection and the correctness of the generated outputs. Our findings reveal that data scientists must identify context-specific "retrieval requirements" through iterative experimentation together with users because they are the ones who can determine correctness. We present an empirical process model describing how data scientists practically elicited these "retrieval requirements" and managed system limitations. This work advances software engineering knowledge by providing insights into the specialized requirements engineering processes for implementing RAG systems in complex domain-specific applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。