测试大模型大小对人机协作信息检索的影响,发现小模型已足够实用。
Seeking Information with RAG-Assistants: Does Model Size Matter in Human-AI Collaborations?

- 用3B、8B、70B三类模型做RAG助手,对比人机协作效果。
- 无论模型大小,人机协作比纯模型表现更好,且差距显著。
- 用户感知的易用性和满意度在不同模型间差异不大,适合实际办公场景。
大量关于大语言模型(LLM)的研究聚焦于提升基准测试性能,但其在真实世界人机协同工作流程中的评估仍显不足。本文在受职场场景启发的多轮信息检索任务中,评估了基于检索增强生成(RAG)的聊天机器人助手的表现,该任务强调遵守本地法规和敏感数据的安全处理。我们考察了112名参与者在使用3B、8B、70B三种不同规模模型的RAG助手时的表现,并与仅使用LLM或LLM+RAG的基线进行对比。结果表明,无论模型规模如何,人机协作相比模型单方表现均有显著提升,证明混合系统在信息检索场景中的价值。然而,用户对易用性和满意度的感知在不同模型间无明显差异,揭示了模型规模、性能与用户感知之间的微妙权衡。本研究强调,在真实多轮交互中评估AI应用,需同时关注可用性、满意度和准确性,而不能仅依赖基准性能。
原文摘要 · Abstract (English)
Much research on LLMs has focused on increasing benchmark performance. However, the evaluation of such models in real-world collaborative human-AI workflows has stayed behind. This work evaluates a chatbot-style assistant based on Retrieval-Augmented Generation (RAG) in a realistic multi-turn information-seeking scenario inspired by workplace settings where compliance with local legislation and secure handling of sensitive data are often key. Specifically, we examine the performance of humans (N=112) assisted by RAG-assistants compared to LLM-only or LLM+RAG baselines. In this setting, we investigate how underlying model size (3B, 8B, and 70B) shapes the human-AI collaborative dynamic and how it influences perceived usability and satisfaction. Results show that the performance gain of human-AI collaboration over the model-only baselines is significant, irrespective of model size, suggesting that hybrid systems are beneficial in information-seeking scenarios. Interestingly, however, perceived usability and satisfaction among participants showed little difference across model sizes. This demonstrates a nuanced trade-off between model size, performance, and user perception. Our work highlights the added value of evaluating AI applications in actual multi-turn interactions with human users, looking at usability and satisfaction besides accuracy, rather than focusing on benchmark performance only.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。