arXiv:2508.19272cs.CL2025-08被引 1

RAGAPHENE让人工模拟真实对话,评测大模型的问答可靠性。

RAGAPHENE: A RAG Annotation Platform with Human Enhancements and Edits

  • 基于聊天界面的标注平台,支持多轮真实对话模拟
  • 40名标注员构建了数千条高质量对话数据
  • 适合评估大模型在复杂场景下的事实准确性

检索增强生成(RAG)在需要准确信息的大型语言模型(LLM)对话中至关重要。大模型可能生成看似正确但包含幻觉的内容。因此,构建能够评估多轮RAG对话的基准测试变得愈发重要。模拟真实世界对话对生成高质量评估基准至关重要。我们提出RAGAPHENE,一个基于聊天的标注平台,使标注员能模拟真实对话以构建和评估大模型。该平台已由约40名标注员成功使用,构建了数千条真实世界对话。

原文摘要 · Abstract (English)

Retrieval Augmented Generation (RAG) is an important aspect of conversing with Large Language Models (LLMs) when factually correct information is important. LLMs may provide answers that appear correct, but could contain hallucinated information. Thus, building benchmarks that can evaluate LLMs on multi-turn RAG conversations has become an increasingly important task. Simulating real-world conversations is vital for producing high quality evaluation benchmarks. We present RAGAPHENE, a chat-based annotation platform that enables annotators to simulate real-world conversations for benchmarking and evaluating LLMs. RAGAPHENE has been successfully used by approximately 40 annotators to build thousands of real-world conversations.

RAG标注平台对话评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。