构建跨领域深度研究评估基准,测试模型的准确性与客观性。
DRACO: a Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity
- 从真实用户请求中提取10个领域任务,覆盖40国信息源。
- 多维度评分:事实准确、分析完整度、表述客观性、引用质量。
- 适合评估大模型在复杂研究任务中的综合表现,开放可用。
我们提出DRACO(Deep Research Accuracy, Completeness, and Objectivity),一个涵盖10个领域的复杂深度研究任务基准。这些任务源自大规模深度研究系统中的匿名真实使用数据,覆盖40个国家的信息来源。任务从去标识化的Perplexity Deep Research请求数据集中采样,经筛选与增强,确保匿名性、开放性、复杂性及可客观评估性,并代表真实世界深度研究的广泛场景。输出依据任务特定评分标准,从四个维度进行评估:事实准确性、分析广度与深度(完整性)、呈现质量(客观性)以及引用质量。DRACO已公开发布于https://hf.co/datasets/perplexity-ai/draco。
原文摘要 · Abstract (English)
We present DRACO (Deep Research Accuracy, Completeness, and Objectivity), a benchmark of complex deep research tasks. These tasks, which span 10 domains and draw on information sources from 40 countries, originate from anonymized real-world usage patterns within a large-scale deep research system. Tasks are sampled from a de-identified dataset of Perplexity Deep Research requests, then filtered and augmented to ensure that the tasks are anonymized, open-ended and complex, objectively evaluable, and representative of the broad scope of real-world deep research use cases. Outputs are graded against task-specific rubrics along four dimensions: factual accuracy (accuracy), breadth and depth of analysis (including completeness), presentation quality (including objectivity), and citation quality. DRACO is publicly available at https://hf.co/datasets/perplexity-ai/draco.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。