测试大模型实时生成报告能力,比传统评测更贴近真实场景。
DynamicBench: Evaluating Real-Time Report Generation in Large Language Models
- 用网页搜索+本地数据库双路径获取最新信息
- 在无文档和有文档场景下均优于GPT4o
- 适合评估需实时更新的AI应用开发人员
传统大模型评测多依赖静态叙事或观点表达,无法反映现代应用中实时信息处理的需求。为此,我们提出DynamicBench,一个评估大模型存储与处理最新数据能力的基准。该基准采用双路径检索架构,融合网络搜索与本地报告数据库,要求具备领域专业知识,确保在特定领域内生成准确报告。通过在提供或不提供外部文档的场景下评估模型,可有效衡量其独立处理最新信息或利用上下文增强的能力。此外,我们还构建了一个先进的动态信息合成报告系统。实验结果表明,该方法在无文档和有文档场景下分别领先GPT4o 7.0%和5.8%,达到当前最佳性能。代码与数据将公开。
原文摘要 · Abstract (English)
Traditional benchmarks for large language models (LLMs) typically rely on static evaluations through storytelling or opinion expression, which fail to capture the dynamic requirements of real-time information processing in contemporary applications. To address this limitation, we present DynamicBench, a benchmark designed to evaluate the proficiency of LLMs in storing and processing up-to-the-minute data. DynamicBench utilizes a dual-path retrieval pipeline, integrating web searches with local report databases. It necessitates domain-specific knowledge, ensuring accurate responses report generation within specialized fields. By evaluating models in scenarios that either provide or withhold external documents, DynamicBench effectively measures their capability to independently process recent information or leverage contextual enhancements. Additionally, we introduce an advanced report generation system adept at managing dynamic information synthesis. Our experimental results confirm the efficacy of our approach, with our method achieving state-of-the-art performance, surpassing GPT4o in document-free and document-assisted scenarios by 7.0% and 5.8%, respectively. The code and data will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。