arXiv:2503.04188cs.CLcs.IR2025-03被引 1

用时间控制工具测试大模型写作助手的知识时效性。

Measuring temporal effects of agent knowledge by date-controlled tool use

  • 设计时间控制工具,模拟不同日期的网络搜索结果
  • 发现模型表现随搜索时间变化明显,且受基础模型和提示方式影响
  • 适合关注大模型可靠性与动态知识评估的研究者

时间演进是知识积累与更新的核心部分。网络搜索常被用作大语言模型(LLM)代理的知识来源,但配置不当会降低响应质量。本文通过使用不同的日期控制工具(DCTs)作为压力测试,评估大型语言模型代理在撰写科学论文摘要时的行为表现。结果表明,搜索引擎的时间特性直接影响代理性能,且该影响可通过选择合适的基模型和采用链式思维提示等显式推理指令缓解。研究强调,代理的设计与评估应采取动态视角,充分考虑外部资源的时间影响,以保障系统可靠性。

原文摘要 · Abstract (English)

Temporal progression is an integral part of knowledge accumulation and update. Web search is frequently adopted as grounding for agent knowledge, yet an improper configuration affects the quality of the agent's responses. Here, we assess the agent behavior using distinct date-controlled tools (DCTs) as stress test to measure the knowledge variability of large language model (LLM) agents. We demonstrate the temporal effects of an LLM agent as a writing assistant, which uses web search to complete scientific publication abstracts. We show that the temporality of search engine translates into tool-dependent agent performance but can be alleviated with base model choice and explicit reasoning instructions such as chain-of-thought prompting. Our results indicate that agent design and evaluations should take a dynamical view and implement measures to account for the temporal influence of external resources to ensure reliability.

大模型知识时效工具使用推理提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。