arXiv:2510.18289cs.CLcs.CY2025-10被引 1

构建食物援助智能代理框架,精准匹配用户需求与资源

Food4All: An Agentic Framework and Benchmark for Food Resource Navigation with Adaptive User Understanding

  • 设计专用搜索工具+多轮对话评估任务,模拟真实求助场景
  • 最强模型达96.33%推荐准确率,但约束条件理解仍有缺陷
  • 适合研究人机交互、辅助弱势群体的具身智能系统开发者

食物援助转介需对话智能体将模糊、嘈杂的求助对话转化为本地有效的资源推荐。我们提出Food4All,一个基于686个印第安纳州结构化食物资源的智能体框架与基准测试。该框架结合专用搜索工具,包含300个多轮评估任务,涵盖单一食物需求、含访问或文件限制的复合案例,以及五类非理想用户行为:不合理要求、冗长回复、急躁、回答不完整、信息不一致。我们评估了六种大语言模型在需求定位、资源检索、最终推荐准确性和交互效率上的表现。尽管最强模型达到96.33%的推荐准确率,诊断显示其在时间安排、资格条件、摄入量和文件约束的定位上仍存在持续失败,且未能保留有效检索结果。行为级分析表明,不同非理想行为对转介流程的不同环节造成压力。Food4All为研究约束敏感场景下工具调用智能体在真实用户交互挑战中的表现提供了可控测试平台。

原文摘要 · Abstract (English)

Food assistance referral requires conversational agents to translate underspecified, often noisy help-seeking dialogues into locally valid resource recommendations. We present Food4All, an agentic food-resource referral framework and benchmark grounded in 686 structured Indiana food resources. Food4All couples a food-specific search tool with 300 multi-turn evaluation tasks spanning single food needs, composite cases with access or document constraints, and five non-ideal user interaction traits: unreasonable demands, rambling responses, impatience, incomplete answers, and inconsistent information. We evaluate six Large Language Models (LLMs) on requirement grounding, resource retrieval, final referral correctness, and interaction efficiency. Although the strongest model achieves 96.33% referral accuracy, our diagnostics reveal persistent failures in grounding schedule, eligibility, intake, and document constraints, as well as failures to preserve valid retrieved resources in the final recommendation. Trait-level analysis further shows that different non-ideal behaviors stress different parts of the referral pipeline. Food4All provides a controlled testbed for studying tool-calling agents in constraint-sensitive food assistance referral under realistic user interaction challenges.

智能代理食物援助对话系统大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。