评测大模型代理在真实社会服务导航中的表现,发现其应对复杂用户交互能力不足。
HoosierHelp: Benchmarking LLM Agents for Social Service Navigation

- 构建交互式基准测试平台,模拟真实用户需求与非理想对话行为。
- 7个大模型在3971项资源中仅部分准确匹配需求,尤其在矛盾或需回退场景下表现差。
- 适合研究社会服务对话系统、大模型鲁棒性及人机交互的学者与开发者。
社会服务导航需要将求助者与满足其需求及特定约束的资源对接。尽管大模型代理为对话式资源导航提供了潜在接口,但现有基准未能体现该场景中的交互复杂性和约束依赖性。我们提出HoosierHelp,一个基于印第安纳州3,971项公共社会服务资源的交互式基准。代理需与模拟用户互动,发出结构化资源搜索请求,处理非理想交互,并从工具返回结果中选择最终资源。通过变化用户的需求结构、约束可满足性及行为模式(如不耐烦、啰嗦、无支持请求、自我矛盾),提升模拟用户的真实性。在240个样本上对7个大模型的实验表明,当前大模型代理在社会服务导航中仍存在显著不可靠性,尤其在需要回退和自我矛盾对话中性能急剧下降,凸显出对更鲁棒代理的需求。
原文摘要 · Abstract (English)
Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints. Although LLM agents offer a promising interface for conversational resource navigation, existing benchmarks do not capture the interaction complexity and constraint-grounding demands of this setting. We introduce HoosierHelp, an interactive benchmark grounded in 3,971 Indiana public social service resources. Agents interact with simulated users, issue structured resource-search calls, handle non-ideal interactions, and select the final resources returned by the tool. HoosierHelp enhances the realism of simulated users by varying their need structure, constraint satisfiability, and behavior patterns, including impatience, rambling, unsupported requests, and self-contradiction. Experiments on 240 samples across seven LLMs show that current LLM agents remain substantially unreliable for social service navigation. Performance drops sharply on fallback-required and self-contradictory conversations, highlighting the need for agents that are more robust to complex and non-ideal user interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。