评测大模型网络代理在文化社交认知上的表现,发现其实际应用中严重不足。
Evaluating Cultural and Social Awareness of LLM Web Agents
- 构建CASA基准,评估代理在网购和社交论坛中的文化敏感性。
- 现有模型在网页代理环境下的认知覆盖率低于10%,违规率超40%。
- 微调加提示可互补提升跨区域泛化与复杂任务应对能力,适合安全合规研究者。
随着大语言模型(LLMs)拓展至真实世界应用中的智能体角色,评估其鲁棒性愈发重要。然而,现有基准常忽略文化与社会意识等关键维度。为此,我们提出CASA基准,用于评估LLM代理在在线购物和社交讨论论坛两类网页任务中对文化与社会规范的敏感度。该方法评估代理识别并恰当响应违规用户查询与观察的能力。同时,我们设计了综合性评估框架,涵盖意识覆盖范围、处理用户请求的帮助性,以及面对误导性网络内容时的违规率。实验表明,当前LLMs在非代理环境中表现优于网页代理环境,代理在网页场景中意识覆盖率不足10%,违规率超过40%。为提升性能,我们探索了提示工程与微调两种方法,发现两者结合可产生互补优势:在文化特异性数据集上微调显著增强代理在不同地区间的泛化能力,而提示则提升其应对复杂任务的能力。这些发现凸显了在开发周期中持续评测LLM代理文化与社会意识的重要性。
原文摘要 · Abstract (English)
As large language models (LLMs) expand into performing as agents for real-world applications beyond traditional NLP tasks, evaluating their robustness becomes increasingly important. However, existing benchmarks often overlook critical dimensions like cultural and social awareness. To address these, we introduce CASA, a benchmark designed to assess LLM agents' sensitivity to cultural and social norms across two web-based tasks: online shopping and social discussion forums. Our approach evaluates LLM agents' ability to detect and appropriately respond to norm-violating user queries and observations. Furthermore, we propose a comprehensive evaluation framework that measures awareness coverage, helpfulness in managing user queries, and the violation rate when facing misleading web content. Experiments show that current LLMs perform significantly better in non-agent than in web-based agent environments, with agents achieving less than 10% awareness coverage and over 40% violation rates. To improve performance, we explore two methods: prompting and fine-tuning, and find that combining both methods can offer complementary advantages -- fine-tuning on culture-specific datasets significantly enhances the agents' ability to generalize across different regions, while prompting boosts the agents' ability to navigate complex tasks. These findings highlight the importance of constantly benchmarking LLM agents' cultural and social awareness during the development cycle.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。