用人格模拟评估智能体检索的可靠性,发现传统测试忽略的失败模式。
AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval

- 基于五大性格维度构建多样化用户群,模拟真实交互环境。
- 发现同一任务在不同人格下通过率差达50%至100%,质量差距0.27点。
- 引入风险分析器量化系统脆弱性,揭示工具层攻击占46%影响力。
智能体信息检索的评估仍局限于固定脚本与统一用户,忽视了自然的人格多样性与对抗性脆弱性。我们提出AgentWorld,一个融合(一)基于大五人格(OCEAN)的性格驱动用户群体与状态化工具使用环境;(二)基于pass$^k$的一致性度量,含结构化故障分类、部分得分与双控交接验证;(三)六种微调格式的阈值评分数据导出;(四)对抗性风险分析器,可捕捉必要中间状态骨架,对四种任务感知扰动类型进行蒙特卡洛推演,并通过$ΔP / ΔT$评分、Dempster–Shafer证据融合与Shapley攻击类别归因量化风险。三个实验验证:跨10种人格(240次评估)的对话分析代理;5个任务×4种人格变体的客服代理;以及5个任务的对抗压力测试,揭示原有轨迹脆弱性(无扰动时$V_{ ext{min}}=0.375$)与工具/基础设施层攻击主导(Shapley: 系统46%,动作38%)。人格差异暴露了统一测试无法发现的失效模式——跨域泄漏、上下文漂移、0.27分质量差距,以及同一任务下人格间通过率从50%到100%的差异;而风险分析器则量化了pass$^k$无法捕捉的轨迹级脆弱性。
原文摘要 · Abstract (English)
Evaluation of agentic information retrieval remains limited to scripted interactions with uniform users, missing both natural personality diversity and adversarial brittleness. We present AgentWorld, a simulation framework combining (i)Big Five (OCEAN) personality-driven user populations with stateful tool-use environments; (ii)the pass$^k$ consistency metric with structured fault classification, partial-credit scoring, and dual-control handoff verification; (iii)score-thresholded training-data export in six fine-tuning formats; and (iv)an adversarial Risk Analyser that snapshots required-intermediate-state spines, branches Monte-Carlo rollouts under four task-aware perturbation types, and quantifies risk via $ΔP / ΔT$ scoring, Dempster--Shafer evidence fusion, and Shapley attack-category attribution. Three experiments demonstrate the framework: a conversational analytics agent across 10 OCEAN personas (240 evaluator judgments); a customer-support agent across 5 tasks $\times$ 4 persona variants; and adversarial stress-testing of 5 tasks revealing pre-existing trajectory brittleness ($V_{\min}=0.375$ without perturbation) and tool/infrastructure-layer attack dominance (Shapley: 46% system, 38% action). Personality variation surfaces failure modes uniform testing cannot expose---cross-domain leakage, contextual drift, a 0.27-point quality gap, and 50% vs. 100% pass-rate across personas on the same task---while the Risk Analyser quantifies trajectory-level brittleness that pass$^k$ alone cannot measure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。