用三层向量模拟真实用户,让大模型评测更可信。
A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation
- 用23个维度构建用户人格向量,分三层次控制行为与情绪。
- 不同人格导致代理目标达成率相差15.8个百分点,效果显著。
- 适合做智能体评估的开发者,尤其关注真实场景测试者。
评估工具增强型大模型智能体需要多样且真实的用户输入,但现有框架多采用扁平角色描述(如“你是一个愤怒的顾客”),在不同场景下生成几乎相同的对话。本文提出一个包含23个可操作维度的三层人格向量:6个分类人口统计特征(管辖区域、年龄、沟通渠道、设备、语言能力、时间可用性)、12个连续行为特质(如耐心、果断性、数字素养)以高斯噪声围绕预设基向量采样,以及5个随场景动态变化的连续情绪状态(挫败感、焦虑、信任、自信、压力)。人格之外,引入4级查询复杂度叠加层,控制语句从直接到刻意模糊。我们在合成数据生成流程中评估该模型,覆盖8个命名人格和3个生产数据集,共生成64,698轮多轮对话。关键发现:(i) 不同人格间代理目标达成率差异达15.8个百分点,证实特质向量能产生可测量的行为差异;(ii) 同一人格在不同场景中因情绪状态响应而表现不同,验证了场景自适应设计的有效性;(iii) 在领域特定项目中,人格敏感性体现在预订流程合规性上,差距达15-20个百分点,准确复现了真实世界难度分布;(iv) 七个规则定义的特质相关性产生可审计的共现模式,无需学习协方差矩阵。该人格模型完全可复现。
原文摘要 · Abstract (English)
Evaluating tool-augmented LLM agents requires diverse, realistic user inputs yet most evaluation frameworks use flat role descriptions ("you are an angry customer") that produce near-identical conversations regardless of the underlying scenario. In this paper, we propose a three-tier persona vector with 23 operationalized dimensions: 6 categorical demographics (jurisdiction, age, channel, device, language proficiency, time availability), 12 continuous behavioral traits (patience, assertiveness, digital literacy, etc.) sampled with Gaussian noise around curated profile base vectors, and 5 continuous emotional states (frustration, anxiety, trust, confidence, stress) that shift in response to scenario context. Orthogonal to the persona, a 4-level query-complexity overlay controls utterance phrasing from direct to deliberately vague. We evaluate the persona model inside a synthetic data generation pipeline across 64,698 multi-turn conversations spanning 8 named profiles and 3 production corpora. Key findings: (i) a 15.8 percentage-point spread in agent goal-achievement across personas confirms trait vectors produce measurably different user behavior; (ii) the same persona behaves differently across scenarios due to scenario-reactive emotional state shifts, validating the scenario-reactive design; (iii) domain-specific projects show persona sensitivity on booking-flow compliance (~15-20 percentage points gap between tier-aware and pressure-test personas), demonstrating the model faithfully reproduces real-world difficulty distributions; (iv) seven rule-described trait correlations produce auditable co-occurrence patterns without requiring learned covariance matrices. The persona model is fully specified for reproduction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。