arXiv:2609.08592cs.AIcs.CL2026-09

用三层向量模拟真实用户,让大模型评测更可信。

A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation

  • 用23个维度构建用户人格向量,分三层次控制行为与情绪。
  • 不同人格导致代理目标达成率相差15.8个百分点,效果显著。
  • 适合做智能体评估的开发者,尤其关注真实场景测试者。

评估工具增强型大模型智能体需要多样且真实的用户输入,但现有框架多采用扁平角色描述(如“你是一个愤怒的顾客”),在不同场景下生成几乎相同的对话。本文提出一个包含23个可操作维度的三层人格向量:6个分类人口统计特征(管辖区域、年龄、沟通渠道、设备、语言能力、时间可用性)、12个连续行为特质(如耐心、果断性、数字素养)以高斯噪声围绕预设基向量采样,以及5个随场景动态变化的连续情绪状态(挫败感、焦虑、信任、自信、压力)。人格之外,引入4级查询复杂度叠加层,控制语句从直接到刻意模糊。我们在合成数据生成流程中评估该模型,覆盖8个命名人格和3个生产数据集,共生成64,698轮多轮对话。关键发现:(i) 不同人格间代理目标达成率差异达15.8个百分点,证实特质向量能产生可测量的行为差异;(ii) 同一人格在不同场景中因情绪状态响应而表现不同,验证了场景自适应设计的有效性;(iii) 在领域特定项目中,人格敏感性体现在预订流程合规性上,差距达15-20个百分点,准确复现了真实世界难度分布;(iv) 七个规则定义的特质相关性产生可审计的共现模式,无需学习协方差矩阵。该人格模型完全可复现。

原文摘要 · Abstract (English)

Evaluating tool-augmented LLM agents requires diverse, realistic user inputs yet most evaluation frameworks use flat role descriptions ("you are an angry customer") that produce near-identical conversations regardless of the underlying scenario. In this paper, we propose a three-tier persona vector with 23 operationalized dimensions: 6 categorical demographics (jurisdiction, age, channel, device, language proficiency, time availability), 12 continuous behavioral traits (patience, assertiveness, digital literacy, etc.) sampled with Gaussian noise around curated profile base vectors, and 5 continuous emotional states (frustration, anxiety, trust, confidence, stress) that shift in response to scenario context. Orthogonal to the persona, a 4-level query-complexity overlay controls utterance phrasing from direct to deliberately vague. We evaluate the persona model inside a synthetic data generation pipeline across 64,698 multi-turn conversations spanning 8 named profiles and 3 production corpora. Key findings: (i) a 15.8 percentage-point spread in agent goal-achievement across personas confirms trait vectors produce measurably different user behavior; (ii) the same persona behaves differently across scenarios due to scenario-reactive emotional state shifts, validating the scenario-reactive design; (iii) domain-specific projects show persona sensitivity on booking-flow compliance (~15-20 percentage points gap between tier-aware and pressure-test personas), demonstrating the model faithfully reproduces real-world difficulty distributions; (iv) seven rule-described trait correlations produce auditable co-occurrence patterns without requiring learned covariance matrices. The persona model is fully specified for reproduction.

智能体评估人格建模可控生成用户模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。