用虚拟人格评估生成模型,更真实反映人类判断多样性。
A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
- 构建虚拟认知角色模拟不同人群评价视角
- 发现角色在持续推理中会逐渐失真,导致评价不一致
- 提出动态调节机制,适合关注AI公平性与可解释性的研究者
当前生成式人工智能的对齐评估主要依赖单一基准框架,将多元的人类判断简化为统计平均值,忽略了文化、人口和情境差异。本文提出一种状态空间受限的仿真评估框架,以结构化的合成认知角色替代单一评估函数,代表多样化的人类观点。我们发现现代生成模型能高一致性地实例化并维持这些评价人格,实现更具包容性的多视角基准测试。然而,进一步分析显示,在连续推理和随机提示扰动下,这些模拟评价者稳定性下降,表现为状态空间漂移与语义不一致。这表明静态对齐约束不足以维持长期评价一致性。因此,我们主张在生成系统中嵌入动态、以生存能力为导向的调控机制,以保持认知模拟的连贯性。通过将基于人格的评估视为潜在表示流形上的结构化动态系统,本研究为更自适应、以人为本且情境敏感的AI评估方法提供了基础。
原文摘要 · Abstract (English)
Current alignment paradigms for generative artificial intelligence rely predominantly on monolithic benchmarking frameworks that reduce the plurality of human judgment to aggregated statistical baselines, thereby obscuring cultural, demographic, and contextual variability in evaluation. We introduce a state-space constrained emulation framework for AI evaluation that replaces singular assessment functions with a structured manifold of synthetic cognitive profiles representing diverse human perspectives. We show that modern generative architectures can instantiate and maintain these evaluative personas with high consistency, enabling a form of pluralistic, perspective-dependent benchmarking that more closely reflects real-world consensus variability. However, we further analyze the stability of these simulated evaluators under sequential inference and stochastic prompt perturbations, revealing systematic degradation in persona coherence that manifests as state-space drift and semantic inconsistency. These findings suggest that static alignment constraints are insufficient for sustaining robust evaluative behavior over time. Instead, we argue for the necessity of embedding dynamic, viability-driven regulatory mechanisms within generative systems to preserve coherent cognitive emulation. By framing persona-based evaluation as a structured dynamical system over latent representation manifolds, this study provides a foundation for more adaptive, human-aligned, and context-sensitive approaches to AI evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。