arXiv:2509.15447cs.CLcs.AI2025-09被引 1

用心理语言学框架精准控制生成数据的人设表现。

PILOT: Steering Synthetic Data Generation with Psychological & Linguistic Output Targeting

  • 将自然语言人设转化为可量化的心理语言学维度评分
  • 提升输出连贯性,轮廓分从0.098升至0.237,主题纯度达0.957
  • 混合模式平衡多样性与一致性,适合需可控生成的场景

生成式AI常通过用户人设引导合成数据生成,但依赖自然语言描述会导致模型误判关键属性,难以精确控制输出。本文提出PILOT(心理与语言输出定向)框架,分两阶段实现可控生成:第一阶段将自然语言人设转换为包含语言与心理维度的标准化多维评分;第二阶段基于这些评分引导大模型沿可测量的变异轴生成内容。在三种先进LLM(Mistral Large 2、Deepseek-R1、LLaMA 3.3 70B)上,使用25个合成人设,在自然语言人设引导(NPS)、基于模板的引导(SBS)和混合引导(HPS)三种条件下评估。结果表明,基于模板的方法显著减少人工感重复,提升输出连贯性,轮廓分由0.098升至0.237,主题纯度从0.773提升至0.957。分析揭示:SBS生成更简洁、主题一致,而NPS词汇多样性更高但可预测性低;HPS在两者间取得平衡,保持多样性同时维持结构一致性。专家语言评估显示,所有引导方式对响应质量无显著差异。

原文摘要 · Abstract (English)

Generative AI applications commonly leverage user personas as a steering mechanism for synthetic data generation, but reliance on natural language representations forces models to make unintended inferences about which attributes to emphasize, limiting precise control over outputs. We introduce PILOT (Psychological and Linguistic Output Targeting), a two-phase framework for steering large language models with structured psycholinguistic profiles. In Phase 1, PILOT translates natural language persona descriptions into multidimensional profiles with normalized scores across linguistic and psychological dimensions. In Phase 2, these profiles guide generation along measurable axes of variation. We evaluate PILOT across three state-of-the-art LLMs (Mistral Large 2, Deepseek-R1, LLaMA 3.3 70B) using 25 synthetic personas under three conditions: Natural-language Persona Steering (NPS), Schema-Based Steering (SBS), and Hybrid Persona-Schema Steering (HPS). Results demonstrate that schema-based approaches significantly reduce artificial-sounding persona repetition while improving output coherence, with silhouette scores increasing from 0.098 to 0.237 and topic purity from 0.773 to 0.957. Our analysis reveals a fundamental trade-off: SBS produces more concise outputs with higher topical consistency, while NPS offers greater lexical diversity but reduced predictability. HPS achieves a balance between these extremes, maintaining output variety while preserving structural consistency. Expert linguistic evaluation confirms that PILOT maintains high response quality across all conditions, with no statistically significant differences between steering approaches.

人设生成可控生成语言学建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。