用小模型生成符合人物设定的调查回复,高效又可靠。
Polypersona: Persona-Grounded LLM for Synthetic Survey Responses
- 用对话式数据流保留人物特征,确保回复一致性
- 1.1B和Phi-2小模型表现接近7B~8B大模型,最高BLEU达0.090
- 适合需要可控合成数据的研究者和评估人员
本文提出PolyPersona,一个用于跨领域生成人物设定驱动调查回复的生成框架。该框架通过资源自适应训练设置,使用4位量化与参数高效的LoRA适配器对小型聊天模型进行指令微调。基于对话式数据流水线,显式保留人物线索,确保生成回复的行为一致性。利用该流程,构建了包含3,568条合成调查回复的数据集,覆盖十大学科领域与433个不同人物设定,支持受控指令微调与系统性多域评估。采用多指标评估套件,结合标准文本生成指标(如BLEU、ROUGE、BERTScore)与专为调查设计的结构连贯性、风格一致性和情感契合度指标。实验表明,小型模型如TinyLlama 1.1B和Phi-2在性能上可媲美7B至8B的基线模型,最高BLEU为0.090,ROUGE-1达0.429。结果表明,人物设定微调使小模型能生成可靠且连贯的合成调查数据。该框架提供一种高效、可复现的调查数据生成方案,支持大规模评估并可通过透明开放协议开展偏差分析。
原文摘要 · Abstract (English)
This paper introduces PolyPersona, a generative framework for synthesizing persona-conditioned survey responses across multiple domains. The framework instruction-tunes compact chat models using parameter-efficient LoRA adapters with 4-bit quantization under a resource-adaptive training setup. A dialogue-based data pipeline explicitly preserves persona cues, ensuring consistent behavioral alignment across generated responses. Using this pipeline, we construct a dataset of 3,568 synthetic survey responses spanning ten domains and 433 distinct personas, enabling controlled instruction tuning and systematic multi-domain evaluation. We evaluate the generated responses using a multi-metric evaluation suite that combines standard text generation metrics, including BLEU, ROUGE, and BERTScore, with survey-specific metrics designed to assess structural coherence, stylistic consistency, and sentiment alignment.Experimental results show that compact models such as TinyLlama 1.1B and Phi-2 achieve performance comparable to larger 7B to 8B baselines, with a highest BLEU score of 0.090 and ROUGE-1 of 0.429. These findings demonstrate that persona-conditioned fine-tuning enables small language models to generate reliable and coherent synthetic survey data. The proposed framework provides an efficient and reproducible approach for survey data generation, supporting scalable evaluation while facilitating bias analysis through transparent and open protocols.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。