用语义相似度让大模型模拟真实消费者评分,效果接近真人。
LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings
- 让大模型先写文字反馈,再通过嵌入相似度转为量表评分。
- 在9300条真人数据上,测试重测信度达真人90%,分布相似度超0.85。
- 既能生成可量化的评分,还能保留有解释性的定性内容。
消费者研究每年耗费企业数十亿美元,却面临样本偏差和规模受限问题。大语言模型(LLMs)可作为替代方案,模拟合成消费者,但直接要求其给出数值评分时,结果分布不真实。本文提出语义相似度评分(SSR)方法:先让LLM生成文本反馈,再通过嵌入相似度将这些文本映射到李克特量表分布。在某领先企业开展的57份个人护理产品调查数据集(共9,300条真人响应)上,SSR实现90%的人类测试-重测信度,且分布相似度KS > 0.85。此外,合成受访者还提供了丰富定性反馈以解释评分。该框架可在保持传统调查指标与可解释性的同时,实现可扩展的消费者研究模拟。
原文摘要 · Abstract (English)
Consumer research costs companies billions annually yet suffers from panel biases and limited scale. Large language models (LLMs) offer an alternative by simulating synthetic consumers, but produce unrealistic response distributions when asked directly for numerical ratings. We present semantic similarity rating (SSR), a method that elicits textual responses from LLMs and maps these to Likert distributions using embedding similarity to reference statements. Testing on an extensive dataset comprising 57 personal care product surveys conducted by a leading corporation in that market (9,300 human responses), SSR achieves 90% of human test-retest reliability while maintaining realistic response distributions (KS similarity > 0.85). Additionally, these synthetic respondents provide rich qualitative feedback explaining their ratings. This framework enables scalable consumer research simulations while preserving traditional survey metrics and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。