用大模型生成产品吸引力数据集,低成本高效替代真实用户评论。
Utilizing Large Language Models to Synthesize Product Desirability Datasets
- 用GPT-4o-mini通过三种组合方式合成1000条产品评论。
- 情感一致性达0.93~0.97,供词法覆盖更广且多样性最高。
- 适合数据稀缺场景,提升测试可扩展性与生产灵活性。
本研究探索利用大语言模型(LLMs)生成用于产品吸引力工具包(PDT)测试的合成数据集,该工具包是评估用户情感与产品体验的关键组件。采用成本较低的gpt-4o-mini,分别通过Word+Review、Review+Word、Supply-Word三种方法各生成1000条产品评论。评估结果表明,所有方法在情感一致性上表现优异,皮尔逊相关系数为0.93至0.97之间。Supply-Word方法在术语覆盖和文本多样性方面最优,但生成成本较高。尽管存在轻微正向情感偏倚,在测试数据有限的情况下,使用LLM生成的合成数据仍具显著优势,包括可扩展性、成本节约及数据生成灵活性。
原文摘要 · Abstract (English)
This research explores the application of large language models (LLMs) to generate synthetic datasets for Product Desirability Toolkit (PDT) testing, a key component in evaluating user sentiment and product experience. Utilizing gpt-4o-mini, a cost-effective alternative to larger commercial LLMs, three methods, Word+Review, Review+Word, and Supply-Word, were each used to synthesize 1000 product reviews. The generated datasets were assessed for sentiment alignment, textual diversity, and data generation cost. Results demonstrated high sentiment alignment across all methods, with Pearson correlations ranging from 0.93 to 0.97. Supply-Word exhibited the highest diversity and coverage of PDT terms, although with increased generation costs. Despite minor biases toward positive sentiments, in situations with limited test data, LLM-generated synthetic data offers significant advantages, including scalability, cost savings, and flexibility in dataset production.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。