用大模型精准量化产品吸引力,结果可解释且成本低。
Evaluating LLM Usage for Efficient and Explainable Numerical and Classified Implicit Sentiment Analysis of Product Desirability

- 直接从用户描述生成数值化情感分和分类标签,无需评分数据。
- 相关系数达0.97,分类准确率94%,优于词典与传统模型。
- 支持低成本部署,输出理由与置信度,提升可解释性。
定性产品反馈能揭示用户深层体验,但其隐含情感难以量化。本文提出一种可扩展且可解释的框架,利用大语言模型(LLM)从此类数据中量化产品吸引力。基于ZORQ与CARMA的两个产品吸引力工具包(PDT)数据集,包含106组受试者术语分组及人工标注金标准,评估了零样本连续数值情感评分与分类任务。在不依赖显式评分的情况下,LLM直接从定性回答生成情感分数,与专家标签高度一致,最高皮尔逊相关系数达0.97,分类准确率最高94%。模型在多种输入形式下仍保持鲁棒性,并持续输出高置信度。相比之下,词典基与Transformer基线未达统计显著水平。其中,GPT-4o-mini表现媲美大模型,成本降低94%。框架还集成模型置信度与人类可读推理说明(xAI),增强可解释性与可信度,适用于产品满意度评估。总体而言,结合PDT问卷与高效LLM进行情感分析,可为产品评估提供丰富的情感分数(数值与分类)及高层级用户印象,助力产品改进与营销策略制定。
原文摘要 · Abstract (English)
Qualitative product feedback can reveal nuanced user experiences, but its implicit sentiment is difficult to measure. This paper presents a scalable and interpretable framework that uses large language models (LLMs) to quantify product desirability from such data. Using two Product Desirability Toolkit (PDT) datasets from ZORQ and CARMA comprising 106 respondent term groupings with gold-standard human annotation, zero-shot continuous numerical sentiment scoring and categorical sentiment classification are evaluated without relying on explicit review scores. Across the datasets, LLMs generated numerical sentiment scores directly from qualitative responses and closely matched expert labels, achieving Pearson correlations up to 0.97 and classification accuracy up to 94%. LLMs maintained robustness even when handling data presented in multiple forms and consistently expressed high confidence. In contrast, lexicon-based and transformer baselines did not produce statistically significant results. Among the models tested, GPT-4o-mini achieved performance comparable to larger models at 94% lower cost, supporting scalable deployment. The framework also incorporates model confidence ratings and human-readable rationale explanations (xAI), improving interpretability, transparency, and trust while supporting practical use in product satisfaction assessment. In general, using the PDT tool as a survey method along with a cost efficient LLM for sentiment analysis has the potential to provide for product evaluation with results that are rich in terms of sentiment scores (both numerical and classified sentiment) and in terms of the high-level user impressions of the product that can be used to identify ideas for product development and improvement, as well as marketing ideas for target audiences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。