arXiv:2509.11777cs.CLcs.LG2025-09

构建工业论坛用户反馈数据集,助力隐私保护下的用户体验分析。

User eXperience Perception Insights Dataset (UXPID): Synthetic User Feedback from Public Industrial Forums

  • 从公开论坛提取7130条合成匿名评论,含多轮对话与元数据。
  • 用大模型标注用户体验、情绪、严重性等,支持多种任务训练。
  • 适合研究工业软件支持中的自然语言处理与用户需求挖掘。

工业论坛中的客户反馈蕴含丰富的实际产品体验信息,但因内容非结构化、领域特定且高质量标注数据稀缺,系统分析仍具挑战。本文提出用户体验感知洞察数据集(UXPID),包含7130条从公开工业自动化论坛中提取的合成与匿名化用户反馈分支。每条JSON记录包含多轮评论、元数据,并由大语言模型(LLM)标注用户体验洞察、用户期望、严重性评分、情感倾向及主题分类。UXPID旨在推动用户需求、用户体验(UX)分析与AI驱动反馈处理的研究,尤其适用于受隐私与许可限制而无法获取真实数据的场景。该数据集支持基于Transformer模型的任务训练与评估,如问题检测、情感分析、需求提取,为工业产品支持与软件工程领域的NLP方法发展提供重要资源。

原文摘要 · Abstract (English)

Customer feedback in industrial forums offers rich but underexplored insights into real-world product experience. Yet systematic analysis remains challenging due to unstructured, domain-specific content and the scarcity of high-quality labeled datasets. This paper presents the User eXperience Perception Insights Dataset (UXPID), a collection of 7130 synthesized and anonymized user feedback branches extracted from a public industrial automation forum. Each JSON record contains multi-post comments enriched with metadata and annotated by a large language model (LLM) for UX insights, user expectations, severity ratings, sentiment, and topic classifications. UXPID is designed to facilitate research in user requirements, user experience (UX) analysis, and AI-driven feedback processing, particularly where privacy and licensing restrictions limit access to real-world data. It supports the training and evaluation of transformer-based models for tasks such as issue detection, sentiment analysis, and requirements extraction in technical forums, providing a valuable resource for advancing NLP methods within industrial product support and software engineering domains.

用户体验NLP数据集工业AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。