arXiv:2508.03710cs.IRcs.LG2025-08

对比五款生成式AI在个性化线下活动推荐中的表现与用户满意度。

Evaluating Generative AI Tools for Personalized Offline Recommendations: A Comparative Study

  • 基于用户画像和干预场景生成非数字化活动建议。
  • 工具间精度、召回率差异显著,最优者F1-score达0.72。
  • 用户满意度与推荐相关性高度相关,适合健康行为干预研究者。

生成式AI工具在个性化推荐中日益重要,但在旨在减少技术使用的健康行为干预领域仍缺乏充分研究。本研究评估了五款最广泛使用的生成式AI工具在针对重复性劳损风险人群推荐非数字活动时的性能与用户满意度。采用目标/问题/度量(GQM)范式,实验基于预设用户档案与干预情景,要求各工具生成线下活动建议。评估涵盖定量指标(精确率、召回率、F1分数、马修斯相关系数)与定性维度(用户满意度与感知推荐相关性)。研究提出两个问题:RQ1关注哪款工具推荐最准确;RQ2分析工具选择对用户满意度的影响。结果表明,不同工具在推荐准确性上存在显著差异,最优工具达到F1-score 0.72,且用户满意度与推荐相关性高度正相关。

原文摘要 · Abstract (English)

Background: Generative AI tools have become increasingly relevant in supporting personalized recommendations across various domains. However, their effectiveness in health-related behavioral interventions, especially those aiming to reduce the use of technology, remains underexplored. Aims: This study evaluates the performance and user satisfaction of the five most widely used generative AI tools when recommending non-digital activities tailored to individuals at risk of repetitive strain injury. Method: Following the Goal/Question/Metric (GQM) paradigm, this proposed experiment involves generative AI tools that suggest offline activities based on predefined user profiles and intervention scenarios. The evaluation is focused on quantitative performance (precision, recall, F1-score and MCC-score) and qualitative aspects (user satisfaction and perceived recommendation relevance). Two research questions were defined: RQ1 assessed which tool delivers the most accurate recommendations, and RQ2 evaluated how tool choice influences user satisfaction.

生成式AI个性化推荐健康干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。