从图文混合的用户反馈中自动提取相关片段,提升产品优化效率。
Leveraging Customer Feedback for Multi-modal Insight Extraction
- 在隐空间融合图文信息,用图像引导文本解码提取相关段落。
- 弱监督数据生成技术构建训练集,模型在未见数据上F1提升14点。
- 适合需要分析用户反馈的企业,尤其适用于多模态评论场景。
企业可从文本和图像等多种形式的用户反馈中获益,以改进产品与服务。然而,在单次处理中提取相关的文本片段与图像仍具挑战。本文提出一种新型多模态方法,将图像与文本信息在隐空间中融合,并通过图像-文本锚定的文本解码器提取相关反馈内容。同时引入弱监督数据生成技术,为该任务构建训练数据。在未见数据上的评估表明,该模型能有效挖掘可操作的洞察,相比现有基线在F1分数上提升14分。
原文摘要 · Abstract (English)
Businesses can benefit from customer feedback in different modalities, such as text and images, to enhance their products and services. However, it is difficult to extract actionable and relevant pairs of text segments and images from customer feedback in a single pass. In this paper, we propose a novel multi-modal method that fuses image and text information in a latent space and decodes it to extract the relevant feedback segments using an image-text grounded text decoder. We also introduce a weakly-supervised data generation technique that produces training data for this task. We evaluate our model on unseen data and demonstrate that it can effectively mine actionable insights from multi-modal customer feedback, outperforming the existing baselines by $14$ points in F1 score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。