arXiv:2509.19952cs.CVcs.AI2025-09

用视频生成用户投诉文本,让难言之痛有话可说。

When Words Can't Capture It All: Towards Video-Based User Complaint Text Generation with Multimodal Video Complaint Dataset

  • 构建多模态投诉数据集ComVID,含1175条视频与对应描述。
  • 提出新任务CoD-V和情感感知的生成模型,提升投诉表达准确性。
  • 适合做用户行为分析、人机交互和情感计算的研究者参考。

尽管已有大量可解释性投诉挖掘工作,但通过文字或视频清晰表达用户关切仍具挑战,常导致问题无法解决。用户往往难以用文字准确描述问题,却能轻松上传展示产品缺陷的视频(如‘最差产品’配5秒耳机右耳罩破损视频)。本文提出新的投诉挖掘任务——视频生成投诉描述(CoD-V),旨在帮助用户表达具体问题(如右耳罩损坏)。为此,我们构建了ComVID数据集,包含1,175个投诉视频及其对应描述,并标注了投诉者的情绪状态。同时提出新的评价指标投诉保留率(CR),区分CoD-V与标准视频摘要任务。为增强效果,引入融合检索增强生成(RAG)的VideoLLaMA2-7b模型,结合用户情绪生成投诉文本。我们在多个预训练及微调版本的视频语言模型上进行综合评估,采用METEOR、困惑度、Coleman-Liau可读性评分等指标。本研究为用户提供以视频表达投诉的新平台,相关数据与资源已开源:https://github.com/sarmistha-D/CoD-V。

原文摘要 · Abstract (English)

While there exists a lot of work on explainable complaint mining, articulating user concerns through text or video remains a significant challenge, often leaving issues unresolved. Users frequently struggle to express their complaints clearly in text but can easily upload videos depicting product defects (e.g., vague text such as `worst product' paired with a 5-second video depicting a broken headphone with the right earcup). This paper formulates a new task in the field of complaint mining to aid the common users' need to write an expressive complaint, which is Complaint Description from Videos (CoD-V) (e.g., to help the above user articulate her complaint about the defective right earcup). To this end, we introduce ComVID, a video complaint dataset containing 1,175 complaint videos and the corresponding descriptions, also annotated with the emotional state of the complainer. Additionally, we present a new complaint retention (CR) evaluation metric that discriminates the proposed (CoD-V) task against standard video summary generation and description tasks. To strengthen this initiative, we introduce a multimodal Retrieval-Augmented Generation (RAG) embedded VideoLLaMA2-7b model, designed to generate complaints while accounting for the user's emotional state. We conduct a comprehensive evaluation of several Video Language Models on several tasks (pre-trained and fine-tuned versions) with a range of established evaluation metrics, including METEOR, perplexity, and the Coleman-Liau readability score, among others. Our study lays the foundation for a new research direction to provide a platform for users to express complaints through video. Dataset and resources are available at: https://github.com/sarmistha-D/CoD-V.

视频生成投诉挖掘多模态情感识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。