对比三款大模型生成影评,发现深度思考模型表现最均衡。
An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3
- 用剧本和字幕输入大模型生成影评,评估语言流畅性与情感一致性。
- DeepSeek-V3生成影评最贴近真实用户评论,情感平衡度最佳。
- 尽管难分辨,但大模型仍缺乏真实影评的情感深度与风格统一。
大型语言模型(LLMs)在文本生成和摘要任务中表现突出,其在产品评论生成中的应用正逐步扩展至电影评论领域。本研究提出一个框架,利用 GPT-4o、DeepSeek-V3 和 Gemini-2.0 三款模型生成影评,并以 IMDb 用户评论为基准进行评估。以电影字幕和剧本作为输入,考察其对生成影评质量的影响。从词汇丰富度、情感极性、相似性及主题一致性等方面对比分析生成结果。结果显示,大模型能生成语法正确、结构完整的影评,但在情感丰富度与风格连贯性上仍明显逊于 IMDb 真实用户评论,表明尚需优化。通过调查问卷让参与者辨别生成影评与真实影评,结果表明生成内容难以区分。其中,DeepSeek-V3 生成的影评最为平衡,最接近真实用户评论;GPT-4o 过度强调正面情绪,而 Gemini-2.0 虽更擅长捕捉负面情绪,但情感强度过强。
原文摘要 · Abstract (English)
Large language models (LLMs) have been prominent in various tasks, including text generation and summarisation. The applicability of LLMs to the generation of product reviews is gaining momentum, paving the way for the generation of movie reviews. In this study, we propose a framework that generates movie reviews using three LLMs (GPT-4o, DeepSeek-V3, and Gemini-2.0), and evaluate their performance by comparing the generated outputs with IMDb user reviews. We use movie subtitles and screenplays as input to the LLMs and investigate how they affect the quality of reviews generated. We review the LLM-based movie reviews in terms of vocabulary, sentiment polarity, similarity, and thematic consistency in comparison to IMDB user reviews. The results demonstrate that LLMs are capable of generating syntactically fluent and structurally complete movie reviews. Nevertheless, there is still a noticeable gap in emotional richness and stylistic coherence between LLM-generated and IMDb reviews, suggesting that further refinement is needed to improve the overall quality of movie review generation. We provided a survey-based analysis where participants were told to distinguish between LLM and IMDb user reviews. The results show that LLM-generated reviews are difficult to distinguish from IMDB user reviews. We found that DeepSeek-V3 produced the most balanced reviews, closely matching IMDb reviews. GPT-4o overemphasised positive emotions, while Gemini-2.0 captured negative emotions better but showed excessive emotional intensity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。