ChatGPT无需微调就能精准识别电影类型,结合海报图像更胜一筹。
Demystifying ChatGPT: How It Masters Genre Recognition
- 用影评预告片文本做零样本/少样本提示,评估模型类型识别能力。
- 未微调的ChatGPT已优于其他大模型,微调后表现最佳,准确率超90%。
- 融合电影海报视觉信息可显著提升预测效果,适合多模态内容分析场景。
ChatGPT的出现引发自然语言处理领域的广泛关注。已有研究证实其在多种下游任务中表现卓越,展现出强大的适应性与应用潜力,但其在类型识别方面的能力仍不明确。本研究基于MovieLens-100K数据集(含18类、1682部电影,每部可关联多个类型),对比分析三种大语言模型(LLMs)的类型预测性能。实验采用影评预告片文本作为输入,设计零样本与少样本提示,结果显示,未经微调的ChatGPT已优于其他模型,微调后的版本表现最优。进一步地,通过提取IMDb电影海报并引入视觉语言模型(VLM),将海报视觉特征融入提示,显著提升了提示质量。结果表明,ChatGPT具备出色的类型识别能力,且结合视觉信息后效果更佳,展现了其在内容相关应用中的巨大潜力。
原文摘要 · Abstract (English)
The introduction of ChatGPT has garnered significant attention within the NLP community and beyond. Previous studies have demonstrated ChatGPT's substantial advancements across various downstream NLP tasks, highlighting its adaptability and potential to revolutionize language-related applications. However, its capabilities and limitations in genre prediction remain unclear. This work analyzes three Large Language Models (LLMs) using the MovieLens-100K dataset to assess their genre prediction capabilities. Our findings show that ChatGPT, without fine-tuning, outperformed other LLMs, and fine-tuned ChatGPT performed best overall. We set up zero-shot and few-shot prompts using audio transcripts/subtitles from movie trailers in the MovieLens-100K dataset, covering 1682 movies of 18 genres, where each movie can have multiple genres. Additionally, we extended our study by extracting IMDb movie posters to utilize a Vision Language Model (VLM) with prompts for poster information. This fine-grained information was used to enhance existing LLM prompts. In conclusion, our study reveals ChatGPT's remarkable genre prediction capabilities, surpassing other language models. The integration of VLM further enhances our findings, showcasing ChatGPT's potential for content-related applications by incorporating visual information from movie posters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。