arXiv:2502.09573cs.CVcs.CL2025-02

优化提示词让GPT零样本识别视频质量,效果显著提升。

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering

  • 通过简化策略和分解聚合式提示工程提升模型表现
  • 在七个视频质量类别上显著降低误判率
  • 无需微调,适合工业界快速部署使用

本研究针对视频内容分类的产业挑战,探索并优化基于GPT的模型在七类关键视频质量上的零样本分类性能。提出一种新的提示词优化与策略精炼方法,证明简化复杂策略可有效减少误报。引入基于分解-聚合的提示工程技巧,优于传统单提示方法。实验基于真实产业问题,表明精心设计提示词可在不进行额外微调的情况下显著提升GPT性能,为视频分类提供高效且可扩展的解决方案。

原文摘要 · Abstract (English)

In this study, we tackle industry challenges in video content classification by exploring and optimizing GPT-based models for zero-shot classification across seven critical categories of video quality. We contribute a novel approach to improving GPT's performance through prompt optimization and policy refinement, demonstrating that simplifying complex policies significantly reduces false negatives. Additionally, we introduce a new decomposition-aggregation-based prompt engineering technique, which outperforms traditional single-prompt methods. These experiments, conducted on real industry problems, show that thoughtful prompt design can substantially enhance GPT's performance without additional finetuning, offering an effective and scalable solution for improving video classification.

视频理解提示工程零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。