医学视频问答中,用GPT4结合视觉与字幕定位答案,提升精准度。
PolySmart @ TRECVid 2024 Medical Video Question Answering
- 用GPT4匹配视频字幕与答案,检索相关视频。
- 通过视觉内容与字幕对齐,定位答案起止时间,平均IoU达9.65。
- 适合医疗视觉问答、多模态模型应用研究者参考。
视频语料库视觉答案定位(VCVAL)包含与问题相关的视频检索及视频中的视觉答案定位。具体而言,我们采用文本到文本的检索方法,基于视频转录文本与GPT4生成的答案之间的相似性,检索与医学问题相关的视频。对于视觉答案定位,通过查询与视频内容及字幕间的对齐,预测答案的起止时间戳。在查询聚焦的教学步骤描述(QFISC)任务中,使用GPT4生成步骤描述:以LLaVA-Next-Video模型生成的视频字幕和带时间戳的字幕作为上下文,要求GPT4为给定医学问题生成步骤描述。我们仅提交一次运行结果,获得F-score为11.92,平均交并比(mean IoU)为9.6527。
原文摘要 · Abstract (English)
Video Corpus Visual Answer Localization (VCVAL) includes question-related video retrieval and visual answer localization in the videos. Specifically, we use text-to-text retrieval to find relevant videos for a medical question based on the similarity of video transcript and answers generated by GPT4. For the visual answer localization, the start and end timestamps of the answer are predicted by the alignments on both visual content and subtitles with queries. For the Query-Focused Instructional Step Captioning (QFISC) task, the step captions are generated by GPT4. Specifically, we provide the video captions generated by the LLaVA-Next-Video model and the video subtitles with timestamps as context, and ask GPT4 to generate step captions for the given medical query. We only submit one run for evaluation and it obtains a F-score of 11.92 and mean IoU of 9.6527.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。