arXiv:2601.01062cs.LGcs.AI2026-01中稿 · WVAQ 2026, WACV 20…

评测视觉语言模型生成多角色播客叙事能力,提出新基准与评估方法。

SPoRC-VIST: A Benchmark for Evaluating Generative Natural Narrative in Vision-Language Models

  • 用合成图像+真实对话训练,再在真实图片上测试模型泛化能力。
  • 微调后模型对话自然度胜过235B基线模型80%以上,叙事更长50%。
  • 采用AI评委和新指标,超越传统文本重合度评价方式。

视觉语言模型在图像描述和视觉问答等任务中取得显著进展,但在生成引人入胜的长篇叙事(尤其是多角色播客对话)方面仍缺乏探索且难以评估。标准指标如BLEU和ROUGE无法捕捉对话自然性、个性和叙事连贯性,常奖励安全但重复的输出。本文提出端到端视觉播客生成流程,对Qwen3-VL-32B模型在4000个图像-对话对数据集上进行微调。关键采用从合成到真实的训练策略:在结构化播客研究语料库(SPoRC)的高质量播客对话与合成图像配对的数据上训练,并在视觉叙事数据集(VIST)的真实照片序列上评估。该设置检验模型从合成数据到真实视觉域的泛化能力。我们构建了超越文本重合度的综合评估框架,使用AI作为评判者(Gemini 3 Pro、Claude Opus 4.5、GPT 5.2)及新指标(平均发言长度、说话人切换率)。实验表明,微调后的32B模型在对话自然度上显著优于235B基线模型(>80%胜率),叙事深度提升50%(发言长度),同时保持相同的视觉对齐能力(CLIPScore: 20.39)。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have achieved remarkable success in descriptive tasks such as image captioning and visual question answering (VQA). However, their ability to generate engaging, long-form narratives -- specifically multi-speaker podcast dialogues -- remains under-explored and difficult to evaluate. Standard metrics like BLEU and ROUGE fail to capture the nuances of conversational naturalness, personality, and narrative flow, often rewarding safe, repetitive outputs over engaging storytelling. In this work, we present a novel pipeline for end-to-end visual podcast generation, and fine-tune a Qwen3-VL-32B model on a curated dataset of 4,000 image-dialogue pairs. Crucially, we use a synthetic-to-real training strategy: we train on high-quality podcast dialogues from the Structured Podcast Research Corpus (SPoRC) paired with synthetically generated imagery, and evaluate on real-world photo sequences from the Visual Storytelling Dataset (VIST). This rigorous setup tests the model's ability to generalize from synthetic training data to real-world visual domains. We propose a comprehensive evaluation framework that moves beyond textual overlap, and use AI-as-a-judge (Gemini 3 Pro, Claude Opus 4.5, GPT 5.2) and novel style metrics (average turn length, speaker switch rate) to assess quality. Our experiments demonstrate that our fine-tuned 32B model significantly outperforms a 235B base model in conversational naturalness ($>$80\% win rate) and narrative depth (+50\% turn length), while maintaining identical visual grounding capabilities (CLIPScore: 20.39).

视觉叙事多模态生成评估基准播客生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。