研究小模型在少量偏好数据下生成幽默/浪漫风格图文描述的极限能力。
Probing the Limits of Stylistic Alignment in Vision-Language Models
- 用少量偏好数据微调小规模视觉语言模型
- 发现只需极少数据即可达到风格表现饱和点
- 为风格化生成提供高效训练基准,适合资源有限的研究者
视觉语言模型越来越多地用于生成特定风格的图像描述,如幽默或浪漫风格。然而,这些基于Transformer的模型在零样本设置下往往难以完成这一主观任务。虽然偏好数据可用于引导模型向目标风格对齐,但此类数据获取成本高,限制了模型能力的充分探索。本文通过研究小型视觉语言模型在幽默和浪漫风格对齐中的数据效率,旨在界定模型性能上限,并确定实现风格饱和所需的最少偏好数据量,从而为模型的能力与局限性提供基准评估。
原文摘要 · Abstract (English)
Vision-language models are increasingly used to generate image captions in specific styles, such as humor or romantic. However, these transformer-based models often struggle with this subjective task in a zero-shot setting. While preference data can be used to align them toward a desired style, such data is expensive to acquire, limiting the ability to explore the models' full capabilities. This work addresses this by studying the data efficiency of aligning small vision-language models to humor and romantic styles. This approach helps to define the performance limits of these models and determine how little preference data is needed to achieve stylistic saturation, benchmarking their capabilities and limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。