评测希腊语图书检索中嵌入与大模型表现,发现混合方法最优。
A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset
- 用混合检索融合词频与语义,提升多类型查询效果。
- 多语言嵌入优于专用希腊语模型,混合方法准确率最高。
- 适合关注跨语言、自然语言查询的图书馆与出版机构。
我们提出CUP,一个包含868条图书目录记录和104个专家标注的分级相关性查询的希腊语图书检索基准。评估了稀疏(BM25)、稠密(sentence-transformers)、混合及大语言模型辅助检索方法在该场景下的表现。多语言嵌入优于希腊语专用模型,混合检索整体表现最佳。查询层面分析显示,BM25在命名实体查询中更优,而稠密与混合方法在自然语言、噪声、跨语言及概念查询上表现更佳。领域感知提示对不同模型影响各异,大语言模型生成目录摘要可提升仅含目录的检索效果,但后处理过滤虽提升早期检索精度,代价较高。总体而言,CUP支持在词汇、语义、噪声及跨语言查询下对希腊语检索的真实世界评估。
原文摘要 · Abstract (English)
We present CUP, a Greek book retrieval benchmark consisting of 868 catalog records and 104 expert-annotated queries with graded relevance judgments. We evaluate sparse (BM25), dense (sentence-transformers), hybrid, and LLM-assisted retrieval methods in this book-search setting. Multilingual embeddings outperform Greek-specific models, while hybrid retrieval performs best overall. A query-level analysis shows that BM25 excels at named-entity queries, while dense and hybrid methods improve natural-language, noisy, cross-lingual, and concept queries. Field-aware prompting has model-specific effects, while LLM TOC summarization improves TOC-only retrieval and LLM post-filtering improves early-stage retrieval at a high cost. Overall, CUP enables real-world evaluation of Greek retrieval across lexical, semantic, noisy, and cross-lingual queries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。