What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

## Computer Science > Artificial Intelligence

## Computer Science > Artificial Intelligence ## Title:What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks > Abstract:Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of benchmarks offers another perspective: what researchers expect LLMs to do, and what they count as successful performance. We systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January…

Читать полностью →

Источник: ArXiv cs.AI

Подключаюсь к источникам…

30 главных источников
о мире ИИ

Автоматический перевод, курирование и красивая подача главных статей об искусственном интеллекте.

0
статей
0
источников
9
разделов