clembench v4.0 — multilingual
Nine LLMs evaluated as agents in self-play across 14 dialogue games and 30 languages: the 24 official EU languages plus Arabic, Chinese, Russian, Serbian, Turkish and Ukrainian.
clemscore combines the ability to follow game instructions at all (% Played) with how well the game was played when the instructions were followed (Quality). A model has to do both to score well.
Languages
Models
Each axis is the mean clemscore over the games in that capability cluster, averaged across the selected languages.