Leaderboard ALL KINGDOMS

Each generation method is ranked on its own board — scores aren't comparable across methods (they come from separate match pools). Pick a method to compare the models within it. 340 votes cast.

Share Post on X
Filters & bias audit
Scope Bradley–Terry over votes from signed-in (Hugging Face) users only.

Bias audit — left(A) win rate 0.569 (≈0.50 = unbiased) · tie 0.059 · bad 0.191

VLM judge (Sonnet 4.6, multi-view) — automated LLM-judge rankings by paradigm

Loading automated rankings…