AI-judge board ALL KINGDOMS
An LLM writes code (e.g. Blender) that builds the 3D model.
These ranks come from a VLM judge, not human votes. A vision-language model (Sonnet 4.6, multi-view) casts the ballots on this board — no person voted on them. It is a separate, automated surface: its scores are never mixed into the human leaderboard, and the two can disagree. For the ranking humans voted for, see the human-vote LLM procedural (code-gen) board →
LLM procedural (code-gen)
Scores are only comparable within a paradigm.| Rank (UB) | Generator | BT score | Votes |
|---|---|---|---|
| 1 | openai/gpt-5.6-sol-pro | 1346.5 | 4 |
| 2 | openai/gpt-5.6-sol | 1163.8 | 68 |
| 3 | z-ai/glm-5.2 | 1110.7 | 74 |
| 4 | anthropic/claude-sonnet-5 | 1096.8 | 60 |
| 5 | moonshotai/kimi-k2.7-code | 1073.7 | 62 |
| 6 | google/gemini-3.1-pro-preview | 1065.7 | 73 |
| 7 | anthropic/claude-opus-4.8 | 1052.9 | 61 |
| 8 | x-ai/grok-4.5 | 1043.2 | 67 |
| 9 | x-ai/grok-4.20 | 1025.0 | 58 |
| 10 | z-ai/glm-4.6v | 1018.9 | 48 |
| 11 | deepseek/deepseek-v4-pro | 1014.2 | 48 |
| 12 | deepseek/deepseek-v3.2 | 1010.0 | 62 |
| 13 | qwen/qwen3.7-plus | 975.7 | 64 |
| 14 | minimax/minimax-m3 | 964.6 | 58 |
| 15 | mistralai/mistral-medium-3-5 | 933.2 | 65 |
| 16 | qwen/qwen3.6-plus | 916.9 | 56 |
| 17 | openai/gpt-5.1 | 898.1 | 67 |
| 18 | meta-llama/llama-4-maverick | 873.6 | 57 |
| 19 | x-ai/grok-4.3 | 872.2 | 32 |
BT = Bradley–Terry score over VLM-judge ballots (multi-view condition), fitted within a single method — scores from different methods come from disconnected match pools and aren't comparable. Votes = judge ballots, not human votes.