AI-judge board ALL KINGDOMS

These ranks come from a VLM judge, not human votes. A vision-language model (Sonnet 4.6, multi-view) casts the ballots on this board — no person voted on them. It is a separate, automated surface: its scores are never mixed into the human leaderboard, and the two can disagree. For the ranking humans voted for, see the human-vote leaderboard →

Agentic 3D

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 openai/gpt-5.6-sol-pro (agentic) 1139.7 20
2 google/gemini-3.1-pro-preview (agentic) 1105.9 46
3 x-ai/grok-4.5 (agentic) 1093.4 49
4 anthropic/claude-opus-4.8 (agentic) 1058.1 38
5 openai/gpt-5.6-sol (agentic) 1053.9 56
6 anthropic/claude-sonnet-5 (agentic) 1002.1 29
7 x-ai/grok-4.20 (agentic) 999.7 35
8 moonshotai/kimi-k2.7-code (agentic) 990.2 25
9 z-ai/glm-4.6v (agentic) 979.5 34
10 minimax/minimax-m3 (agentic) 970.4 10
11 qwen/qwen3.7-plus (agentic) 948.7 32
12 openai/gpt-5.1 (agentic) 919.1 28
13 meta-llama/llama-4-maverick (agentic) 882.7 35
14 mistralai/mistral-medium-3-5 (agentic) 870.2 8
15 qwen/qwen3.6-plus (agentic) 829.0 7

Image→3D reconstruction

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 Hunyuan3D v3 1114.9 93
2 TRELLIS 2 1103.3 40
3 Hunyuan3D 3.1 1102.9 102
4 Meshy 6 1073.7 52
5 Rodin/Hyper3D 1027.0 104
6 TRELLIS via fal 1026.8 104
7 SAM 3D 995.0 42
8 TRELLIS via Replicate 967.4 76
9 InstantMesh 955.2 8
10 Hunyuan3D v2 952.4 67
11 Pixal3D 920.9 50
12 TripoSR 844.4 84

LLM procedural (code-gen)

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 openai/gpt-5.6-sol-pro 1346.5 4
2 openai/gpt-5.6-sol 1163.8 68
3 z-ai/glm-5.2 1110.7 74
4 anthropic/claude-sonnet-5 1096.8 60
5 moonshotai/kimi-k2.7-code 1073.7 62
6 google/gemini-3.1-pro-preview 1065.7 73
7 anthropic/claude-opus-4.8 1052.9 61
8 x-ai/grok-4.5 1043.2 67
9 x-ai/grok-4.20 1025.0 58
10 z-ai/glm-4.6v 1018.9 48
11 deepseek/deepseek-v4-pro 1014.2 48
12 deepseek/deepseek-v3.2 1010.0 62
13 qwen/qwen3.7-plus 975.7 64
14 minimax/minimax-m3 964.6 58
15 mistralai/mistral-medium-3-5 933.2 65
16 qwen/qwen3.6-plus 916.9 56
17 openai/gpt-5.1 898.1 67
18 meta-llama/llama-4-maverick 873.6 57
19 x-ai/grok-4.3 872.2 32

Text→3D (native)

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 Tripo H3.1 text 1143.3 95
2 Hunyuan3D v3 text 1053.9 87
3 Hunyuan3D 3.1 text 1023.5 81
4 Tripo P1 text 996.0 97
5 Rodin text via Replicate 928.0 72
6 Rodin text via fal 923.4 90
7 Meshy v6 text 783.8 2

BT = Bradley–Terry score over VLM-judge ballots (multi-view condition), fitted within a single method — scores from different methods come from disconnected match pools and aren't comparable. Votes = judge ballots, not human votes.