AI-judge board ANIMALS

These ranks come from a VLM judge, not human votes. A vision-language model (Sonnet 4.6, multi-view) casts the ballots on this board — no person voted on them. It is a separate, automated surface: its scores are never mixed into the human leaderboard, and the two can disagree. For the ranking humans voted for, see the human-vote leaderboard →

Agentic 3D

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 google/gemini-3.1-pro-preview (agentic) 1240.0 10
2 z-ai/glm-4.6v (agentic) 1154.4 6
3 openai/gpt-5.6-sol-pro (agentic) 1137.8 12
4 anthropic/claude-sonnet-5 (agentic) 1118.0 12
5 openai/gpt-5.6-sol (agentic) 1096.9 16
6 x-ai/grok-4.5 (agentic) 1062.8 12
7 x-ai/grok-4.20 (agentic) 996.3 12
8 qwen/qwen3.7-plus (agentic) 991.7 12
9 anthropic/claude-opus-4.8 (agentic) 984.0 10
10 minimax/minimax-m3 (agentic) 952.2 8
11 openai/gpt-5.1 (agentic) 850.4 10
12 meta-llama/llama-4-maverick (agentic) 798.5 8
13 qwen/qwen3.6-plus (agentic) 790.4 4
14 moonshotai/kimi-k2.7-code (agentic) 771.5 4
15 mistralai/mistral-medium-3-5 (agentic) 703.1 4

Image→3D reconstruction

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 Meshy 6 1223.5 12
2 Hunyuan3D v3 1186.3 16
3 Rodin/Hyper3D 1096.0 16
4 TRELLIS via Replicate 1042.7 8
5 SAM 3D 1017.4 8
6 TRELLIS via fal 1006.2 16
7 Hunyuan3D 3.1 1000.6 14
8 Hunyuan3D v2 860.8 16
9 Pixal3D 814.6 12
10 TRELLIS 2 773.7 6

LLM procedural (code-gen)

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 openai/gpt-5.6-sol 1249.8 16
2 anthropic/claude-opus-4.8 1183.6 10
3 moonshotai/kimi-k2.7-code 1165.8 14
4 z-ai/glm-5.2 1138.3 14
5 anthropic/claude-sonnet-5 1123.4 16
6 deepseek/deepseek-v3.2 1051.8 15
7 google/gemini-3.1-pro-preview 1042.1 14
8 z-ai/glm-4.6v 1026.6 14
9 deepseek/deepseek-v4-pro 1011.8 13
10 x-ai/grok-4.5 980.6 16
11 qwen/qwen3.7-plus 979.4 16
12 x-ai/grok-4.20 930.3 16
13 mistralai/mistral-medium-3-5 922.6 15
14 minimax/minimax-m3 907.7 15
15 meta-llama/llama-4-maverick 905.2 13
16 openai/gpt-5.1 863.9 14
17 qwen/qwen3.6-plus 829.4 13

Text→3D (native)

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 Tripo P1 text 1075.3 12
2 Rodin text via Replicate 1056.0 16
3 Rodin text via fal 1042.8 12
4 Hunyuan3D v3 text 1012.8 16
5 Tripo H3.1 text 984.0 12
6 Hunyuan3D 3.1 text 842.2 8

BT = Bradley–Terry score over VLM-judge ballots (multi-view condition), fitted within a single method — scores from different methods come from disconnected match pools and aren't comparable. Votes = judge ballots, not human votes.