These ranks come from a VLM judge, not human votes.
A vision-language model (Sonnet 4.6, multi-view) casts the ballots on this
board — no person voted on them. It is a separate, automated surface: its
scores are never mixed into the human leaderboard, and the two can disagree.
For the ranking humans voted for, see the
human-vote leaderboard →
Agentic 3D
Scores are only comparable within a paradigm.
Rank (UB)
Generator
BT score
Votes
1
openai/gpt-5.6-sol (agentic)
1152.7
20
2
x-ai/grok-4.5 (agentic)
1118.2
16
3
google/gemini-3.1-pro-preview (agentic)
1070.5
14
4
anthropic/claude-opus-4.8 (agentic)
1062.8
14
5
qwen/qwen3.7-plus (agentic)
1013.7
10
6
x-ai/grok-4.20 (agentic)
1003.4
16
7
moonshotai/kimi-k2.7-code (agentic)
983.8
7
8
anthropic/claude-sonnet-5 (agentic)
966.7
8
9
qwen/qwen3.6-plus (agentic)
927.7
3
10
z-ai/glm-4.6v (agentic)
892.2
16
11
meta-llama/llama-4-maverick (agentic)
883.5
12
12
openai/gpt-5.1 (agentic)
870.7
12
13
mistralai/mistral-medium-3-5 (agentic)
723.5
2
Image→3D reconstruction
Scores are only comparable within a paradigm.
Rank (UB)
Generator
BT score
Votes
1
TRELLIS 2
1205.8
18
2
Hunyuan3D 3.1
1155.0
60
3
Meshy 6
1143.5
18
4
Hunyuan3D v3
1084.8
52
5
TRELLIS via fal
1001.3
60
6
Hunyuan3D v2
995.3
40
7
Rodin/Hyper3D
986.3
60
8
InstantMesh
959.4
8
9
TRELLIS via Replicate
929.7
40
10
SAM 3D
911.6
18
11
Pixal3D
909.0
18
12
TripoSR
861.6
56
LLM procedural (code-gen)
Scores are only comparable within a paradigm.
Rank (UB)
Generator
BT score
Votes
1
openai/gpt-5.6-sol-pro
1341.9
4
2
openai/gpt-5.6-sol
1210.1
24
3
google/gemini-3.1-pro-preview
1189.8
36
4
z-ai/glm-5.2
1132.7
34
5
x-ai/grok-4.20
1074.9
18
6
anthropic/claude-opus-4.8
1048.6
36
7
moonshotai/kimi-k2.7-code
1035.0
22
8
anthropic/claude-sonnet-5
1026.4
18
9
deepseek/deepseek-v4-pro
1018.5
14
10
minimax/minimax-m3
997.3
19
11
deepseek/deepseek-v3.2
971.8
22
12
x-ai/grok-4.5
956.3
23
13
qwen/qwen3.6-plus
947.3
23
14
mistralai/mistral-medium-3-5
943.1
24
15
qwen/qwen3.7-plus
932.8
22
16
meta-llama/llama-4-maverick
894.9
20
17
z-ai/glm-4.6v
893.7
20
18
x-ai/grok-4.3
883.3
32
19
openai/gpt-5.1
879.1
33
Text→3D (native)
Scores are only comparable within a paradigm.
Rank (UB)
Generator
BT score
Votes
1
Tripo H3.1 text
1179.2
55
2
Hunyuan3D v3 text
1047.5
55
3
Hunyuan3D 3.1 text
1044.1
57
4
Tripo P1 text
998.4
57
5
Rodin text via fal
935.1
58
6
Rodin text via Replicate
831.5
50
BT = Bradley–Terry score over VLM-judge ballots (multi-view condition), fitted
within a single method — scores from different methods come from disconnected
match pools and aren't comparable. Votes = judge ballots, not human votes.