AI-judge board FUNGI

These ranks come from a VLM judge, not human votes. A vision-language model (Sonnet 4.6, multi-view) casts the ballots on this board — no person voted on them. It is a separate, automated surface: its scores are never mixed into the human leaderboard, and the two can disagree. For the ranking humans voted for, see the human-vote leaderboard →

Agentic 3D

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 mistralai/mistral-medium-3-5 (agentic) 1181.7 2
2 openai/gpt-5.1 (agentic) 1118.9 6
3 openai/gpt-5.6-sol-pro (agentic) 1074.5 8
4 google/gemini-3.1-pro-preview (agentic) 1074.3 22
5 x-ai/grok-4.5 (agentic) 1069.0 21
6 moonshotai/kimi-k2.7-code (agentic) 1043.5 14
7 anthropic/claude-opus-4.8 (agentic) 1039.7 14
8 z-ai/glm-4.6v (agentic) 1027.1 12
9 x-ai/grok-4.20 (agentic) 974.5 7
10 minimax/minimax-m3 (agentic) 902.0 2
11 meta-llama/llama-4-maverick (agentic) 901.4 15
12 openai/gpt-5.6-sol (agentic) 891.3 20
13 qwen/qwen3.7-plus (agentic) 887.2 10
14 anthropic/claude-sonnet-5 (agentic) 858.6 9

Image→3D reconstruction

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 TRELLIS 2 1110.0 16
2 Hunyuan3D v3 1099.3 25
3 TRELLIS via fal 1082.5 28
4 SAM 3D 1080.3 16
5 Rodin/Hyper3D 1011.0 28
6 Hunyuan3D 3.1 1010.7 28
7 Pixal3D 977.6 20
8 TRELLIS via Replicate 972.4 28
9 Meshy 6 967.5 22
10 Hunyuan3D v2 853.8 11
11 TripoSR 806.1 28

LLM procedural (code-gen)

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 z-ai/glm-4.6v 1176.4 14
2 x-ai/grok-4.5 1117.3 28
3 anthropic/claude-sonnet-5 1106.5 26
4 z-ai/glm-5.2 1076.9 26
5 openai/gpt-5.6-sol 1063.2 28
6 x-ai/grok-4.20 1043.0 24
7 moonshotai/kimi-k2.7-code 1011.8 26
8 deepseek/deepseek-v3.2 1003.2 25
9 deepseek/deepseek-v4-pro 997.0 21
10 qwen/qwen3.7-plus 983.2 26
11 anthropic/claude-opus-4.8 964.0 15
12 minimax/minimax-m3 958.9 24
13 openai/gpt-5.1 931.2 20
14 mistralai/mistral-medium-3-5 924.4 26
15 google/gemini-3.1-pro-preview 901.7 23
16 qwen/qwen3.6-plus 897.2 20
17 meta-llama/llama-4-maverick 821.8 24

Text→3D (native)

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 Rodin text via Replicate 1333.1 6
2 Tripo H3.1 text 1093.4 28
3 Hunyuan3D 3.1 text 1088.6 16
4 Hunyuan3D v3 text 1008.6 16
5 Tripo P1 text 949.2 28
6 Rodin text via fal 783.6 20
7 Meshy v6 text 751.0 2

BT = Bradley–Terry score over VLM-judge ballots (multi-view condition), fitted within a single method — scores from different methods come from disconnected match pools and aren't comparable. Votes = judge ballots, not human votes.