AI-judge board FUNGI

An LLM writes code (e.g. Blender) that builds the 3D model.

These ranks come from a VLM judge, not human votes. A vision-language model (Sonnet 4.6, multi-view) casts the ballots on this board — no person voted on them. It is a separate, automated surface: its scores are never mixed into the human leaderboard, and the two can disagree. For the ranking humans voted for, see the human-vote LLM procedural (code-gen) board →

LLM procedural (code-gen)

Scores are only comparable within a paradigm.
Rank (UB) Generator BT score Votes
1 z-ai/glm-4.6v 1176.4 14
2 x-ai/grok-4.5 1117.3 28
3 anthropic/claude-sonnet-5 1106.5 26
4 z-ai/glm-5.2 1076.9 26
5 openai/gpt-5.6-sol 1063.2 28
6 x-ai/grok-4.20 1043.0 24
7 moonshotai/kimi-k2.7-code 1011.8 26
8 deepseek/deepseek-v3.2 1003.2 25
9 deepseek/deepseek-v4-pro 997.0 21
10 qwen/qwen3.7-plus 983.2 26
11 anthropic/claude-opus-4.8 964.0 15
12 minimax/minimax-m3 958.9 24
13 openai/gpt-5.1 931.2 20
14 mistralai/mistral-medium-3-5 924.4 26
15 google/gemini-3.1-pro-preview 901.7 23
16 qwen/qwen3.6-plus 897.2 20
17 meta-llama/llama-4-maverick 821.8 24

BT = Bradley–Terry score over VLM-judge ballots (multi-view condition), fitted within a single method — scores from different methods come from disconnected match pools and aren't comparable. Votes = judge ballots, not human votes.