AI-judge board FUNGI
An LLM writes code (e.g. Blender) that builds the 3D model.
These ranks come from a VLM judge, not human votes. A vision-language model (Sonnet 4.6, multi-view) casts the ballots on this board — no person voted on them. It is a separate, automated surface: its scores are never mixed into the human leaderboard, and the two can disagree. For the ranking humans voted for, see the human-vote LLM procedural (code-gen) board →
LLM procedural (code-gen)
Scores are only comparable within a paradigm.| Rank (UB) | Generator | BT score | Votes |
|---|---|---|---|
| 1 | z-ai/glm-4.6v | 1176.4 | 14 |
| 2 | x-ai/grok-4.5 | 1117.3 | 28 |
| 3 | anthropic/claude-sonnet-5 | 1106.5 | 26 |
| 4 | z-ai/glm-5.2 | 1076.9 | 26 |
| 5 | openai/gpt-5.6-sol | 1063.2 | 28 |
| 6 | x-ai/grok-4.20 | 1043.0 | 24 |
| 7 | moonshotai/kimi-k2.7-code | 1011.8 | 26 |
| 8 | deepseek/deepseek-v3.2 | 1003.2 | 25 |
| 9 | deepseek/deepseek-v4-pro | 997.0 | 21 |
| 10 | qwen/qwen3.7-plus | 983.2 | 26 |
| 11 | anthropic/claude-opus-4.8 | 964.0 | 15 |
| 12 | minimax/minimax-m3 | 958.9 | 24 |
| 13 | openai/gpt-5.1 | 931.2 | 20 |
| 14 | mistralai/mistral-medium-3-5 | 924.4 | 26 |
| 15 | google/gemini-3.1-pro-preview | 901.7 | 23 |
| 16 | qwen/qwen3.6-plus | 897.2 | 20 |
| 17 | meta-llama/llama-4-maverick | 821.8 | 24 |
BT = Bradley–Terry score over VLM-judge ballots (multi-view condition), fitted within a single method — scores from different methods come from disconnected match pools and aren't comparable. Votes = judge ballots, not human votes.