Leaderboard PLANTS
An LLM writes code (e.g. Blender) that builds the 3D model.
Human voting is currently focused on the commercial image→3D and text→3D models, so these scores are paused where the votes so far left them. The AI-judge board still ranks this modality.
Ranked by human votes · 340 cast · see the AI-judge board →
How ranking works
Rank groups models whose 95% bootstrap CIs overlap into the same rank (they are not statistically separable). BT = Bradley–Terry score (Elo-scaled); the bar shows the 95% CI with the point estimate marked. Every board ranks a SINGLE method — scores from different methods come from disconnected match pools and aren't comparable, so there is no cross-method ranking. Elo updates live per vote. Recompute in Admin.
LLM procedural (code-gen)
16 generators| Rank | Generator | BT score | 95% interval BT -181.1–2557.4 | Trend | Votes | Status |
|---|---|---|---|---|---|---|
| 1 | openai/gpt-5.6-sol-pro | 2111.6 | [1095.8, 2557.4] | 1 | 29 more votes → firm ► | |
| openai/gpt-5.6-sol | 1412.5 | [875.7, 1911.7] | 7 | 23 more votes → firm ► | ||
| 1 | x-ai/grok-4.5 | 1695.1 | [698.5, 2197.4] | 3 | 27 more votes → firm ► | |
| 1 | openai/gpt-5.1 | 1309.1 | [1025.6, 1645.0] | 1 | 29 more votes → firm ► | |
| 1 | anthropic/claude-opus-4.8 | 1181.8 | [647.0, 1822.7] | 6 | 24 more votes → firm ► | |
| 1 | google/gemini-3.1-pro-preview | 1159.4 | [578.9, 1896.6] | 6 | 24 more votes → firm ► | |
| 1 | qwen/qwen3.7-plus | 1106.7 | [464.6, 1575.4] | 2 | 28 more votes → firm ► | |
| 1 | x-ai/grok-4.3 | 1061.9 | [375.0, 1735.9] | 3 | 27 more votes → firm ► | |
| 1 | anthropic/claude-sonnet-5 | 1047.0 | [548.2, 1443.9] | 6 | 24 more votes → firm ► | |
| 1 | z-ai/glm-5.2 | 986.1 | [495.1, 1404.3] | 10 | 20 more votes → firm ► | |
| 1 | x-ai/grok-4.20 | 892.6 | [476.8, 1302.7] | 2 | 28 more votes → firm ► | |
| 1 | moonshotai/kimi-k2.7-code | 861.2 | [263.7, 1481.5] | 4 | 26 more votes → firm ► | |
| 1 | deepseek/deepseek-v4-pro | 800.8 | [346.8, 1271.1] | 3 | 27 more votes → firm ► | |
| 1 | mistralai/mistral-medium-3-5 | 646.9 | [139.6, 1391.6] | 4 | 26 more votes → firm ► | |
| 1 | deepseek/deepseek-v3.2 | 640.0 | [119.8, 1139.0] | 3 | 27 more votes → firm ► | |
| 3 | z-ai/glm-4.6v | 476.0 | [-48.3, 820.2] | 5 | 25 more votes → firm ► | |
| 1 | meta-llama/llama-4-maverick | 412.3 | [-181.1, 1151.9] | 2 | 28 more votes → firm ► |
95% credible interval · point estimate · trend · firm a rank backed by 30+ votes; below that, Status counts the votes still needed · click a row for detail
Filters & bias audit
Bias audit — left(A) win rate 0.569 (≈0.50 = unbiased) · tie 0.059 · bad 0.191
VLM judge (Sonnet 4.6, multi-view) — automated LLM-judge rankings by paradigm
Loading automated rankings…