Coverage & governance FUNGI

How much evidence backs each ranking — a generator's rank is only as trustworthy as the votes behind it, and confidence firms up as vote counts grow. Published in full so a rank can be read with its uncertainty, not in isolation.

107 Agentic 3D
130 Image→3D reconstruction
185 LLM procedural (code-gen)
99 Text→3D (native)

Generator coverage — 53 generators with arena outputs, by total Mode-A votes

GeneratorTasksOutputsA-votesConfidenceIn arena
Tripo H3.1 text model 16 16 242 firm ✓
Hunyuan3D v3 text model 16 16 241 firm ✓
Tripo P1 text model 15 15 229 firm ✓
Meshy v6 text model 14 14 219 firm ✓
Rodin text via fal model 14 14 218 firm ✓
Hunyuan3D 3.1 text model 13 13 195 firm ✓
SAM 3D model 15 16 158 firm ✓
Rodin text via Replicate model 11 11 145 firm ✓
Meshy 6 model 11 11 142 firm ✓
Pixal3D model 11 11 125 firm ✓
Hunyuan3D v3 model 12 15 115 firm ✓
Hunyuan3D 3.1 model 12 15 112 firm ✓
Hunyuan3D v2 model 11 12 95 firm ✓
Rodin/Hyper3D model 11 15 91 firm ✓
TRELLIS via Replicate model 11 12 75 firm ✓
TRELLIS via fal model 8 9 62 firm ✓
TRELLIS 2 model 6 6 57 firm ✓
TripoSR model 7 7 42 firm ✓
anthropic/claude-sonnet-5 model 13 13 16 provisional ✓
deepseek/deepseek-v3.2 model 14 14 16 provisional ✓
anthropic/claude-opus-4.8 (agentic) model 11 11 15 provisional ✓
openai/gpt-5.6-sol model 14 14 15 provisional ✓
x-ai/grok-4.20 model 14 14 15 provisional ✓
google/gemini-3.1-pro-preview (agentic) model 12 12 14 provisional ✓
openai/gpt-5.6-sol (agentic) model 12 12 14 provisional ✓
anthropic/claude-opus-4.8 model 12 12 12 provisional ✓
mistralai/mistral-medium-3-5 model 12 12 12 provisional ✓
z-ai/glm-5.2 model 12 12 12 provisional ✓
anthropic/claude-sonnet-5 (agentic) model 9 9 11 provisional ✓
deepseek/deepseek-v4-pro model 10 10 10 provisional ✓
moonshotai/kimi-k2.7-code model 10 10 10 provisional ✓
qwen/qwen3.7-plus model 13 13 10 provisional ✓
x-ai/grok-4.5 model 13 13 10 provisional ✓
x-ai/grok-4.5 (agentic) model 9 9 10 provisional ✓
z-ai/glm-4.6v (agentic) model 8 8 10 provisional ✓
meta-llama/llama-4-maverick model 6 6 9 provisional ✓
google/gemini-3.1-pro-preview model 11 11 8 provisional ✓
qwen/qwen3.7-plus (agentic) model 8 8 8 provisional ✓
x-ai/grok-4.20 (agentic) model 7 7 8 provisional ✓
z-ai/glm-4.6v model 8 8 8 provisional ✓
minimax/minimax-m3 model 8 8 7 provisional ✓
minimax/minimax-m3 (agentic) model 6 6 7 provisional ✓
openai/gpt-5.1 (agentic) model 6 6 7 provisional ✓
meta-llama/llama-4-maverick (agentic) model 6 6 6 provisional ✓
mistralai/mistral-medium-3-5 (agentic) model 4 4 5 provisional ✓
openai/gpt-5.6-sol-pro (agentic) model 4 4 5 provisional ✓
moonshotai/kimi-k2.7-code (agentic) model 4 4 4 provisional ✓
qwen/qwen3.6-plus model 4 4 4 provisional ✓
x-ai/grok-4.3 model 4 4 4 provisional ✓
openai/gpt-5.1 model 6 6 3 provisional ✓
InstantMesh model 1 1 1 provisional ✓
openai/gpt-5.6-sol-pro model 1 1 1 provisional ✓
qwen/qwen3.6-plus (agentic) model 1 1 1 provisional ✓

How models enter & compete

  • Entry — every output is added by an administrator or via the public submission queue, then moderated before it enters the arena. There is no private or preferential pre-testing: a model is not quietly trialled and published only if it scores well.
  • Fair matchmaking — pairs are drawn by a sampler that biases toward the least-compared outputs, so coverage spreads evenly rather than concentrating votes on a favoured few. Every model is drawn from the same pool by the same sampler.
  • No silent deprecation — outputs are not removed to flatter the board; the arena has no deprecation mechanism, so what competed stays on the record.
  • Exclusions are principled, not selective — only reference ground-truth scans and untextured geometry-only outputs are held out of the Mode-A perceptual pool (they confound a visual-preference vote); they remain in the objective Mode-B board. Excluded generators are flagged in the table above.

Reading a rank

A generator's rank is only as trustworthy as the votes behind it. Below 30 total Mode-A votes a rank is marked provisional; at or above it, firm. Bradley–Terry confidence intervals on the leaderboard make the same uncertainty visible per generator.

Current limits, stated plainly: the arena is in an internal evaluation phase, so vote volume is low and most ranks read provisional — ranks will firm up as voting scales. Mode-A votes here are internal and not yet from a public pool.

Task coverage — 6 active tasks

"Mode-B" = objective scoring against held-out ground-truth is available for that task.

TaskCategoryTierGeneratorsOutputs A-votesJudge votesMode-BMode-C
Morchella esculenta — single-image → 3D reconstruction Fungi moderate 29 29 105 0 ✓ —
Amanita muscaria — single-image → 3D reconstruction Fungi moderate 28 28 86 0 ✓ —
Hericium erinaceus — single-image → 3D reconstruction Fungi hard 26 28 75 0 ✓ —
Boletus edulis — single-image → 3D reconstruction Fungi easy 26 26 91 0 ✓ —
Lycoperdon perlatum — single-image → 3D reconstruction Fungi easy 26 26 86 0 ✓ —
Trametes versicolor — single-image → 3D reconstruction Fungi hard 17 17 76 0 ✓ —

Mode-C — botanical-trait accuracy experimental

Generator-level mean botanical accuracy — a 3D model graded against a literature-sourced per-taxon trait rubric by a calibrated VLM trait-checker. Only trait classes that have passed the human↔VLM agreement gate count toward the score. No class has passed the gate yet, so this axis is experimental: shown for inspection, not yet a ranking signal.

No scored outputs yet. Per-output scorecards live at /trait/<id>.

Machine-readable: /api/coverage.json.