LLM Build Leaderboard

Each game on our roadmaps is built with multiple LLMs. This page tracks how each model performs — speed, quality, cost, and bugs — so you can pick the right model for your next build.

27 benchmark runs
ScoreModelProviderGameTimeTokensCostBugsNotes
95Mistral SmallMistral APIAI Dodge12s1,754$00Cleanest output — zero bugs, 12s, all features working.
95DeepSeek V4 ProDeepSeek API3D Space Dodge2s8,335$0.07500Fastest build (~2s). Largest output. Bot mode. Premium cost ($0.075).
92DeepSeek V4 Flash (0.4 temp)DeepSeek API3D Space Dodge30s8,000$0.00100Two prompts. Complete game + bot mode on first attempt. 680 lines.
91Tencent Hy3 PreviewOpenRouter3D Space Dodge5s7,459$00Both prompts. Bot mode. Large output (16.1 KB). Fast free-tier inference.
90DeepSeek V4 Flash (0.7 temp)DeepSeek APIAI Dodge68s8,399$0.00120Largest output (429 lines, 15,487 chars). Full feature set.
90Xiaomi Mimo V2.5OpenRouter3D Space Dodge2s7,970$0.00500Both prompts. Bot mode. 13.3 KB output. Competitive performance ($0.005).
90DeepSeek V4 FlashDeepSeek APIMemory Match45s7,300$0.00100First PixiJS build. Bot mode with perfect-memory AI. Sound effects and leaderboard.
90DeepSeek V4 FlashDeepSeek APISliding Puzzle68s5,300$0.00080Vanilla Canvas 15-puzzle. BFS bot solver, keyboard support, solvability check. Zero-dependency game.
89Nemotron 3 Ultra (550B)OpenRouter3D Space Dodge2s7,350$0.00300550B model. ~2s inference. Bot mode working. $0.003 total.
88owl-alphaOpenRouter (free)AI Dodge53s2,024$00Best free-tier output (8,272 chars). No bugs.
88Mistral SmallMistral API3D Space Dodge21s6,689$00Two prompts. Clean Three.js output. Bot mode working. Fast build.
88DeepSeek V4 FlashDeepSeek API3D Ball Balance45s8,300$0.00100Three.js tilt-platform game. Full physics simulation, coin collection, bot mode, glow effects. 329 lines.
87DeepSeek V4 FlashDeepSeek APIAsteroids12s8,200$0.00120Arena match winner (8.73/10 avg). Momentum physics, asteroid splitting, screen wrapping, wave progression, mobile touch controls. 609 lines.
85owl-alphaOpenRouter (free)3D Space Dodge6s8,913$00Two prompts. Fastest build (6s total). Bot mode working.
82DeepSeek V4 Flash (0.4 temp)DeepSeek APIAI Dodge14s2,077$0.00053Interactive build across 5 prompts. Fastest cloud time.
82Gemma-4-31BOpenRouter (free tier)3D Space Dodge2s4,032$00Both prompts. Bot mode working. 31B model, fast free-tier inference.
80poolside/laguna-m.1OpenRouter (free tier)3D Space Dodge2s7,796$00PASS after fix: Model used MeshBasicMaterial (no emissive) instead of MeshStandardMaterial. One-word fix. Bot mode working.
78openai/gpt-oss-20bLocal — Mac Mini M4AI Dodge23s1,823$02Best local model. Fast (23s), functional. Speed + double-R bugs.
76Nemotron 3 UltraOpenRouter (free)AI Dodge86s2,015$01Most compact (132 lines). May lack edge wrapping.
74Qwen 3.5 9BLocal — Mac Mini M4AI Dodge190s3,210$01Best local code quality. Clean structure, few bugs.
72Llama 3.1 8BLocal — Mac Mini M4AI Dodge55s1,220$01Most reliable local fallback. Sparse but functional.
70Gemma-4-12b-qatLocal — Mac Mini M4AI Dodge281s3,563$00Slowest by far (281s). Output adequate but not proportional to time.
70Gemma-4-12b-qatLocal — Apple Mac Mini M4 (16GB)3D Space Dodge691s8,431$00Slowest benchmark (11.5 min). Both prompts succeeded. Bot mode working.
68Gemma-4-12b-coder-fableLocal — Mac Mini M4AI Dodge112s1,463$01Shortest working output (3,451 chars). Missing some features.
68Qwen 3.5 9BLocal — Apple Mac Mini M4 (16GB)3D Space Dodge462s5,435$06DEGRADED: 6 const reassignment errors at runtime. WebGL renders (680x480). Game runs but errors degrade gameplay. 40% tokens on CoT reasoning.
65DeepSeek V4 ProDeepSeek APIAI Dodge101s8,399$0.00120Spent tokens on reasoning. Only 50 lines actual game code.
60GPT-OSS-120BOpenRouter (free tier)3D Space Dodge2s3,292$00Prompt 1 OK. Prompt 2 failed — no bot mode, glow, or boundary. Instruction gap.