GPT-5.6 Sol
No generated artifact is being simulated
GPT-5.6 Sol
OpenAI
Compare complete games, apps, landing pages, data products, and defensive security labs by challenge, sector, or model. Versioned evidence keeps capability, reliability, risk, visual quality, latency, and cost separate.
The API exposes status for withheld runs, never synthetic model aggregates. A live ranking requires a complete verified release using the clean external core-v2 catalog, explicit approval, and an exact report-and-artifact evidence seal. The 20 public prompts and their canonical repetition-1 responses are disclosed only after that gate; there is no reserved payload-membership index. Private payloads, full artifacts, and tool traces remain outside Git and Vercel.
Core report JSONCompare one-attempt games, product interfaces, landing pages, data products, and defensive security labs. Every output is judged against the same published stress contract; empty cards are explicit run slots, never simulated results.
Games
extreme stressMinecraft-style browser game
Gameplay completeness
35pts
Interaction quality
20pts
Runtime resilience
15pts
Visual craft
15pts
Accessibility and responsiveness
15pts
Evaluation gates
GPT-5.6 Sol
No generated artifact is being simulated
GPT-5.6 Sol
OpenAI
GPT-5.6 Terra
No generated artifact is being simulated
GPT-5.6 Terra
OpenAI
Claude Opus 5
No generated artifact is being simulated
Claude Opus 5
Anthropic
Claude Sonnet 5
No generated artifact is being simulated
Claude Sonnet 5
Anthropic
Gemini 3.1 Pro Preview
No generated artifact is being simulated
Gemini 3.1 Pro Preview
Gemini 3.6 Flash
No generated artifact is being simulated
Gemini 3.6 Flash
Grok 4.5
No generated artifact is being simulated
Grok 4.5
xAI
DeepSeek V4 Pro
No generated artifact is being simulated
DeepSeek V4 Pro
DeepSeek
Qwen 3.7 Max
No generated artifact is being simulated
Qwen 3.7 Max
Qwen
Llama 4 Maverick
No generated artifact is being simulated
Llama 4 Maverick
Meta
Mistral Large 2512
No generated artifact is being simulated
Mistral Large 2512
Mistral AI
Keeping the tracks separate prevents a polished interface from masking weak reasoning, unsafe actions, or unreliable task execution.
Model-only tasks with deterministic checks wherever possible. The capability score excludes agent execution. JavaScript checks require explicit use of a disposable, network-denied Docker sandbox; they never execute inside the application process.
Tool use and recovery through bounded, read-only virtual tools and explicit budgets. Agent score and reliable agent success are reported separately.
Screenshots and generated interfaces are scored separately against a visual rubric. Design never contributes to Core and remains unranked until a live, rubric-scored release exists.
Open the Design suiteThe clean core-v2 secret pack and public commitments are prepared, but no live ranking has been run. Release still needs provider credentials, a candidate and two independent judge models, a finite approved worst-case budget, the complete 96-task run, verification, and an exact approval evidence seal. Synthetic reports validate only the pipeline and never establish a ranking.
Choose a capability to compare models, then inspect canonical public task responses. Agent and Design remain separate from Overall Core.
Equal scores share the same rank.
Weighted capability across the seven model-only tracks. Agent and Design remain separate.
Rankings, model names, winners, and responses remain hidden until a complete live release passes publication, approval, and integrity checks. You can still explore every section above without synthetic results being presented as a model comparison.