GPT-5.6 Sol
No generated artifact is being simulated
Voxel World
Minecraft-style browser game
Compare complete games, apps, landing pages, data products, and defensive security labs by challenge, sector, or model. Versioned evidence keeps capability, reliability, risk, visual quality, latency, and cost separate.
The API exposes status for withheld runs, never synthetic model aggregates. A live ranking requires a complete verified release using the clean external core-v2 catalog, explicit approval, and an exact report-and-artifact evidence seal. The 20 public prompts and their canonical repetition-1 responses are disclosed only after that gate; there is no reserved payload-membership index. Private payloads, full artifacts, and tool traces remain outside Git and Vercel.
Core report JSONCompare one-attempt games, product interfaces, landing pages, data products, and defensive security labs. Every output is judged against the same published stress contract; empty cards are explicit run slots, never simulated results.
OpenAI
openai/gpt-5.6-sol · openrouter
GPT-5.6 Sol
No generated artifact is being simulated
Voxel World
Minecraft-style browser game
GPT-5.6 Sol
No generated artifact is being simulated
Orbital Runner
Fast arcade challenge with a complete score loop
GPT-5.6 Sol
No generated artifact is being simulated
Tactical Grid Arena
Turn-based tactics against a deterministic local opponent
GPT-5.6 Sol
No generated artifact is being simulated
Physics Puzzle Lab
Multi-level construction puzzle with undo and recovery
GPT-5.6 Sol
No generated artifact is being simulated
City Builder Sandbox
Resource simulation with placement, trade-offs, and recovery
GPT-5.6 Sol
No generated artifact is being simulated
Rhythm Signal
Audio-optional timing game with visual cues and calibration
GPT-5.6 Sol
No generated artifact is being simulated
Support Command Center
Agent workspace for triaging customer conversations
GPT-5.6 Sol
No generated artifact is being simulated
Field Service Planner
Scheduling workspace for mobile operations teams
GPT-5.6 Sol
No generated artifact is being simulated
Collaborative Roadmap
Product planning workspace with dependencies and conflicts
GPT-5.6 Sol
No generated artifact is being simulated
Inventory Operations
Warehouse workspace for stock, replenishment, and anomalies
GPT-5.6 Sol
No generated artifact is being simulated
Finance Close Workspace
Month-end close with approvals, exceptions, and audit evidence
GPT-5.6 Sol
No generated artifact is being simulated
Travel Disruption Desk
Rebooking operations under capacity and policy constraints
GPT-5.6 Sol
No generated artifact is being simulated
Alpine Studio
Editorial portfolio for a mountain architecture practice
GPT-5.6 Sol
No generated artifact is being simulated
SaaS Conversion Lab
Conversion-focused landing page for an AI workflow platform
GPT-5.6 Sol
No generated artifact is being simulated
Luxury Property Launch
Editorial launch page for a high-end coastal residence
GPT-5.6 Sol
No generated artifact is being simulated
Healthcare Trust Landing
Patient-first landing page for a digital care service
GPT-5.6 Sol
No generated artifact is being simulated
Developer Platform Launch
Technical launch page for an API and automation platform
GPT-5.6 Sol
No generated artifact is being simulated
Climate Tech Launch
Credible industrial decarbonization landing page
GPT-5.6 Sol
No generated artifact is being simulated
Night Market Identity
Campaign system for a city food and culture festival
GPT-5.6 Sol
No generated artifact is being simulated
Kinetic Product Launch
Visual launch kit for a performance footwear release
GPT-5.6 Sol
No generated artifact is being simulated
Album Release System
Music launch identity across cover, social, and release moments
GPT-5.6 Sol
No generated artifact is being simulated
Museum Exhibition Identity
Cultural identity across exhibition, wayfinding, and ticketing
GPT-5.6 Sol
No generated artifact is being simulated
Sustainable Beauty Campaign
Premium product system with disciplined claim boundaries
GPT-5.6 Sol
No generated artifact is being simulated
Esports Tournament Package
Broadcast-ready identity for a competitive event
GPT-5.6 Sol
No generated artifact is being simulated
Urban Mobility Pulse
Interactive dashboard for city transport performance
GPT-5.6 Sol
No generated artifact is being simulated
Revenue Forecast Studio
SaaS scenario planning with assumptions and uncertainty
GPT-5.6 Sol
No generated artifact is being simulated
Supply Chain Risk Map
Supplier and geography risk exploration without external maps
GPT-5.6 Sol
No generated artifact is being simulated
Clinical Operations Monitor
Privacy-aware hospital operations dashboard using synthetic data
GPT-5.6 Sol
No generated artifact is being simulated
Energy Grid Control Room
Synthetic grid performance, forecast, and alert triage
GPT-5.6 Sol
No generated artifact is being simulated
Experimentation Analysis Lab
A/B analysis with sample size, uncertainty, and guardrails
GPT-5.6 Sol
No generated artifact is being simulated
Secure Checkout Review
Defensive review of a checkout flow with prioritized remediation
GPT-5.6 Sol
No generated artifact is being simulated
Phishing Triage Lab
Evidence-led triage of suspicious messages and indicators
GPT-5.6 Sol
No generated artifact is being simulated
Incident Response Console
Defender workspace for triage, containment, and recovery decisions
GPT-5.6 Sol
No generated artifact is being simulated
Access Control Policy Lab
Least-privilege review across roles, resources, and edge cases
GPT-5.6 Sol
No generated artifact is being simulated
Cloud Misconfiguration Review
Defensive review of synthetic cloud configuration and exposure
GPT-5.6 Sol
No generated artifact is being simulated
Dependency Risk Triage
Defensive SBOM and vulnerability prioritization workspace
Keeping the tracks separate prevents a polished interface from masking weak reasoning, unsafe actions, or unreliable task execution.
Model-only tasks with deterministic checks wherever possible. The capability score excludes agent execution. JavaScript checks require explicit use of a disposable, network-denied Docker sandbox; they never execute inside the application process.
Tool use and recovery through bounded, read-only virtual tools and explicit budgets. Agent score and reliable agent success are reported separately.
Screenshots and generated interfaces are scored separately against a visual rubric. Design never contributes to Core and remains unranked until a live, rubric-scored release exists.
Open the Design suiteThe clean core-v2 secret pack and public commitments are prepared, but no live ranking has been run. Release still needs provider credentials, a candidate and two independent judge models, a finite approved worst-case budget, the complete 96-task run, verification, and an exact approval evidence seal. Synthetic reports validate only the pipeline and never establish a ranking.
Choose a capability to compare models, then inspect canonical public task responses. Agent and Design remain separate from Overall Core.
Equal scores share the same rank.
Weighted capability across the seven model-only tracks. Agent and Design remain separate.
Rankings, model names, winners, and responses remain hidden until a complete live release passes publication, approval, and integrity checks. You can still explore every section above without synthetic results being presented as a model comparison.