CresciBench Core

See what every model builds under pressure

Compare complete games, apps, landing pages, data products, and defensive security labs by challenge, sector, or model. Versioned evidence keeps capability, reliability, risk, visual quality, latency, and cost separate.

Public evidence contract

The API exposes status for withheld runs, never synthetic model aggregates. A live ranking requires a complete verified release using the clean external core-v2 catalog, explicit approval, and an exact report-and-artifact evidence seal. The 20 public prompts and their canonical repetition-1 responses are disclosed only after that gate; there is no reserved payload-membership index. Private payloads, full artifacts, and tool traces remain outside Git and Vercel.

Core report JSON
Evidence explorer

See what every model builds under pressure

Compare one-attempt games, product interfaces, landing pages, data products, and defensive security labs. Every output is judged against the same published stress contract; empty cards are explicit run slots, never simulated results.

Models
11
Stress tests
36
Output slots
396
Published
0
  • One attempt
  • 12k token cap
  • Offline artifact
  • Responsive
  • Failure states
  • 100-point rubric
Core release: emptyOpen official leaderboard

OpenAI

GPT-5.6 Sol

openai/gpt-5.6-sol · openrouter

Input
$5/ 1M
Output
$30/ 1M
textreasoningcodevisiontool-use
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Voxel World

Minecraft-style browser game

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

high stress

Orbital Runner

Fast arcade challenge with a complete score loop

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Tactical Grid Arena

Turn-based tactics against a deterministic local opponent

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Physics Puzzle Lab

Multi-level construction puzzle with undo and recovery

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

City Builder Sandbox

Resource simulation with placement, trade-offs, and recovery

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

high stress

Rhythm Signal

Audio-optional timing game with visual cues and calibration

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Support Command Center

Agent workspace for triaging customer conversations

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

high stress

Field Service Planner

Scheduling workspace for mobile operations teams

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Collaborative Roadmap

Product planning workspace with dependencies and conflicts

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Inventory Operations

Warehouse workspace for stock, replenishment, and anomalies

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Finance Close Workspace

Month-end close with approvals, exceptions, and audit evidence

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Travel Disruption Desk

Rebooking operations under capacity and policy constraints

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

high stress

Alpine Studio

Editorial portfolio for a mountain architecture practice

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

SaaS Conversion Lab

Conversion-focused landing page for an AI workflow platform

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

high stress

Luxury Property Launch

Editorial launch page for a high-end coastal residence

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Healthcare Trust Landing

Patient-first landing page for a digital care service

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Developer Platform Launch

Technical launch page for an API and automation platform

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Climate Tech Launch

Credible industrial decarbonization landing page

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

high stress

Night Market Identity

Campaign system for a city food and culture festival

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

high stress

Kinetic Product Launch

Visual launch kit for a performance footwear release

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

high stress

Album Release System

Music launch identity across cover, social, and release moments

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Museum Exhibition Identity

Cultural identity across exhibition, wayfinding, and ticketing

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Sustainable Beauty Campaign

Premium product system with disciplined claim boundaries

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Esports Tournament Package

Broadcast-ready identity for a competitive event

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Urban Mobility Pulse

Interactive dashboard for city transport performance

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Revenue Forecast Studio

SaaS scenario planning with assumptions and uncertainty

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Supply Chain Risk Map

Supplier and geography risk exploration without external maps

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Clinical Operations Monitor

Privacy-aware hospital operations dashboard using synthetic data

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Energy Grid Control Room

Synthetic grid performance, forecast, and alert triage

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Experimentation Analysis Lab

A/B analysis with sample size, uncertainty, and guardrails

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Secure Checkout Review

Defensive review of a checkout flow with prioritized remediation

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Phishing Triage Lab

Evidence-led triage of suspicious messages and indicators

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Incident Response Console

Defender workspace for triage, containment, and recovery decisions

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Access Control Policy Lab

Least-privilege review across roles, resources, and edge cases

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Cloud Misconfiguration Review

Defensive review of synthetic cloud configuration and exposure

Run not started
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

extreme stress

Dependency Risk Triage

Defensive SBOM and vulnerability prioritization workspace

Run not started

Three surfaces, three different questions

Keeping the tracks separate prevents a polished interface from masking weak reasoning, unsafe actions, or unreliable task execution.

Core

Model-only tasks with deterministic checks wherever possible. The capability score excludes agent execution. JavaScript checks require explicit use of a disposable, network-denied Docker sandbox; they never execute inside the application process.

Agent

Tool use and recovery through bounded, read-only virtual tools and explicit budgets. Agent score and reliable agent success are reported separately.

Design

Screenshots and generated interfaces are scored separately against a visual rubric. Design never contributes to Core and remains unranked until a live, rubric-scored release exists.

Open the Design suite

Clean external release contract

The clean core-v2 secret pack and public commitments are prepared, but no live ranking has been run. Release still needs provider credentials, a candidate and two independent judge models, a finite approved worst-case budget, the complete 96-task run, verification, and an exact approval evidence seal. Synthetic reports validate only the pipeline and never establish a ranking.

Public disclosure
20 complete public tasks; the 76 exposed reserved tasks are compromised, synthetic/dev-only, and never eligible for a live release.
Public result gate
Live, complete, publishable, explicitly approved, and exact report plus artifact-evidence match.
Code execution
Explicit opt-in only: digest-pinned, network-denied Docker containers enforce CPU, memory, process, time, output, and filesystem limits with no host mounts.
Bounded I/O
Provider bodies cap at 8 MiB. Attachments and artifacts require explicit canonical roots, realpath containment, size checks, and no symlinks.

Results by section

Choose a capability to compare models, then inspect canonical public task responses. Agent and Design remain separate from Overall Core.

Equal scores share the same rank.

DesignSeparate visual suite · not ranked until a live rubric-scored release exists

Overall Core

Weighted capability across the seven model-only tracks. Agent and Design remain separate.

Awaiting live release

Awaiting live release

Rankings, model names, winners, and responses remain hidden until a complete live release passes publication, approval, and integrity checks. You can still explore every section above without synthetic results being presented as a model comparison.

CresciBench Core by CrescitalyCapability · Agent · Reliability · Risk · Efficiency