CresciBench Core

See what every model builds under pressure

Compare complete games, apps, landing pages, data products, and defensive security labs by challenge, sector, or model. Versioned evidence keeps capability, reliability, risk, visual quality, latency, and cost separate.

Public evidence contract

The API exposes status for withheld runs, never synthetic model aggregates. A live ranking requires a complete verified release using the clean external core-v2 catalog, explicit approval, and an exact report-and-artifact evidence seal. The 20 public prompts and their canonical repetition-1 responses are disclosed only after that gate; there is no reserved payload-membership index. Private payloads, full artifacts, and tool traces remain outside Git and Vercel.

Core report JSON
Evidence explorer

See what every model builds under pressure

Compare one-attempt games, product interfaces, landing pages, data products, and defensive security labs. Every output is judged against the same published stress contract; empty cards are explicit run slots, never simulated results.

Models
11
Stress tests
36
Output slots
396
Published
0
  • One attempt
  • 12k token cap
  • Offline artifact
  • Responsive
  • Failure states
  • 100-point rubric
Core release: emptyOpen official leaderboard

Games

extreme stress

Voxel World

Minecraft-style browser game

Playable browser buildKeyboard and pointer controlsResponsive HUD and pause stateSource artifact
Pressure applied
Interaction depthOffline runtimePerformance pressureState recoveryResponsive layout
Browser game rubric
Version 1 · 100 points

Gameplay completeness

35pts

Interaction quality

20pts

Runtime resilience

15pts

Visual craft

15pts

Accessibility and responsiveness

15pts

Evaluation gates

  • The playable build starts offline without external requests or console errors.
  • Movement, camera, world interaction, pause, and restart states all remain operable.
  • Keyboard focus, instructions, HUD, and status feedback remain legible at mobile and desktop widths.
Output slot

GPT-5.6 Sol

No generated artifact is being simulated

GPT-5.6 Sol

OpenAI

Run not started
openai/gpt-5.6-sol
Output slot

GPT-5.6 Terra

No generated artifact is being simulated

GPT-5.6 Terra

OpenAI

Run not started
openai/gpt-5.6-terra
Output slot

Claude Opus 5

No generated artifact is being simulated

Claude Opus 5

Anthropic

Run not started
anthropic/claude-opus-5
Output slot

Claude Sonnet 5

No generated artifact is being simulated

Claude Sonnet 5

Anthropic

Run not started
anthropic/claude-sonnet-5
Output slot

Gemini 3.1 Pro Preview

No generated artifact is being simulated

Gemini 3.1 Pro Preview

Google

Run not started
google/gemini-3.1-pro-preview
Output slot

Gemini 3.6 Flash

No generated artifact is being simulated

Gemini 3.6 Flash

Google

Run not started
google/gemini-3.6-flash
Output slot

Grok 4.5

No generated artifact is being simulated

Grok 4.5

xAI

Run not started
x-ai/grok-4.5
Output slot

DeepSeek V4 Pro

No generated artifact is being simulated

DeepSeek V4 Pro

DeepSeek

Run not started
deepseek/deepseek-v4-pro
Output slot

Qwen 3.7 Max

No generated artifact is being simulated

Qwen 3.7 Max

Qwen

Run not started
qwen/qwen3.7-max
Output slot

Llama 4 Maverick

No generated artifact is being simulated

Llama 4 Maverick

Meta

Run not started
meta-llama/llama-4-maverick
Output slot

Mistral Large 2512

No generated artifact is being simulated

Mistral Large 2512

Mistral AI

Run not started
mistralai/mistral-large-2512

Three surfaces, three different questions

Keeping the tracks separate prevents a polished interface from masking weak reasoning, unsafe actions, or unreliable task execution.

Core

Model-only tasks with deterministic checks wherever possible. The capability score excludes agent execution. JavaScript checks require explicit use of a disposable, network-denied Docker sandbox; they never execute inside the application process.

Agent

Tool use and recovery through bounded, read-only virtual tools and explicit budgets. Agent score and reliable agent success are reported separately.

Design

Screenshots and generated interfaces are scored separately against a visual rubric. Design never contributes to Core and remains unranked until a live, rubric-scored release exists.

Open the Design suite

Clean external release contract

The clean core-v2 secret pack and public commitments are prepared, but no live ranking has been run. Release still needs provider credentials, a candidate and two independent judge models, a finite approved worst-case budget, the complete 96-task run, verification, and an exact approval evidence seal. Synthetic reports validate only the pipeline and never establish a ranking.

Public disclosure
20 complete public tasks; the 76 exposed reserved tasks are compromised, synthetic/dev-only, and never eligible for a live release.
Public result gate
Live, complete, publishable, explicitly approved, and exact report plus artifact-evidence match.
Code execution
Explicit opt-in only: digest-pinned, network-denied Docker containers enforce CPU, memory, process, time, output, and filesystem limits with no host mounts.
Bounded I/O
Provider bodies cap at 8 MiB. Attachments and artifacts require explicit canonical roots, realpath containment, size checks, and no symlinks.

Results by section

Choose a capability to compare models, then inspect canonical public task responses. Agent and Design remain separate from Overall Core.

Equal scores share the same rank.

DesignSeparate visual suite · not ranked until a live rubric-scored release exists

Overall Core

Weighted capability across the seven model-only tracks. Agent and Design remain separate.

Awaiting live release

Awaiting live release

Rankings, model names, winners, and responses remain hidden until a complete live release passes publication, approval, and integrity checks. You can still explore every section above without synthetic results being presented as a model comparison.

CresciBench Core by CrescitalyCapability · Agent · Reliability · Risk · Efficiency