Design benchmark — unranked

Same prompt, one attempt per model, saved as raw output, extracted HTML, screenshot, and render telemetry. Design stays separate from Core and does not produce a general model ranking.

Runner
npm run bench:visual -- --provider openrouter --models all --limit 20

No one-shot run artifacts yet

No approved live run or synthetic fixture is available. Run the benchmark command above to create an artifact set under data/crescibench/one-shot-ui-benchmark/prompt.md.