← Back to stats

Case studies

Real prompts, real models, no cherry-picking - every result shown, including the failures.

Building a SaaS landing page

Building a SaaS landing page

Same prompt, 10+ models, real deploys - who actually ships a working site.

Anatomy of a failed build

Anatomy of a failed build

Half the models failed on the same task - five different ways, some of them looked like success.

Cloning apple.com, guided vs. plain prompt

Cloning apple.com, guided vs. plain prompt

A little extra guidance in the prompt made every model that succeeded noticeably faster and cheaper.

Kanban board, guided prompt + multi-model auto-fix

Kanban board, guided prompt + multi-model auto-fix

Only 33% worked on the first try - a real browser check plus a second opinion from another model got that to 100%.

A multi-step, multi-API landing page

A multi-step, multi-API landing page

Research a topic, then build a page that actually fetches from 3 real APIs - all 4 models finished, one needed a fixer swap.

An overnight free-model marathon

An overnight free-model marathon

13 fun tasks, 4 free models, 52 runs - only 2% landed on the first try, but the auto-fixer saved nearly every one.

Does planning first actually help?

Does planning first actually help?

Counterintuitive: a plan-first step made success rate slightly worse, not better, across 48 runs.

The 10 most common coding tasks

The 10 most common coding tasks

6 models, 10 everyday task types - a real detector gap was found, fixed, then re-run to 100%.