Benchmarks

What an app costs to build with AI

Claude built the same apps in each stack: small apps from scratch, full-stack apps, and changes to existing projects of 11, 42 and 102 components. Every app was run and used in a simulated browser. The cost is the real API bill: the spec in the prompt, every retry and the thinking tokens.
Claude Haiku 4.5−38%cost per working app vs React + TypeScript
ArtScript$0.0122 · 43/52
React + TS$0.0197 · 42/52
Svelte 5$0.0244 · 41/52
Vue 3$0.0228 · 41/52
SolidJS$0.0206 · 42/52
Claude Sonnet 5.5−56%cost per working app vs React + TypeScript
ArtScript$0.0071 · 52/52
React + TS$0.0160 · 52/52
Svelte 5$0.0148 · 52/52
Vue 3$0.0160 · 52/52
SolidJS$0.0152 · 52/52
Claude Opus 5.5−53%cost per working app vs React + TypeScript
ArtScript$0.0188 · 52/52
React + TS$0.0404 · 51/51
Svelte 5$0.0354 · 52/52
Vue 3$0.0370 · 52/52
SolidJS$0.0372 · 52/52
Read the numbers honestlyTwo or three runs per task is an early signal, not a definitive benchmark. With the smallest model, Haiku, ArtScript solved fewer of the new-app tasks than React; on modifications of large projects with focused context it is close to even. The raw data of every run is in the repository, and anyone can reproduce it with npm run eval.
Runtime speedIn a local run of js-framework-benchmark's own harness, ArtScript is in the group of Solid, Svelte and Vue on every operation (fastest of them at selecting a row, about 1.1× Solid at updating every 10th row) and still behind at creating 10,000 rows (1.15× Solid). Apps ship about 6 KB compressed.The full table and how to reproduce it

How it was measured

  • The 2026-10-02 measurement (the numbers above): 50 tasks, three models, five stacks, two runs per task. Before it, a pilot ran ArtScript alone once per model (USD 2.34, files 2026-10-02T16-1*); what the models tripped on was fixed in the compiler (accepting what they write, clearer errors, a for statement) and the measurement started from that version. Haiku then solved 76/100 ArtScript runs, against 86–89 in the other stacks; two more rounds of the same kind of fixes followed, and its ArtScript cells were run again each time (80/100, then 90/100, the one reported). React, Svelte, Vue and Solid ran once: their toolchains didn't change. Sonnet's and Opus's ArtScript cells ran once, between those rounds. Every answer of every run is in the raw data, and --replay re-checks them with the current compiler.
  • Each task is the same functional request for every stack: small apps created from scratch (some full-stack), small modifications, and modifications to larger generated projects (11, 42 and 102 components). Claude gets the task, returns files, and the harness validates them: ArtScript with its compiler, React and SolidJS with strict tsc, Svelte and Vue with their compilers. Errors are fed back, up to 3 attempts.
  • In modification tasks each stack may use its cheapest edit format: ArtScript an art patch, React and Svelte search/replace edit blocks (like a coding agent's Edit tool). Full files are also accepted. Runs before 2026-10-01 had no edit formats: every stack returned full files.
  • Cost is computed from the real usage the API returns: the ArtScript spec in the system prompt, retries and thinking tokens (billed as output) all count.
  • ArtScript's system prompt includes its spec (1.2K–3.2K tokens depending on the run date; 3.2K in the 2026-10-02 measurement), which is served from the prompt cache after the first request; the "without prompt cache" column prices those tokens at the full input rate.
  • Since 2026-10-01 every app is also run and used like a person would: it's mounted in a simulated browser (happy-dom) and a stack-agnostic check clicks, types and reads the screen (e.g. adds and completes todos, reloads the page to check data persisted on the server). A failed check is fed back to Claude like a compiler error. Earlier runs only checked that code compiled and typechecked.
  • Full-stack tasks: the other stacks also write their own server.ts (Node http, no dependencies, data in memory); ArtScript uses api, which also persists to SQLite. Svelte is validated without TypeScript type checking of .svelte files, which favors it.
  • Since 2026-10-01 18:47, ArtScript modification tasks get docs/SPEC-EDIT.md (~800 tokens) instead of the full spec; creation tasks keep the full spec. Earlier runs sent the full spec everywhere.
  • Each run uses the ArtScript spec as of its date (in English since 2026-10-02).
  • A model refusal (Opus, on the photo-upload task: twice in React, once in Vue) is not a result: it is left out, and re-run when the run is resumed.
  • Task prompts (and the feedback given to the model) are in Spanish; they are the fixed dataset these numbers were measured on.
  • App JS: each working app bundled with esbuild (minified, production mode) and compressed with brotli: the JavaScript a browser downloads. ArtScript's includes its runtime; React's includes React DOM; Svelte's includes its client runtime.
  • Two runs per task is a signal, not a definitive benchmark: the smaller the model, the more its results move between runs (Haiku's ArtScript cells went 76, 80, 90 out of 100 as the compiler improved, and part of that is noise). Reproduce it with npm run eval; npm run eval -- --dry-run checks the harness and every reference app offline, and --replay <results.json> re-checks stored answers with the current compiler.