Case Studies / Where the Specification Lived

Experiment Case Study

Where the Specification Lived

Two agents built the same Traffic Simulator in the same repository, one from a 38-word prompt and one from a 615-word specification. A blinded review of both artifacts shows what each prompt bought, and why prompt length and specification load are different variables.

Agent-assisted developmentSpecificationBlinded review

Question

How much of an agent's specification has to arrive in the prompt when the agent works inside a mature repository? Two agents were given the same task in the same repository: finish the Traffic Simulator that the site listed as in development. One prompt was 38 words. The other was 615. The working thesis was that prompt length and specification load are different variables, because much of the effective specification may already live in the repository.

Where specification can live

A mature repository already specifies a great deal: AGENTS.md, tests, neighbouring implementations, types, names, documentation, route machinery, and visual conventions. What a prompt adds is the residual specification on top of that environment.

So the question is not whether short prompts beat long ones. It is where each part of the specification lived, and which copy of it each agent obeyed.

Experimental Setup

Held constant

Same repository and base commit (ac5fe45), same task, and the same standing instructions: work in an isolated worktree, commit in coherent chunks, do not push. Both sessions ran Claude Opus 5.5 in Claude Code 2.1.289, and the prompts were submitted five seconds apart.

Varied

Only the invocation prompt. The Simple prompt asked the agent to inspect the repository and existing simulations, implement the simulator completely, and verify it. The Detailed prompt restated the placeholder's concept, required specific behaviours, named conventions to follow, and set testing, validation, and reporting requirements.

Not controlled

One pair of runs, not a crossover. Agent variance and tool-use luck are uncontrolled. The two agents first collided in the same worktree, each overwrote one of the other's files, told the other, and moved to separate worktrees. The experiment is quasi-controlled, not lab-clean.

The two prompts

Simple prompt (38 words, 2 paragraphs)

Implement the traffic simulator on my personal site. Work in your own isolated worktree. You can commit in coherent chunks, but do not push.

Inspect the repository and existing site/simulations, then implement it completely and verify your work.

Detailed prompt (615 words, 15 paragraphs)

Implement the traffic simulator on my personal site. Work in your own isolated worktree. You can commit in coherent chunks, but do not push.

The Traffic Simulator already exists conceptually in the site as an in-development simulation at `/simulations/traffic-simulator`. Its existing description establishes the core idea: multi-lane traffic flow with signal timing, vehicle queuing, throughput visualization, and the emergence of congestion/coordination from those rules. Turn that placeholder into a finished, playable simulation.

Before implementing anything, inspect the repository thoroughly enough to understand the current architecture and conventions. Read and follow `AGENTS.md` and any applicable repository instructions. Treat the repository as the source of truth: inspect the existing simulations—especially the most recently implemented ones—and reuse established patterns where appropriate rather than designing an unrelated mini-app.

The personal site is a React + Vite + TypeScript frontend using React Router. Simulations belong under the Simulations area rather than Projects or Case Studies. The site has centralized simulation/route metadata and existing mechanisms around routing, metadata, sitemap/static route generation, styling, testing, and build validation. Discover their current form rather than assuming filenames or structure from this prompt.

Build the traffic simulation as an actual interactive system, not a canned animation. The underlying state and rules should produce the observed traffic behavior. At minimum, the finished experience should meaningfully expose:

- multi-lane vehicle flow;
- traffic signals and their timing;
- vehicles stopping, queuing, and proceeding according to the simulation state;
- congestion emerging under unfavorable conditions rather than being scripted;
- throughput or other useful live traffic metrics/visualization;
- enough user control to experiment with the system and see how changing traffic conditions or signal behavior affects the outcome.

Use your judgment for the exact road/intersection model, controls, parameters, visualization, presets, and simulation mechanics. Prefer a small coherent model with understandable behavior over superficial complexity. The simulator should fit the site's existing idea of simulations as interactive explorations of rules, state, feedback, and emergence.

Treat the simulation model and rendering/UI as distinct concerns where practical so the important rules can be tested independently. Avoid broad refactors and unnecessary dependencies. Preserve the site's established visual language and interaction conventions rather than introducing a new design system.

Integrate the simulator completely into the current site. That includes whatever the current architecture actually requires for the route, simulation registry/card, availability/status, page metadata, sitemap/static route handling, navigation or discovery surfaces, and documentation. Remove or update obsolete “In development”/unavailable treatment once the simulator is genuinely playable.

Make it usable across the site's supported desktop and mobile layouts. Preserve light/dark theme behavior if applicable. Pay attention to accessibility, keyboard-usable controls where appropriate, reduced-motion behavior, resizing, cleanup of animation/timers, and avoiding runaway CPU work or needless React rerenders.

Add meaningful automated tests for the simulation's important invariants and deterministic logic rather than only testing that components render. Use the repository's established testing approach. Test integration/registry behavior where that is already conventional.

Before finishing, inspect the diff as a whole and run the relevant repository validation suite. At minimum, ensure typechecking, linting, the production build, existing smoke/regression tests, and the new simulator tests pass. If browser-level testing is supported in the repository, exercise the finished simulator there as well, including its route, primary controls, responsive behavior, themes, and console/runtime errors.

Do not weaken or delete existing tests just to make the implementation pass. Do not modify unrelated parts of the site. Do not push anything.

You may commit the implementation in coherent chunks. When finished, leave the worktree clean and report:

1. what you built and the important design decisions;
2. the files/areas changed;
3. the tests and validation you ran and their results;
4. the commits you created;
5. any remaining caveats or deliberately deferred improvements.

What Each Agent Built

The two experimental artifacts at phone width, captured from production builds of their final commits. Left: the Simple artifact's signal corridor, folded into rows. Right: the Detailed artifact zoomed to one intersection in rush hour.
What the blinded evaluation measured
Simple promptDetailed prompt
Prompt length38 words615 words
Wall time, prompt to closing report26 min 55 s56 min 34 s
Change size27 files, +2,622 / −9 lines30 files, +3,235 / −10 lines
Model tests1426, plus a browser regression script
Typecheck, lint, build, smoke suitePassPass
30-minute invariant probes51 configurations, no violations60 configurations, no violations
PhenomenaSignal coordination, merge bottleneck, ring-road stop-and-go waveSignal coordination in both directions, oversaturation, fixed vs actuated control
Single-rule breakages the suite missed25 of 51, about 5 of them no-opsRed, yellow, intersection, sequence, and actuation rules caught; metric definitions and some parameters missed
Correctness defect foundRed-light crossings after abrupt timing editsNone; two rules never fire in its presets
Documented claims re-measuredAll reproduce at the default seed; one holds in 6 of 10 seedsCoordination and actuation claims hold across 10 seeds; capacity stated as ~1,500 veh/h, measured 1,440

Rows are comparable only where the same measurement was applied to both. The mutation counts are not: each suite was tested against breakages of its own model's rules.

How the Artifacts Were Judged

What the agents reported

Each run ended with a careful, specific closing report. The Simple run listed tests for “red lights obeyed” and said an independent code review it ran “found no engine bugs.” The Detailed run reported that the eastbound approach “saturates at about 1,500 vehicles/hour,” and that switching off each driver rule in turn made a test fail.

Those reports were self-reports: each author's account of its own work, checked by reviewers it chose itself. They are evidence of what was claimed. They are not validation, so the artifacts were judged directly.

Evaluate blind

The evaluator compared the two branches without knowing which prompt produced which. Branch names, commit counts, elapsed time, and verbosity were excluded as quality signals.

Use both simulators

Every scenario and control was exercised in production builds at desktop and phone widths, in both themes and with reduced motion, while watching for console errors.

Probe past the shipped tests

Independent runs held every control at its extremes for 30 simulated minutes and checked invariants at every step. Claims in each project's documentation were re-measured, across ten random seeds where seeds applied.

Break rules on purpose

Each important model rule was disabled or altered one at a time, and the shipped test suite was run against each change. A suite that still passes with a rule removed is not constraining that rule.

No composite score

The artifacts were compared dimension by dimension. Breadth, rigour, and fit pull in different directions, and a single number would hide exactly that.

What the Evaluation Found

Both passed every repository gate

Typecheck, lint, production build, and the smoke suite passed for both. Neither broke an invariant in the 30-minute probe runs: no overlapping vehicles, no reversing, and no invalid numbers.

Breadth on one side, depth on the other

The Simple artifact shows three structurally different phenomena. The Detailed artifact goes further into one of them, signal control: both travel directions, cross traffic, safe transitions, and actuation.

The time-space diagram made emergence visible

In the Simple artifact, signal coordination, the merge queue, and the ring road's backward-travelling wave all appear as shapes. The Detailed artifact reports comparable effects as numbers in tiles and tables.

Test strength diverged

Against the Simple suite, 25 of 51 single-rule breakages went unnoticed; about five of those change nothing the model can show. The Detailed suite caught breakages of its red-light, yellow, intersection, signal-sequence, and actuation rules, and missed changes to how metrics were defined and to a few parameters.

A correctness bug its author did not report

Editing signal timing mid-run in the Simple artifact switched lights straight from green to red. In 49 abrupt edits there were 12 crossings on red, 11 of them more than 2 s after the light turned. Its red-light test allowed crossings up to 2 s into red, and its own review had reported no engine bugs.

Rules the presets never exercise

The Detailed artifact's rules against blocking an intersection never fired in its shipped presets: zero times in 52.2 million checks over 30-minute runs. Its report disclosed that its spillback test forces the situation, but its page description still promised spillback.

The Simple artifact's time-space diagrams for the same arriving traffic, at phone width. Each dot is one vehicle at one moment: time runs left to right, position bottom to top, so rising streaks are moving traffic and the lines at 250, 550, and 850 m are the signals' stop lines. With a green wave, platoons that clear the first signal mostly meet green at the next ones. With a reverse wave, they stop again at every signal, and the stops appear as bright bands on each stop line.
The jam comes from the same car-following rule every driver uses; nothing scripts it. Left alone, this ring also jams by itself after about 15 to 20 simulated minutes, which the clip does not show.
It shows the deeper signal model: two streets competing for one intersection through yellow and all-red. Its rules against blocking the intersection exist but do not fire in this preset.

Outcome

What changed before publication

What this says about repository-resident specification

The Simple prompt said almost nothing about this repository, and the Simple run still reconstructed nearly all of its integration requirements: the route, simulation card, metadata, sitemap, static route shell, smoke tests, theme tokens, reduced-motion handling, and the existing split between model and renderer. That part of the specification came from the repository.

The Detailed prompt asked for things the repository did not require on its own, such as meaningful tests of the model's invariants and browser-level checks, and the Detailed run delivered them: a stronger test suite and a browser regression script. Those are task-specific non-negotiables, and the invocation is where they had to live. Its safer signal transitions were not in the prompt; they were the run's own design choice.

The Detailed prompt also listed signals, queuing, and throughput as minimum requirements, and the Detailed run built deeply around exactly those, while the Simple run, given no list, ranged wider. One pair of runs cannot show that the list caused the narrower scope. It is consistent with detail consuming decisions that could have stayed open.

Nor does this show that the repository always suffices. The Simple run shipped a red-light bug that its own tests and review missed, and the repository contained no traffic model from which to inherit the rule that a light never skips yellow.

Limitations

The principle

Put durable invariants in the repository, where every task inherits them. Put the task's non-negotiables in the invocation. Leave genuinely open decisions open. Then judge what comes back by the artifact, not the report. This experiment does not show that shorter prompts are better. It shows that the useful question about a prompt is which of those three places each requirement belongs in.

Evidence

Repository provenance / Public

Simple-prompt artifact

The Simple run's final commit, preserved unchanged by the tag traffic-sim-experiment-simple. The post-evaluation fixes were made on top of it.

View artifact

Repository provenance / Public

Detailed-prompt artifact

The Detailed run's final commit, preserved unchanged by the tag traffic-sim-experiment-detailed as the comparison artifact.

View artifact

Evaluation record / Public

Blinded evaluation record

The evaluation written before the prompt mapping was revealed, the exact prompts and closing reports, probe and mutation results, captured media, and the scripts that produced them.

View artifact

Local system evidence / Private

Session transcripts

The two authoring sessions' transcripts are the source of the prompts, timings, models, and closing reports quoted here.

Note: The transcripts stay local; the evaluation record reproduces the prompts and closing reports verbatim.