Isolated and disposable
One container and one git worktree per task. When the PR is opened the sandbox is destroyed — nothing persists but the code and the test evidence.
Plan v2 is approved. Three coder agents pick up the first wave of chess tasks — each in its own disposable sandbox — write the tests before the code, then prove the pieces work together in a full-stack preview environment.
| Test | What it proves | In the chess build | Runs in |
|---|---|---|---|
| Unit | One function behaves, in milliseconds, with no network or database. | The clock only ticks for the side to move (fake timers). | Sandbox |
| Property-based | A rule holds for thousands of generated inputs, not just the cases someone thought of. | Random legal games never reach an illegal board. | Sandbox |
| Integration | Your code works with real dependencies — not mocks. | Matchmaking queue survives a restart, against a real Postgres container. | Sandbox |
| Contract | Services agree on message shapes, so one side can't silently break the other. | Every WebSocket move message matches the schema. | Preview env |
| End-to-end | A real user journey works through real browsers. | Two browsers play a game; Black sees White's move. | Preview env |
| Load | It stays fast under realistic traffic. | 500 virtual players, move latency p95 < 150 ms (AC-1). | Preview env |
| Static + security | Lint, types, SAST and dependency scans. | Run by agents here, then re-run independently as gates. | Sandbox → Verify |
One container and one git worktree per task. When the PR is opened the sandbox is destroyed — nothing persists but the code and the test evidence.
No production credentials, ever. Outbound traffic is limited to an allowlist such as the package registry, so a confused agent can't reach anything that matters.
Agents write failing tests from the acceptance criteria before writing code. Green tests mean the spec is met — not just that the code compiles.