# Agents explore, tests remember

> For agents: start with the [agent guide](https://ratstack.sh/llms.txt). Every page is Markdown by default; add `Accept: text/html` for HTML.

Free workshop: [how to burn a trillion tokens and get good results](/tokenmaxx#interested).

[Joel Hooks, October 1, 2026](https://x.com/joelhooks/status/2105660110068699302):

> "run up to 50 agent users using luna:max and have them click through various scenarios. chaos testing. go nuts, fix all the bad stuff"
>
> real engineering

[André Staltz replied](https://x.com/andrestaltz/status/2105716416444010671):

> @joelhooks And when those agents find good reproduction steps for a bug, convert that into playwright tests.

The proposal gives agent users the exploratory work: click through scenarios and find failures. A good reproduction then becomes a deterministic Playwright regression test. Exploration discovers a path; the test remembers it without needing the explorer to find it again. This is a direction, not a claim that rat-stack already runs fifty agent users or has this Playwright workflow built.

## Rat-stack's reading

These connections are our reading, not claims Staltz made in that reply:

- [Model-based testing](/lore/model-based-testing) generates sequences of actions and checks the system against a model. Agent users can propose action sequences for exploration; that alone is not model-based testing, because it supplies neither the model nor the assertions.
- [Property-based testing](/lore/property-based-testing) shrinks a failure to a small counterexample. Keep the smallest useful failing case from exploration, rather than every step the agent happened to take.
- [The fence](/lore/the-fence) puts a correction into an enforceable check. A reproduced mistake becomes a test when it exposes a genuine behavior gap, not a test that merely pins the patch. See [tests that earn their place](/lore/tests-that-earn-their-place).

## A human-driven example

A human-driven signup proof found a five-second intake timeout and a confirmation email's `/confirm` link returning 404. These were not discoveries by agent users, and the resulting tests are Effect tests, not Playwright tests.

[Commit 222edd8](https://github.com/joelhooks/rat-stack/commit/222edd833ccc19f41da236f6556a5a09dace1582) raised the timeout to twenty seconds and made timeout request an immediate retry. Its [intake test](https://github.com/joelhooks/rat-stack/blob/222edd833ccc19f41da236f6556a5a09dace1582/packages/core/test/interest-intake.test.ts) uses `TestClock` and a hanging fake HTTP client to check that retry outcome. Its [route test](https://github.com/joelhooks/rat-stack/blob/222edd833ccc19f41da236f6556a5a09dace1582/apps/mischief/test/interest.test.ts) checks that GET and POST to `/confirm` return 308 and that the GET redirect preserves the token query. The discovery became checks that do not depend on repeating the live proof.

## Sources

1. [joel ⛈️ on X: ""run up to 50 agent users using luna:max and have them click through various scenarios. chaos testing. go nuts, fix all the bad stuff" real engineering" / X](https://x.com/joelhooks/status/2105660110068699302)
   Joel Hooks. Exploration proposal; used for agents clicking through scenarios to find failures. Accessed 2026-10-01.

2. [André Staltz on X: "@joelhooks And when those agents find good reproduction steps for a bug, convert that into playwright tests." / X](https://x.com/andrestaltz/status/2105716416444010671)
   André Staltz. Reply; used for converting good reproductions into repeatable tests. Accessed 2026-10-01.

3. [Commit 222edd8](https://github.com/joelhooks/rat-stack/commit/222edd833ccc19f41da236f6556a5a09dace1582)
   GitHub joelhooks/rat-stack. Signup fix; used for the human-driven timeout and confirmation-route example. Accessed 2026-10-01.

## Lore on this page

- [Model-based testing](/lore/model-based-testing)

## Linked from

- Agent guide → An Effect stack so pure (aspirational) Kit Langton will blush. → [Read page](https://ratstack.sh/llms.txt)

- Change log → What changed in the files served here, newest first. → [Read page](https://ratstack.sh/log)

- Code snippets quote a pinned source → Shipped design brief: pinned Git objects, typed diagnostics, and one build-only snippet model for both views. → [Read page](https://ratstack.sh/lore/code-snippets)

- Full agent guide → An Effect stack so pure (aspirational) Kit Langton will blush. → [Read page](https://ratstack.sh/llms-full.txt)

- log.md → The generated change log as Markdown. → [Read page](https://ratstack.sh/log.md)

- Rat Stack lore | rat-stack → Short, source-grounded notes on the ideas and decisions behind rat-stack. → [Read page](https://ratstack.sh/lore)
