Automated Tests
In short
Tests are generated per feature rather than per file, run on demand on the Build Server, and reported as which features are covered and passing, so the answer to "did that change break anything" is a run rather than an opinion.
Tests are the project’s regression net. Every change made from here on risks breaking something that worked yesterday, and a growing system stops fitting in anyone’s head long before it stops growing. Tests are how that risk is checked automatically rather than remembered.
The Automated Tests page appears in your project from the Build stage, when there is a real system to test.
Coverage is per feature, not per file
Test coverage in most tools means “what percentage of lines ran” — a number that says nothing about whether your invoices actually work. Sapilon reports coverage against the features your project declares: which of them have tests, at which layers, and whether those tests passed on the last run. “Invoices — 12 passing” is a sentence you can act on.
The layers behind it are the ordinary ones: unit tests for logic, API tests for endpoints, database tests for the schema, UI tests for components, and end-to-end tests that drive real journeys through the app.
Who writes them
The agent does, on request. Tests are not generated on every turn — that would be slow and expensive for a change to a label — so the page lets you ask for them where you want them: for a feature with no coverage, or for a gap in a feature that has some.
They can also come from a task’s plan. When a plan changes what a feature does, it asks whether to add automated tests. If you say yes, its last step writes tests that check the decisions you agreed, not just that the feature exists.
They are ordinary project source, kept in your repository alongside the code. A technical user can read and edit them through the normal code surfaces; there is no separate test editor.
Running them
Runs are started by you, from the Automated Tests page, or by a checkup — the checkup’s test check is the same run. They execute on your Build Server, so a run costs you nothing extra to host, and take from under a minute to a few minutes depending on how much there is.
When a run finishes, failures are listed with what broke, said in words rather than in the test runner’s shorthand: “the API crashed: 500 Internal Server Error instead of 200 OK”, not “expected 500 to be 200”. Open one and the side panel has the rest:
- Overview — what that kind of failure usually means, what the test expected and what it got, and the line of the test that caught it.
- Error — the test runner’s own output and stack trace, exactly as reported, to copy.
- Code — the whole test, with the line where it failed highlighted.
- Server output — when an API test got a server error back, what the API itself logged while answering it, if the run recorded that. It is where the cause is — a stack trace, a failed query — because all the test ever sees is the status code.
Handing a failure to the Agent is usually the fastest way through it; the test itself is often the clearest description of what the code was supposed to do.
After the Agent has worked on a failure, the finished turn offers Run the tests again. It takes you back to this page with the run already starting, scoped to the feature and layer that went red — because the only thing that can say a failure is gone is another run, not a look at the app.
When a test doesn’t make sense
A red test means the app and the test disagree. It does not mean the app is wrong: the test may describe a journey nobody asked for. That happens most with generated end-to-end tests, which are written from what the code appears to do, and are named after the journey the agent inferred — a name you never chose, describing a promise you never made.
So every failure offers three things, not one:
- Explain and fix — the test is right; change the app until it passes. The one exception is a test that asks for something the app cannot do by design (say, a reply body on a response that by definition has none). The agent corrects that test instead, and tells you it did.
- What does this test check? — have the test read back to you in plain language, with where it came from and whether the app or the test is the thing that is wrong. Nothing is changed; you get a recommendation and decide.
- Remove test — the test describes something this project never promised. It is deleted as an ordinary change you can undo, and nothing else is touched.
Removing a test is a real answer, not a defeat. A test you cannot read and did not ask for costs you a red suite forever, and a red suite hides the next real regression. What it does cost is coverage: nothing watches that behaviour afterwards, and the feature’s cell on the grid can drop back to a gap.
What the page tells you
The card at the top of the Automated Tests page says where your tests stand, in one word, and what to do next:
- Healthy — every feature is covered, and the last run passed on your code as it is now. Nothing to do.
- Waiting — your move: some features are missing tests, so generate them, or your tests haven’t run on your latest code, so run them.
- Creating — the agent is writing tests.
- Running — your tests are running on the Build Server.
- Failed — some tests failed, or the last job didn’t finish. Fix in Agent hands every failure to the Agent in one go, and the finished turn offers the run that checks it.
Creating and Running happen in the background, so you can keep working. The card beside it has the numbers: how many features are covered, how many tests are still to write, and how many tests the last run passed, failed and skipped.
What tests are not
They are not a substitute for using the app yourself. A green suite means nothing regressed against what was written down, not that the product is right — which is why Build also asks you to run the app end to end, and why Rehearsal exists.