A Green Checkmark Is Not Proof

A passing check is useful once I can say what it checked, what it skipped, and what question still needs an answer.

laptop screen displaying colorful code

A passing check is useful once I can say what it checked, what it skipped, and what question still needs an answer.

I've learned to get suspicious when a technical result collapses into one green mark. Green is comforting. It is also remarkably good at making several different questions look like the same question.

Did the code compile? Did the tests pass? Did the page work in a browser? Did the delivered behavior match what somebody will actually use? Those checks overlap, but they are not interchangeable. I don't need every possible check for every change. I do need to stop borrowing confidence from one layer to cover another.

Here is the scope checklist I use when a result looks reassuringly simple.

Start with the claim

Before I run a check, I write down the claim it can support. A compiler can tell me whether the code satisfies a particular set of language and build rules. It cannot tell me whether I built the right behavior. A unit test can tell me whether its assertions passed. It cannot tell me whether the suite forgot the awkward path that prompted the change.

This sounds like fussy wording until something passes. Then the label becomes the memory of the work, and “build passed” has a habit of turning into “change verified” by lunchtime.

My first pass is blunt:

  • What exact question does this check answer?
  • Which inputs and configuration did it use?
  • What important behavior sits outside its scope?
  • Would an empty, skipped, or cached run still look successful?
  • What result would make me stop?

That last one matters. If I have not decided what failure looks like, I can usually explain away whatever appears. I am talented that way. Most builders are.

Compile proves the thing can be built

A successful compile or package step matters. It catches broken references, type errors, missing assets, and configuration mistakes that prevent an artifact from existing. When it fails, there is little point debating the finer qualities of the interface.

But a build result is narrow. It doesn't prove the application starts with the expected configuration. It doesn't prove a route returns useful data. It doesn't prove a user can finish the interaction the change was meant to improve.

I keep the build result named as a build result. I also check whether the command actually included the component I changed. Monorepos, optional packages, feature flags, and generated assets all create ways for a command to succeed while quietly walking around the relevant code. The green mark may be accurate. The scope may be wrong.

Tests prove their assertions

Tests earn the same precision. I look at the command, discovered test set, skips, warnings, and final status. A line that says “passed” isn't enough if the runner found nothing or excluded the area under review.

Then I ask whether the assertions reach the behavior I care about. A test may prove that a function returns the expected object while missing that the interface never calls it. Another may prove the happy path while leaving error handling untouched. That is not a criticism of automated tests. It's a reminder that the suite is an argument somebody wrote, not an oracle that wandered in from the mountains.

When a change fixes a specific class of failure, I look for a check that would have failed before the fix. If there isn't one, I want a clear reason. Sometimes the right verification is at another layer, especially for visual behavior or integration boundaries. Pretending every useful observation belongs in a unit test just produces strange unit tests.

Give the browser one small job

For interface work, I use a small browser check with a named path and visible outcome. Suppose a settings page now reveals an extra field after a checkbox is selected. I load that page, select the checkbox, confirm the field appears, enter a value, and watch for relevant console or network errors.

That check doesn't prove the whole settings area works. It doesn't prove every browser behaves the same way. It proves that one representative path worked under the conditions I used. That's enough to catch a surprising amount: missing client assets, stale selectors, runtime exceptions, or a request that failed while the screen looked mostly fine.

I write down what I actually did. “Browser checked” is how a glance at the first screen becomes folklore about the entire product. A useful note says which path ran and what counted as success. If I didn't exercise the save step, I don't imply that I did.

The browser can also disagree with the tests. Good. Disagreement tells me the layers are measuring different things, which is exactly why I kept them separate.

Check what gets delivered

The last pass follows the behavior to the point where it matters. That could mean opening the packaged application, calling the configured route, or reading back a saved value. The check depends on the change, and sometimes it simply isn't warranted. A documentation edit doesn't need an elaborate runtime ceremony.

For behavior that crosses boundaries, though, I want to see the result after those boundaries have had a chance to interfere. Build tooling can omit an asset. Configuration can point at the wrong endpoint. A successful save request can still produce a value the next read doesn't return. Each layer can be healthy on its own while the delivered path is broken between them.

So I finish with four plain entries: build, tests, browser path, delivered behavior. Each entry says passed, failed, or not run, plus the narrow claim behind it. I don't average them into a confidence score. If the browser path wasn't checked, the blank stays visible.

That makes the decision less theatrical and more useful. I can ship a small change without pretending it survived checks nobody ran, or hold it because the one untested boundary is the whole point of the change. Either way, the green marks keep their proper jobs instead of joining forces and writing a much happier story than the work supports.