The Dashboard Is Not the System

A green runtime indicator answered a narrow question. The browser was asking the one that mattered.

graphs of performance analytics on a laptop screen

A green runtime indicator answered a narrow question. The browser was asking the one that mattered.

The dashboard was green. The browser was showing a 502.

I had both open at once, so I couldn't blame a delayed refresh or a stale screenshot. According to the fleet view, the service was running. According to the route a person would actually use, it wasn't available. Both observations were accurate. I'd simply presented the first one as if it contained the second.

The initial dashboard had been built around data that was easy to collect. It could ask the container runtime whether a workload existed and whether its process was running. Those are useful facts. A stopped container deserves attention, and a restart loop is worth seeing. But “running” describes a process boundary. It doesn't tell me whether the request path reaches that process, whether the proxy has a valid upstream, or whether the application can return a usable response.

That distinction sounds painfully obvious now. It wasn't obvious in a neat grid of green status marks. Presentation can give weak evidence more authority than it's earned. Once a runtime fact is labeled “healthy,” the eye stops asking what was actually measured. Mine certainly did.

I started the second pass from the outside. For each delivered route, the check made an HTTP request through the same public-facing path a browser would take. That immediately exposed multiple 502 responses sitting behind green runtime state. The containers had not lied. The dashboard had combined “process is running” and “service is reachable” into one state, then selected the more reassuring label.

The first useful correction was to stop collapsing those observations. Runtime health and endpoint health became separate fields. A workload could now be running while its route failed. That combination was no longer an awkward edge case; it was a valid state the interface had to show plainly.

This changed more than the color of a badge. It changed the questions available during diagnosis. A failed runtime check points toward the workload or its host. A running workload paired with a failing endpoint points toward a different path: routing, proxy configuration, application readiness, or something between the request and the process. The dashboard didn't need to diagnose all of those causes. It needed to preserve enough truth that an operator would start in the right neighborhood.

There was another flaw hiding in the original aggregate. A fleet-level total could report that most services were healthy, but there was no reliable path from that number back to the route observations that produced it. The count looked precise while being difficult to audit. If one service appeared in more than one place, or if a route had disappeared from discovery, the summary could drift without offering a clue about why.

So the second correction was traceability. Each endpoint result belonged to a specific delivered route, and the aggregate was derived from those results. Clicking or inspecting the summary could lead back to the route, its last observation, and the type of check performed. “Eight healthy” was no longer a free-floating claim. It meant eight named routes had returned acceptable results under a defined check.

That route-level model also forced a useful decision about unknowns. If the system knew about a workload but had no route to test, it could not call the service reachable. If a route had not been checked recently, its last success could not remain green forever. Missing evidence needed its own visible state. Otherwise absence would quietly inherit the color of the last convenient fact.

At this point, I was tempted to make the dashboard smarter than it needed to be. We could've added elaborate scoring, weighted health, or a single calculated confidence number. That would've recreated the original problem with better arithmetic. A composite score still hides which layer failed. Two plain observations were more useful: what the runtime reported, and what the route delivered.

There's a broader design lesson here, though it's easy to make it sound grander than the work. Operational interfaces are full of proxies. Queue depth stands in for flow. Process state stands in for service health. A successful command stands in for a completed outcome. We need proxies because no interface can display the whole system. Trouble starts when the label forgets the boundary of the measurement.

I now ask a blunt question when reviewing a status surface: what exact observation earns this word? If the word is “available,” a running process is incomplete evidence. If it's “current,” a timestamp needs a freshness rule. If it's “complete,” there should be a delivered result to inspect. I'm not chasing semantic perfection. I'm trying to keep the interface from upgrading a narrow fact into a broad promise.

The corrected dashboard did not make every service green. It made the disagreement legible. A route could be failing while its container continued to run, and the summary showed that state without rounding it into either success or total failure. Every aggregate could be traced back to the route checks beneath it. The browser and the dashboard were finally describing the same system at different layers.