The Thin Control Plane Is Usually Enough

A useful control plane coordinates intent and evidence without becoming a second execution engine or a counterfeit source of truth.

a computer screen with a bunch of data on it

A useful control plane coordinates intent and evidence without becoming a second execution engine or a counterfeit source of truth.

Operational software has a habit of gaining weight. A small surface begins by showing work from several systems. Then it gets its own statuses, retry logic, history, and eventually its own version of records that already exist elsewhere. Before long, nobody's quite sure what happens when it disagrees with the systems doing the work.

I've had better results with a thin control plane. The human ledger keeps human commitments. Source systems keep their authoritative records. Execution stays with the components built to execute. The control layer coordinates those boundaries and makes reconciliation visible.

Desired state is a request with provenance

Desired state describes what should be true, who requested it, and under which policy. It isn't proof that the change happened. For human-managed work, the durable request may remain in the existing work ledger. The control plane can reference it, interpret the relevant fields, and present the action without cloning the entire record.

That reference matters. When someone changes priority or closes the work in its normal home, the coordinating surface shouldn't keep an independent truth alive. It should observe the new intent or admit that its view is stale.

A minimal contract can stay small:

WorkRef {
  source_key
  desired_state
  observed_state
  evidence_ref
  observed_at
  policy_result
}

The contract joins information. It doesn't pretend to own every field that produced it.

Observed state belongs to evidence

Observed state says what the execution system most recently reported. That might be a completed job, a delivered configuration, or a failed attempt with diagnostic evidence. The control plane stores enough reference and timing information to explain its view. The detailed artifact stays with the system that generated it.

This avoids a trap I've fallen toward more than once: copying rich execution state into the coordinator until it becomes the only readable history. Once that happens, losing the control plane means losing the operating record, and every executor has to conform to its private database.

An observation without time is dangerous. So is one without a route back to evidence. “Complete” needs an observed-at value and an evidence reference, even if the interface keeps those details quiet.

Reconciliation should be boring and explicit

The reconciliation loop compares intent with observation, evaluates policy, and chooses a narrow next step. It may enqueue permitted work, request human attention, or wait for newer evidence. It should not smuggle a second implementation of each executor into the control service.

Conceptually, the loop is simple:

for item in active_work:
    desired = read_intent(item.source_key)
    observed = read_evidence(item.evidence_ref)
    decision = evaluate(desired, observed, policy)
    publish(decision, with_provenance=True)

The difficult work is in clear adapters and policy evaluation, not in making the loop clever. Reconciliation has to tolerate duplicate observations and repeated runs. It also has to expose disagreement. If the ledger says closed while execution reports pending, silently choosing one makes the coordinator look calm at the expense of truth.

I treat divergence as useful state. It gives an operator a specific question: is the intent outdated, is the evidence late, or did execution fail to report correctly? That's a much better starting point than a mysteriously yellow badge.

When the control plane is unavailable

Thinness becomes testable during an outage. Operators should still be able to find the human commitment in the ledger, inspect the source system, and use its native controls where policy permits. Executors shouldn't fail merely because the coordinating view is dark.

Some convenience disappears. Cross-system queues may be unavailable, and automated reconciliation may pause. That's an acceptable degraded mode if work and evidence remain intact. On return, the control plane should rebuild its view from references and current observations rather than assuming its last cache is current.

Recovery needs an ordering rule. The returning coordinator should read current intent and evidence before releasing queued decisions made from its old view. Otherwise it can create an execution surge based on requests that were already changed or resolved through native systems. I pause reconciliation long enough to refresh references; missed intervals aren't automatically work owed.

This requirement discourages hidden ownership. If the only copy of an action, status, or evidence lives inside the coordinator, it is no longer thin regardless of how small the codebase looks.

Policy is the boundary, not an execution engine

The control plane does own something important: the decision about which transitions are permitted through this coordinating path. It can determine that a request is complete enough, that evidence is fresh, or that a proposed action needs review. I record policy output beside the decision so an operator can tell why work advanced or stopped.

Those decisions should remain reproducible from the referenced inputs and policy version.

Owning policy doesn't require absorbing the machinery that performs every action. A deployment tool can deploy. A records system can retain its record. A human ledger can hold assignments and decisions. The control plane supplies the connective tissue: references, comparison, policy, and a legible next state.

The result has less to impress in a diagram and fewer split-brain arguments to settle in production. I can replace the coordinating interface, pause it, or rebuild its derived view without migrating the organization’s commitments or teaching every execution system a new home. If it can't survive that treatment, it wasn't thin. If it can, I leave it alone and let the source systems keep their jobs.