All insights

Operating document · redacted

What you wished were available (it was always there)

3:12am on the incident bridge, and nothing captured what normal looked like.

Redacted copy (v4). The rule applied here is: preserve why the system is designed this way; remove what an outsider could use to reconstruct it.

Replaced with placeholders: hostnames, addresses, accounts and IDs, service and repository names, the products the estate runs on, the size of the host population, and the specific primitives that enforce each trust boundary. Retained: every rule, the reasoning behind it, and the incidents that produced it — including that each layer establishes its own facts rather than trusting an upstream claim, without naming the mechanism that performs it.

General technologies are kept where they carry the lesson rather than identify the estate: containers, Compose, nginx, SwiftUI, Unix filesystem conventions.

The unredacted original is held privately.

Recorded 2026-09-07, from the agent-update session that produced <findings-doc>.md.

The framing

Not “what should we log” but what would we want already recorded at 3am, with something clearly wrong — because at that point it is too late to add it.

That session is the worked example. The failure was live and reproducible the whole time, and diagnosing it still took hours: the sequence had to be reconstructed from fragments across five separate stores, and the decisive fact only came from a temporary traceback.print_exc() added to a file on a production host.

The problem was not missing logs. Every layer logged. The problem was outcomes recorded without the context to correlate them.

The constraint this must respect

No single component sees the whole chain, by design. <api-service> decides authorization and never touches the host. The gateway resolves identity and cannot know what <api-service> decided. The host executes and knows nothing of the principal. That separation is what makes the trust boundaries real, and it must not be softened for the convenience of debugging.

So the goal is not a central log with a view of everything. It is a correlator: one identifier threaded through, so the pieces can be assembled afterwards without any component being handed visibility it should not have.

That distinction is the whole design. A system that logs richly but cannot correlate produces exactly what this session hit — a pile of true statements that take hours to line up.

Why no component sees the whole chain

Three separate reasons, worth keeping distinct because a proposal can satisfy one and violate another.

1. A component that sees the whole chain can be lied to about all of it.

The gateway takes the peer's identity from a kernel-verified property of the connection itself — its own docstring says “never from any HTTP header.” If it accepted an upstream assertion of who the caller is, compromising <api-service> would compromise every host, because <api-service> would be telling each host what to believe.

Instead every layer derives its facts independently. <api-service> verifies the caller's assertion itself. The gateway resolves identity from what is presented on its own connection. <job-exec> re-validates the action against its own list rather than trusting what it was handed. Compromising one layer yields what that layer can do, not what it can claim.

2. Visibility is itself a capability.

A component that can see the whole chain necessarily holds the material to reconstruct it — principals, decisions, targets. That makes it worth attacking, and it means a read-only compromise yields the entire operational picture. Keeping each layer's knowledge to what it needs means there is no single thing worth stealing.

The estate applies this elsewhere. <deploy-command> receives only deploy <app> <sha>; it cannot be told which stack, image or endpoint, because those are fixed in its own code. A caller that knew more could ask for more.

3. Independent accounts are what make a claim checkable.

If one component asserts the whole story, nothing can contradict it. When four components each record their own view, a discrepancy between them is itself a signal — the gateway reporting a resolved identity with no matching <api-service> authorization would be evidence of something wrong that no single log could show.

This is exactly why a correlation ID is safe where shared visibility is not. An opaque token lets an investigator line up four independent accounts. It does not let any component read another's.

The trade-off is real, and this session was the bill: hours spent reconstructing a sequence that was happening live. That is the cost of the property, not a defect in it — and it is why the answer is threading an identifier rather than pooling the data.

What already works, and is worth saying

More of this exists than the session's difficulty suggested:

  • <api-service>'s audit table captured the requesting principal, action, target host and outcome for the successful update, one second before the services cycled. That is the right fact at the right layer.
  • The gateway logs identity resolution — resolved '<host-id>' -> <local-identity> (supplementary groups resolved); account validation passed — which is precisely what its layer knows and nothing more.
  • Job IDs already exist on both sides.
  • The uniform external denial is deliberate and should stay. The gateway normalizes every refusal so a caller cannot distinguish policy from parse from transport; the detail lives in the protected log. That is correct, and it is not in tension with anything below — the ask is richer logs, never a richer response.

What is missing

1. A correlation ID that survives every boundary

A job ID exists in <api-service> and in the host's own store, but it does not appear in the host's system journal, and nothing ties a single operator action to the records each layer wrote about it.

One identifier, generated once at the app or <api-service>, carried through every hop and written into every log line about that operation. Then one incident is one query, not six timestamp comparisons across five stores.

This gives away nothing. An opaque ID tells a component only that two events are related — not who the principal was, not what was decided elsewhere.

2. Refusal reasons at the point of refusal

Three layers refused this operation and none said why:

Layer Said Knew
<installer>.py:350 "unavailable or interrupted" FileNotFoundError on a known path — no delta published
<installer>.py:308 "enrollment did not complete" which of three conditions failed
<agent-update>.py:217 "refused or failed" Refused: private material must be root-only: <private-job-store>

The third is the root cause of the whole session. It was one string away from being obvious, and instead cost hours.

These are internal logs, not responses. Saying why in the log does not weaken the uniform external denial.

3. The rejected value, not just the rejection

Which action, which host, which base revision did not match, which path failed a permission check. Never the credential, and never anything the external response should not carry — but the log should name what was wrong, not just that something was.

4. Device, source IP and user agent

See finding 7 of the companion document. Capture, do not gate; gating is device-management policy and lives outside this codebase.

The test for anything proposed

Would it have shortened this session?

  • Device and IP: no — but they close fraud cases, which is a different and real need.
  • A threaded correlation ID: yes, substantially.
  • Honest refusal reasons: yes, almost entirely. The root cause was a one-line message away throughout.

What this is not

  • Not a central log aggregating what each component knows. The separation is the design.
  • Not a richer external response. The uniform denial stays.
  • Not device gating.

What this looked like in practice, on a day when it was needed, is in the appendix.