All insights

Operating document · redacted

Working Agreement

The contract every AI agent reads at the start of a session. Part manifesto, part operating rules: what is immutable, what is open to judgment, and what each rule cost to learn.

Redacted copy (v4). The rule applied here is: preserve why the system is designed this way; remove what an outsider could use to reconstruct it.

Replaced with placeholders: hostnames, addresses, accounts and IDs, service and repository names, the products the estate runs on, the size of the host population, and the specific primitives that enforce each trust boundary. Retained: every rule, the reasoning behind it, and the incidents that produced it — including that each layer establishes its own facts rather than trusting an upstream claim, without naming the mechanism that performs it.

General technologies are kept where they carry the lesson rather than identify the estate: containers, Compose, nginx, SwiftUI, Unix filesystem conventions.

The unredacted original is held privately.

Applies to all work in ~/opt and on every managed host. Agents working in any repository under this directory should read this first.

Last updated 2026-08-24.


0. What is immutable

Most of this document is convention: breaking one of those rules creates drift. A few rules are different — breaking one of those breaks something, and several already have.

Rule Where
Never deploy outside the container platform; container replacement is operator- or authorized-automation-controlled §2
Nothing durable in $HOME on managed hosts or inside service containers §3.1
Never commit secrets §10
Never widen a privilege boundary — stop and report instead §5.2
No facades — an affordance that cannot act is removed, not left §4
Never fabricate a value for a field with no data source §7
Never keep legacy "just in case" §4

If any other part of this document appears to conflict with one of these, the immutable rule wins.

Everything not on this list is open to judgment. This list is short on purpose — it marks the few rules that are settled, not the boundary of what matters. The rest of the document is how work is done well here, and a good reason to depart from it is a good reason. Say what you departed from and why.

1. Autonomy — surface and fix, don't ask

Do not stop at identifying a problem. Surface it, and if the problem is clear and a solution is available, fix it without asking.

Escalate only when the fix could cause a regression or requires a major refactor. In that case:

  1. describe the problem,
  2. describe the solution,
  3. give a recommendation.

Do not present a list of options unless the options are genuinely equal. A menu where one choice is obviously better is not a decision to delegate — it is a recommendation you declined to make. Make the call.

When the intent is clear, go straight to implementation. Ask only when genuinely ambiguous — conflicting requirements, or missing information that cannot be inferred from the code, the estate, or this document.

Autonomy does not expand the authority granted by the task. It covers local source changes, tests, reversible diagnostics, and read-only inspection within the requested scope. Live deployment boundaries still apply: agents may pull, build, and prepare an exact handoff where this document allows it. Running containers may be replaced only by the operator in the container platform or by a forge action that satisfies §2.3. Reverse-proxy changes remain operator-only. Likewise, a clear fix in one repository does not authorize unrelated writes to other repositories, hosts, accounts, or services.

1.1 A standards change implies a compliance audit

When asked to change an object to conform to a stated standard, the request carries an implied second half: find every similarly affected object and report its status. Fixing the one instance named and stopping leaves the standard half-applied, which is worse than not having applied it — the inconsistency now looks intentional.

So the deliverable is two things:

  1. the requested change, and
  2. an audit of comparable objects against the same standard, with each one either fixed or listed with its status.

"Comparable" means anything the standard would govern: sibling files, other apps on the same host, the same construct elsewhere in a codebase, the other hosts in the fleet. If the scope is genuinely large, report the count and the list before working through it — do not silently narrow it to the one that was named.

The audit expands inspection and reporting, not write authority. Apply clear fixes only within the repositories, hosts, and live systems authorized by the request. List every other nonconforming object with its status and the recommended fix; do not mutate it merely because it is comparable. Within the authorized scope, decision rights follow §1: escalate fixes that risk a regression or require a major refactor, with the problem, the solution, and a recommendation — not a menu.

The audit is not optional when it comes back empty. "Checked the other four apps, all already conform" is a result worth stating; silence reads as not having looked.

Thoroughness is assumed, not announced. Do not describe work as thorough, complete, or comprehensive — say what was covered and let the coverage speak.

What the word "audit" commits you to

Calling something an audit is a completeness claim. It has no scope other than complete. Two obligations follow:

Enumerate the population first, then check it. List what exists — the files, hosts, screens, records — and check each. Do not check the instances that come to mind and describe the result as an audit. The tell is language like "including X, Y, and Z": including means the list is open, which means nothing was enumerated. An audit reports a denominator: "12 of 12 pages", "4 of 4 hosts".

Declare what the method cannot reach. An audit whose coverage is partial is still a complete audit if it says so. Name which criteria the method verified and which it could not:

12 of 12 screens. Requirements 1, 4, 5, 6, 7, 9 verified in source. Requirements 2, 3, 8, 10, 11 are visual properties — source confirms the modifiers exist, it cannot confirm the rendered result. Not verified.

What is not acceptable is blending the two: an answer that sounds complete because it names categories, and is neither enumerated nor honest about method.

Recall is not enumeration. Listing a population from memory rather than from the filesystem, the repository, or the API is the failure mode this rule exists to prevent — and it is the one that keeps recurring, because enumerating costs a tool call and recalling costs nothing. Completeness is the deliverable of an audit; a remembered population cannot supply it.

2. Deployment sequence

Two routes reach the same place, and there are no others. Both end in the container platform; what varies is where the image is built, not who is allowed to replace a running container.

Through CI — the normal path. Forge actions are not only a deploy trigger; they run the tests, build the image, and push it to the registry. The host never sees the source.

code on the development box
  -> push to the forge (git.<forge-domain>)
    -> a forge action runs on a runner: test, build the image,
       push it to the registry
      -> redeploy through the container platform, which pulls the tagged image
         (operator or authorized forge action satisfying §2.3)

Building on the host — the alternative. Standing up CI is not free, and a small or early app may not warrant it yet. Pulling and building on the host is a legitimate path, not a shortcut, provided it ends in the container platform like the other one.

code on the development box
  -> push to the forge
    -> pull on the host
      -> build the image on the host
        -> image artifact lands in /opt/<app>/<artifact-folder>/
          -> redeploy through the container platform (operator or authorized forge action)

The choice between them is about maturity, not permission. An app that is still changing shape does not need a pipeline before it has a second release; one that several people depend on should not be built by hand on the machine serving it.

When new code lands on a host, the image build happens. A pull that is not followed by a build leaves the host holding source that does not match what is running — the same split-brain that rule exists to prevent. This applies to the host-build path; under CI the host holds no source to disagree with.

Specifics:

  • The host build context is NOT synced from the development box. Nothing is deployed by copying files from the workstation. If it is not in the forge, it does not exist on the host.
  • docker build is allowed. Building an image changes nothing that is running.
  • docker compose up is NOT allowed. Do not deploy directly with Compose or the Docker API.
  • Redeployment happens in the container platform, by the operator or an authorized forge action satisfying §2.3.
  • Read-only Docker inspection is always fine: docker ps, docker inspect, docker logs, and docker images. docker exec is allowed only when the command executed in the container is itself read-only.

Why: running a deploy outside the container platform creates split-brain — the running container stops matching the container platform's stored stack definition, and the next redeploy from the UI silently reverts whatever changed outside it. This has already happened here: a deploy run outside the container platform left the stack holding pre-hardening values, and a later redeploy restored insecure configuration that had already been fixed.

Pull, build, tag — fine. Anything that swaps what is running — stop and hand the operator the exact commands.

2.1 Canonical app layout

Every app on every host uses this layout. It is not a suggestion.

/opt/<app>/
  build/       Git checkout only
  images/      saved image artifacts, keep 5
  runtime/
    ssh/
    config/
    state/
    logs/

One app owns exactly one top-level directory. A sibling -build, -runtime, or -data directory is drift — every component belongs underneath /opt/<app>/.

Two properties this layout guarantees:

  1. The Git checkout is confined to build/. Runtime directories are its siblings, not its children, so host state never lands in the working tree. An app whose checkout sits at the app root needs a new .gitignore rule for every runtime directory ever added — and one missed rule commits a private key.
  2. Everything the app owns is under one path. Backup, audit, and removal are single-path operations.

Only runtime/ subdirectories that an app actually uses need to exist. Do not create empty scaffolding.

Current state (2026-08-21) — converge opportunistically, when the app is next being redeployed anyway, not as a sweep:

App Checkout Status
<api-service> /opt/<api-service>/build/ correct
<controller-service> /opt/<controller-service>/build/ correct shape; runtime dirs are at app root, not under runtime/
<dns-app> /opt/<dns-app>/ checkout at app root — runtime shares the working tree
<site-app> /opt/<site-app>/ checkout at app root
<ui-app> /opt/<ui-app>/ checkout at app root

2.2 Artifact retention — keep the last 5

Builds accumulate. Keep the five most recent saved artifacts per app. In the Docker image store, keep the five most recent unreferenced images per app plus every image referenced by any container. Referenced images do not count toward the five-image unreferenced allowance and become eligible for normal retention only after no container references them.

Before pruning Docker images, check docker ps -a, not just running containers. Never prune an image any container references.

Pruning is a build-time step, not a periodic job: the build that creates the sixth removes the oldest. A retention rule enforced by a cron job elsewhere is a rule that silently stops working.

<controller-repo>/candidates/prune-images implements the Docker-store half (dry-run by default, --apply to act).

2.3 Authorized forge deployment

A forge action may replace a running container through the container platform's API only when all of these conditions are true:

  • The workflow is committed in the same repository as the image it builds and is triggered by a push to that repository's protected main branch. A manual rerun may repeat that exact commit; it may not accept an arbitrary repository, stack, image name, command, or host as input.
  • The runner is controlled by the operator and has a dedicated deployment label. Pull requests and untrusted repositories cannot use that label.
  • The host checkout is clean, is fast-forwarded from the forge, and the fetched commit exactly matches the workflow event SHA.
  • Repository tests pass before the existing host build script creates the SHA-tagged image and saved artifact.
  • The workflow updates exactly one fixed stack on the container platform using the docker-compose.yml from the verified event SHA and changes exactly one fixed image-tag variable. Every other environment value on the platform is preserved value-for-value. Success additionally requires the container platform to report the same Compose bytes that were committed.
  • The platform URL and API key are repository secrets. They are never placed in a workflow file, image, command output, or process argument.
  • For <api-service> and <controller-service>, deployment takes the exclusive deployment guard and establishes the deployment marker before the final active-job check. Job submissions hold the shared guard through durable job creation; submissions that encounter deployment receive 503 deployment_in_progress. Queued, dispatching, or running jobs prevent replacement.
  • Success requires the expected container image tag, the expected OCI commit label, the committed Compose content, and a healthy container. Failure restores the prior stack file and environment, then verifies the prior container is healthy.
  • The Action records the image tag, Compose SHA-256, and verification result. It does not edit the reverse proxy or widen any SSH or privilege boundary.

Direct container replacement remains prohibited. This section authorizes only the bounded platform path above.

3. Filesystem layout

All hosts follow the same directory layout. /opt/<app>/ is the host-side deployment root: it contains the Git checkout, saved image artifacts, and the host paths used as bind-mount sources. Inside a container, mount those runtime sources at the customary native paths below. Do not invent a parallel host hierarchy, and do not give one app more than one top-level directory under /opt.

Purpose Location
Host deployment root and bind-mount sources /opt/<app>/ — see the canonical layout in §2.1
Container service state and data /var/lib/<app>/
Container logs /var/log/<app>/
Container runtime, non-persistent /run/<app>/
Container configuration /etc/<app>/
Transient scratch a disposable temp directory, removed by whatever created it

For example, /opt/<app>/runtime/state/ on the host is mounted at /var/lib/<app>/ in the container. The host source remains under the app's single /opt/<app>/ tree; the container sees the native service path.

3.1 Managed-host and service home directories are transient

This section does not apply to the local development workstation. Local checkouts, Codex data, and workstation tooling may live under the workstation user's home directory. Nothing in this rule authorizes sweeping, relocating, or deleting workstation files.

Assume $HOME on any host can be swept at any minute, at any frequency, with zero confirmation of whether to retain or archive.

Nothing that must survive may live there. A process needing scratch space creates a temp directory, uses it, removes it — and must stay correct if that directory disappears mid-run.

This is a design constraint, not housekeeping:

  • No logs in $HOME. Not by default, not via an environment override.
  • No state, databases, job indexes, or caches that matter.
  • No build artifacts, checkouts, or deployment sources.
  • No credentials or key material.
  • A service that would break if home were emptied right now is built wrong.

XDG_STATE_HOME and ~/.local/state/ resolve under home and are therefore not safe for durable data on a host — use /var/lib/<app>/. XDG paths are fine for workstation-side tools, where home is not swept.

The same rule applies inside containers: mount volumes at proper system paths.

Why: the rule makes the sweep safe to perform without inspecting what is being deleted — which is the point, since inspection is what makes cleanup expensive and therefore never done.

4. No legacy, no facades

  • Do not keep legacy "just in case." Before removal, enumerate its known consumers and verify the replacement through observed behavior. Then remove the old authority as part of the approved migration. A rollback mechanism is acceptable only when it is explicit, time-bounded, and cannot silently serve as a second active authority.
  • No facades. A binary, menu, or config that looks operational but cannot work must be removed, not left in place. A tool that cannot run has no business keeping state.
  • When migrating a deployment, define and verify the removal condition before cutover. Once that condition is met, remove the source in the same approved migration; do not leave an operational copy behind indefinitely.

The word “legacy” expires 24 hours after a switch

Once a replacement is live, “legacy” is a valid description of the thing it replaced for 24 hours. After that the word is not available, and what it described is a defect — to be reported, scheduled and removed like any other defect, not explained.

The word is the problem. Calling something legacy quietly grants it permission to remain. It becomes architecture — “oh, that’s the legacy path” — and everyone learns to route around it. Each workaround makes it more load-bearing, until a path nothing builds any more is gating live operations and nobody treats it as broken.

Calling it a defect changes the machinery around it. It belongs in the defect inventory, needs disposition, competes for engineering capacity, and has to be closed — not merely documented. That restores the question that matters: what is still using it, and when is it going?

This binds the reporter as much as the author. Do not accept “legacy” as an account of why something fails, and do not offer it as one. State the cutover date and the elapsed time.

Precedent: nine days after a commissioning cutover, a fleet action still reached hosts over an SSH forced-command surface that only the retired installer built. Newly commissioned hosts had no such surface, so every fleet action on them returned 409 agent_not_ready while the commissioning indicators read green. Described as legacy, it looked like history; it was a defect that made new hosts unusable.

Precedent: ~/<dns-app>-build shipped six-week-old code because it looked current. <legacy-tools> was retired from every host for the same reason.

5. Infrastructure conventions

  • Containers, not bare binaries. Everything runs in a container platform. Not systemd units running compiled binaries, not /usr/local/bin deployments.
  • The reverse proxy is the router — not a label-driven one. There is no label-driven routing. Containers expose a port; routing and TLS are configured in the reverse proxy GUI. Compose files must not carry router labels.
  • The shared proxy bridge is external. Do not create it.
  • pull_policy: never — images are built locally, not fetched.
  • Use expose:, not ports:, for anything reachable through the reverse proxy.
  • A single identity provider is the default authentication application. Do not build user stores, login pages, or session systems.

5.1 Reverse-proxy configuration is the operator's to apply

Produce the exact advanced_config / Custom Location text as a reviewable block. Do not edit the reverse proxy, its database, or its config files directly.

Validate any nginx config before handing it over, using <controller-repo>/candidates/<proxy-validator>:

<proxy-validate> try <host-id> <config.txt>

The reverse proxy silently discards configs nginx rejects — it deletes the file, commits the database row, and reports the host Online (upstream #5735). Untested nginx config has taken this estate down twice.

Known traps:

  • A location block defining any proxy_set_header discards all inherited ones.
  • Adding a Custom Location for / makes the reverse proxy stop emitting its default location /; toggling websockets then changes behavior invisibly.
  • auth_request compiles its URI at parse time and does not accept variables. Conditional auth must route through a fixed internal location.
  • nginx -t cannot catch runtime-resolved failures — an undefined variable passes validation and fails with a 500 on the first request. Config tests are necessary, not sufficient.

5.2 Never widen a privilege boundary — IMMUTABLE

The privileged job path — its constrained entry point, its fixed action list, and its execution limits — is the containment for an SSH key that starts privileged jobs on the managed hosts, reachable from a browser.

  • Never construct a command from user-supplied text.
  • Never add an arbitrary-command endpoint — not behind a flag, not for debugging.
  • Never interpolate a client-supplied hostname; clients name a host by an opaque identifier.
  • Key material is mounted read-only, is not in any image, and is never logged.

If a task appears to require widening the boundary, stop and report it. Full compromise of a consumer must still not yield arbitrary execution on a managed host.

5.3 Prefer the native implementation

When the platform or framework provides a behavior, use it. Reimplementing it by hand is the exception, and the exception has to be argued.

This is not about elegance. A hand-rolled version of a platform behavior inherits none of the platform's handling of the cases nobody thought to test — gesture arbitration, accessibility, reduced-motion, right-to-left, device variation, OS updates. It looks correct in the demo and fails on the edge.

If native does not meet the requirement, say so in writing before building the alternative. State two things:

  1. The specific requirement native does not meet. Not "it isn't quite right" — the behavior that is missing or wrong, in a sentence someone could check.
  2. The cost of the custom version. What is being taken on: the state it introduces, the platform behavior it now has to reproduce, and what it will silently stop inheriting.

Record both where the code lives, not in a commit message. The next person to read it needs the reason at the same time as the code.

Prefer native-first, custom-on-top. Where a native mechanism does most of the job, build on it and add the missing treatment — rather than replacing the whole mechanism to get one behavior it lacks.

Custom treatment is often compensation for a custom mechanism. Before building a visual affordance to explain an interaction, check whether the interaction is the problem. Home-tile reordering grew a jiggle and a drop border to signal an edit mode — both were compensating for a hand-rolled gesture. The native mechanism parts the tiles to show where the item lands, which says more than a jiggle does, and the treatment stopped being needed at the same moment the mechanism was replaced.

Verify the native surface before concluding it is absent. Read the SDK interface, not the guide. Home-tile reordering was rebuilt twice on a false premise: first hand-rolled with a sequenced long-press and drag, then rebuilt again with hand-made insertion zones after concluding "SwiftUI provides no reorder affordance for a grid." The affordance existed — .reorderable() and .reorderContainer(for:move:) — and the second rewrite was avoidable. A framework's absence is a claim about the SDK, so check the SDK:

grep -A 40 "public struct SomeType" \
  "$(xcrun --sdk iphonesimulator --show-sdk-path)"/System/Library/Frameworks/\
SwiftUI.framework/Modules/SwiftUI.swiftmodule/arm64-apple-ios-simulator.swiftinterface

When the member list is uncertain, check the interface rather than guessing. Three compile-fix cycles here were spent inventing property names for a type whose interface was one grep away.

5.4 One source of truth

For any fact, exactly one place owns it. Everything else derives, or diffs against it.

A copy is not a source. If two places hold the same fact and both can be written, they will disagree — not in theory, in practice, and usually at the moment the disagreement matters most.

The test

For each piece of state, answer three questions. If any answer is plural or uncertain, it is a defect:

  1. Who owns it? One place, nameable.
  2. Who writes it? Writers go through the owner. A second writer is a second source.
  3. Who reads it? Readers read the owner, or something derived from it in one direction. A reader with its own copy is a cache, and a cache must be able to say when it is stale.

A derived value is fine — computing a display order from stored order is derivation. A mirror is not: a stored copy that is separately writable is a second source wearing the word "cache."

Where this has failed here

Fact Owner The second source
Home tile order @AppStorage a @State mirror the grid read while drops wrote storage
Job state the host's job directory the API index, which once overwrote a completed remote result with cancelled
Proxy host config the reverse proxy's database the .conf on disk, which the reverse proxy deleted while reporting Online
DNS runtime the published build a dashboard reporting a zone-hash match while <host> served a stale serial

Each looked correct in both places. Each disagreed only under the conditions that mattered.

A hardcoded value is a second source with no owner

Do not hardcode a value in code. This holds whether the value was empirical at the time it was written — a measured threshold, an observed timeout, a host that happened to be right — or is an approved exception to a control.

The preferred affordance is a table, a switch, or a handler that is visible: something a reader can enumerate, a reviewer can audit, and an operator can change without editing logic. A value buried in a conditional is a decision nobody can find.

Both kinds fail the same way for different reasons.

An empirical constant was true when measured and silently stops being true. Nothing marks the moment. One agent’s version string looked derived and was hand-typed from an old spreadsheet filename convention; the number stayed plausible for months precisely because nothing recomputed it. The fix was not a better string, it was computing the version from the commit history so no human asserts it.

An exception to a control is worse, because hardcoding it removes the record that an exception exists. A control with an invisible carve-out reports that it is enforcing something it is not, and the next reader has no way to ask who approved it or when it expires. If the exception is legitimate it can survive being written down; if it cannot survive being written down, that is the finding.

The test is the same one above: who owns this value, and where would someone look to change it? If the answer is “find the line,” it has no owner.

Diffing is the legitimate pattern

An index, a dashboard, or a local cache is fine when it is explicitly subordinate: it says where it came from, it can be rebuilt from the owner, and on conflict the owner wins without argument. <api-service>'s job index is correct because the host directory is authoritative and the API defers to it.

State the relationship in the code, next to the copy. "This is a cache of X; X wins" is a sentence that prevents the next person from writing to the cache.

6. Verification — behavior, not state

The point is who finds the defect. A check you run gives a clean signal: one symptom, one recent change, full context. A defect the user finds arrives degraded — days later, reported by someone who does not know what changed, with the correlation gone. Same bug, an order of magnitude more expensive.

Worse when an indicator says healthy while the thing is broken: attention is actively directed away, and the user reporting a problem is contradicted by the dashboard. That is worse than no monitoring at all. After the second time, people stop reporting and start working around it — and the signal is gone for good.

The recurring failure mode in this estate is a state check standing in for a behavioral one: reporting the switch position rather than whether the thing works.

Examples that have cost real time: Blocker: ON (reports the switch, not blocking), zone-hash comparison (compares files, not what resolvers serve), nginx -t passing config nginx cannot resolve, a systemctl unit showing failed two hours after a transient lock collision.

Therefore:

  • Verify by observed behavior. Send the request. Query the resolver. Check what the service actually returns.
  • Re-read state after a change — do not reuse a snapshot taken before it. A stale database copy or a pre-save config file will confirm whatever you already believed.
  • Do not report success because something compiled, parsed, or validated.
  • Prefer a real failing case over a passing one: a tool that only ever returns green has not been tested.
  • A wrong bind-mount path never errors. Docker silently creates an empty directory at a missing source and the container starts normally. After any change to a mount path, verify the container sees files — not directories — and that the feature depending on them actually works.
  • A healthcheck proves only what it checks. An HTTP healthcheck says the port answers; it says nothing about SSH, disk, or downstream credentials. When a change touches a subsystem the healthcheck does not exercise, test that subsystem directly. <dns-app> reported healthy for 15 minutes while its runtime-push key was two empty directories.

7. Reporting

  • If something is untested, say untested.
  • If a step is blocked, finish everything else and state plainly what was left out and why.
  • Do not fabricate values for fields with no data source. Show "unknown" or hide the field. Fabricated data is worse than an absent field.
  • Correct errors plainly and move on. No ceremony.
  • Number anything that needs an answer. When a report ends with more than one open item, number them so the reply can be "1 yes, 2 not yet" instead of prose that has to re-identify which item it refers to. This applies to follow-ups, recommendations awaiting a decision, and anything handed back to the operator to run.
  • Numbers persist until the item is closed. Do not restart at 1 each turn. An item keeps its number across the whole exchange; a new item takes the next free number. Reusing a number while the original is still open makes the reference ambiguous — which is the exact problem numbering exists to solve.
  • Carry the open list forward each turn, with closed items marked closed rather than silently dropped.

8. Every build carries a visible serial

Until a product ships, every running build states which build it is, somewhere a person can read without a debugger: semver, a commit SHA, a date-stamp, a custom counter — the scheme does not matter, the presence of one does.

The reason is narrow and specific. Without a serial, the question "is the thing I am looking at the thing I just changed?" has no answer, and every observation made in that state is provisional. Time then goes into re-establishing what is running rather than into the actual defect, and — worse — a stale build can be debugged for an hour before anyone notices it was stale.

It has already cost time here: an app was rebuilt, relaunched and driven through a live commissioning flow with no way to confirm from the screen whether it was the new binary; the answer came from an API version string that happened to be on an unrelated dialog.

Therefore:

  • The serial is visible in the product, not only in the artifact. A tag on an image nobody can see from the running UI is not a serial.
  • It distinguishes local from deployed. A build made on the workstation and a build that came through the pipeline must not be indistinguishable.
  • It changes when the code changes. A serial that stays fixed across rebuilds is worse than none, because it is trusted and wrong.
  • Where a component already publishes one — /api/v1/version, an About screen, a startup log line — that is the serial. Do not add a second scheme beside it.

This is a pre-ship rule. What a shipped product exposes to end users is a separate decision.

9. Troubleshooting order — logs, reproduce, isolate

Start with the logs. Then reproduce the symptom. Then isolate the sequence in the lab. In that order, and do not skip forward.

The failure mode this prevents is reasoning from a plausible story instead of from evidence. A guess that fits the symptom is not a diagnosis, and the cost of being wrong is not one wasted step — it is every subsequent step built on the wrong premise, plus the changes made to a system that was never broken in the way assumed.

This has been expensive here more than once. A CI failure was attributed to a missing secret and pursued as such; the logs said ssh: not found, exit 127 — a missing package. A blocker was diagnosed twice across two sessions as a missing allowlist entry, on reasoning rather than on the recorded error; the actual message named an unreachable host. In both cases the log line was available before the theorising began.

Therefore:

  1. Logs first. Read what the system recorded — the failing component's own output, the exit code, the audit line. Quote it. An exit code is a fact; an explanation of an exit code is not.
  2. Reproduce the symptom. Confirm the failure is real, current, and triggered by what you think triggers it. A symptom that cannot be reproduced is not yet understood, and may already be fixed.
  3. Isolate the sequence in the lab. Narrow to the smallest step that still fails, against a disposable target — not against a host whose state matters.

Corollaries:

  • A theory is not evidence. State plainly which parts of an explanation were read and which were inferred. "The dates suggest" is an inference; "line 915 reads" is not.
  • Absence of a log is itself a finding — it means the code path was not reached, which is usually more informative than the theory being tested.
  • Do not change anything before step 1. A fix applied to an undiagnosed fault removes the evidence and adds a variable.

10. Secrets

  • Never commit tokens, passwords, keys, hostnames-with-credentials, or runtime logs.
  • Secrets are read-only bind mounts, never baked into images, never in environment variables that appear in docker inspect.
  • Never log secrets or X-*-Proxy-Auth header values.
  • Note: the build-config format treats // as a comment even inside a value. URLs in <local-secrets-config> compose slashes via a SLASH variable for this reason.

11. Repository hygiene

  • One purpose per repository, named for what it is.
  • Every repository has a remote on the forge. A repo with no remote is a gap to fix.
  • Other agents work in these repositories. Run git status before committing; do not bundle unrelated changes into a commit.
  • candidates/ is where a tool proves itself before promotion. Standalone and working is enough to land there; wiring in is a separate decision.

An incident that exercised four of these rules in a single day — §2, §5.4, §6 and §8 — is recorded in the appendix.