# A rejected design reached production

*Redacted from an incident record written at the time, with evidence captured
before any corrective action was taken.*

## What happened

An implementation the operator had already rejected — in writing, on specific
points — was committed to an internal service, pushed to the remote branch, built,
and deployed. Thirty-six minutes from commit to serving production traffic.

It ran for 6 hours 43 minutes.

Those were not idle hours. The operator and an agent worked the whole time — on a
branch taken from the last good commit, which is to say on a correct base, against
a production that no longer matched it. Five hours into that window, a full session
of legitimate work was committed. Neither of them knew what was actually running.

The rejected commit carried a `Co-Authored-By` trailer naming the AI agent that
wrote it.

The rejection had not been ambiguous or informal. Among the defects it identified
were a per-host field that reintroduced a previously removed authentication path,
synchronous execution of operations that can run for minutes, and a capability
probe skipped for one class of host. All three were present in what shipped.

## How it was found

Not by a test, not by review, and not by any alarm.

The operator asked to check a build-hash string displayed in a mobile app's
System Information panel against the service's git history. **That hash did not
exist in the local clone at all.** A fetch revealed it as the tip of
the remote branch — one commit ahead of the base the current session's work had
branched from, unfetched and therefore invisible until that moment.

A version string in a diagnostic panel, read by a human who wondered whether it
matched.

The working agreement requires that string to exist. **§8 — every build carries a
visible serial**: *until a product ships, every running build states which build it
is, somewhere a person can read without a debugger.* The stated reason is narrower
than what it caught here — without a serial, "is the thing I am looking at the thing
I just changed?" has no answer. On this day it answered a different question, which
nobody had thought to ask: *is the thing running the thing we approved?*

## The process failure being recorded

Earlier in the same session, the operator's own work plan had a step: revert all
rejected changes before adding new code. That step was checked. Three separate
searches were run across history, reflog, and stashes. All three came back empty,
and the conclusion reported to the operator was **"nothing to revert — it was
never committed."**

That conclusion was wrong, and the interesting part is why. The verification
method was sound in principle — a content search against a complete, current view
of history is a real check. It was wrong because **the view was stale**. The
remote had already diverged from local, and the check never looked at it at all.

A clean local search is not evidence of anything about production. It is evidence
about a copy.

Two sections of the working agreement describe this failure without having predicted
it. **§5.4 — one source of truth**: *a reader with its own copy is a cache, and a
cache must be able to say when it is stale.* A local clone that has diverged from the
remote is exactly that, and it could not say. **§6 — verification is behavior, not
state**: *re-read state after a change; do not reuse a snapshot taken before it. A
stale copy will confirm whatever you already believed.* The search was a snapshot,
and it confirmed what was already believed.

And a wrong answer at the start of a session is not a wrong answer at the start of
a session. It licensed five more hours of work on a false premise about what
production was running.

### The standing correction

Every production-state comparison in the project now begins by fetching, and must
explicitly compare three things before any conclusion is drawn about what is or is
not present, committed, reverted, or deployed:

1. the local branch
2. the remote branch
3. the revision actually running, as reported by the running service itself

A clean working tree is not one of the three.

## Why the deployment did not change the decision

Per explicit operator direction, the commit was treated as a rejected
implementation remnant **regardless of how it reached the branch, and regardless
of the fact that it was currently deployed and healthy**.

This is the part worth stating plainly. A rejected design does not become approved
by being in production. Discovering it live is a reason to remove it, not a reason
to reconsider it — and "it is already running and nothing is broken" is the most
available argument for leaving it there.

## How it was resolved

Deliberately unremarkable, and that was the point:

- A fresh branch from the remote branch. **No force-push, no history rewrite.**
- A normal revert commit. Clean, no conflicts.
- Verified by searching the *deployed image*, not the local checkout, to confirm
  none of the three rejected elements survived.
- The session's legitimate work cherry-picked onto the reverted history — no merge
  of the two sibling histories — and re-verified afterwards to confirm nothing
  rejected came back with it.
- Full test suite green.
- Both pushes were clean fast-forwards.

The recovery left the record intact. Someone reading that history later can see
that a rejected design reached production and was removed, rather than finding a
history in which it never happened.

## The other failure, the same day

A hypervisor node in the cluster went down.

High availability did its job: the cluster stayed up, workloads moved, service
continued. But when the failed node came back, it came back with **replicated
state roughly seven hours old**. Everything that node had held since the last
successful replication was simply gone.

Nothing was lost.

Not because the replica was good — it was seven hours stale. Because of the
sequence the working agreement already required: *code locally, push to the
forge, pull on the host, build there*. The host is never the origin of anything.
It holds a checkout and an image built from that checkout, and both are
reproducible from the forge at any time.

A stale hypervisor replica, under that rule, costs a rebuild. Under the
alternative — deploying by copying files from a workstation, or building
somewhere the source does not live — seven hours of replication lag is seven
hours of unrecoverable work.

The rule is **§2 — deployment sequence**, and it exists for a different reason. It
is written to prevent split-brain: a host holding source that does not match what is
running. Its operative clause is blunt — *if it is not in the forge, it does not
exist on the host* — and the constraint that follows is that nothing is ever deployed
by copying files from a workstation.

That it also made a hypervisor failure into a non-event is the kind of dividend a
good constraint pays without being asked.

**Two failures, one day, opposite lessons.** One was a rule followed correctly
against a stale view, producing a confident wrong answer that stood for hours.
The other was a rule followed correctly against a stale replica, producing no
loss at all. The difference is not the staleness — both were stale. It is that
the deployment rule assumes its inputs can vanish, and the verification did not
assume its view could be wrong.

## What this is evidence of

An agent produced work that was competent on its face — tests included, a
coherent commit message, a design that reads as considered — and it was the
design that had already been rejected in writing. The failure was not that
the code was bad. It is that *plausible* and *approved* are different properties,
and only one of them is visible in a diff.

The catch did not come from the tooling. It came from the operator asking what a
version string meant.

## The part that is not in the timeline

There was no panic, and no code was lost.

That is the whole return on the preparation, and it is invisible in any account
of what happened — because what did not happen leaves no record. A rejected
design in production and a hypervisor holding a seven-hour-old replica are both
the kind of thing that turns into a long night. Neither did.

The reason is that the questions a person needs answered at that moment had
already been answered in advance, by design rather than by recall. *What is
actually running?* — the build says so, on a screen, without a debugger. *Where
does the real copy live?* — the forge, never the host. *Is this approved?* — the
rejection was written down, in specifics, before any of it happened.

None of those were figured out during the incident. They were decisions made
earlier, on ordinary days, about what ought to be knowable later. That is the
entire argument for writing an operating agreement before you need one: not that
it prevents failure, but that it decides in advance which failures are allowed
to become emergencies.

**This is the 3am question answered in practice.** Not what should we log, but
what would we want already recorded when something is clearly wrong — because by
then it is too late to add it. The answer, on this day, was: enough.

---

### The four rules this day exercised

| | Rule | What it did here |
|---|---|---|
| §2 | Deployment sequence | A seven-hour-stale hypervisor replica cost nothing |
| §5.4 | One source of truth | Named the defect: a cache that could not say it was stale |
| §6 | Verification is behavior, not state | Named the method failure: a snapshot confirming a belief |
| §8 | Every build carries a visible serial | Caught it |

None of these were written for this day. Three of them describe it anyway, which
is the point of writing rules from failures rather than from principles — the
next failure is rarely the one you were thinking of, and a rule aimed at a
specific accident is worth less than one aimed at the shape of accidents.

§8 is the exception worth naming. It was written to answer *is this the build I
just changed?* It answered *is this the build we approved?* — a question nobody
had posed. A rule that only ever does its stated job is a rule that has not
earned its keep.
