The failure that cost me the most this year never broke anything.

A data feed in my own system stopped working at the end of June. I didn’t find out until late September. For twelve weeks the job that depended on it ran every day, wrote fresh files, exited cleanly, and put a green light on my dashboard — because when the live feed died, the job fell back to a slower source I’d built as a backup, and the backup worked.

That is what a fallback is supposed to do. It is also why nobody noticed.

Once is an accident

The day I found that one, I went looking for the same shape elsewhere, and found three more before the day was out.

A scoring job that reported OK on every run while scoring nothing. The individual failures were being caught and logged, and none of them ever reached the exit code, so the run as a whole looked fine.

A report that fired a failure alert on every run for a condition that wasn’t a fault at all — it needed history that predated the data, and that history was never going to appear. Eleven identical alerts about one standing fact, sitting in the channel where real failures were supposed to show up.

And a document fetch that had only ever retrieved the index entry, never the document itself. Everything downstream ran on the thinner data and reported low confidence, which is technically honest and practically invisible.

Four failures. Not one of them crashed.

Freshness checks can’t see this

I had monitoring. It checked what most monitoring checks: did the job run, and is the output recent. A fallback defeats that by construction. Its entire purpose is to keep the output recent when the real path is gone.

So the check that should have caught the problem was the check the fallback was designed to pass.

The fallbacks were right

I want to be fair to the design, because the lesson isn’t “don’t build fallbacks.” Every one of those backup paths was the correct call. A system that stops dead when one source hiccups is worse than one that degrades gracefully, and I’d build every one of them again.

The mistake was narrower. I let them succeed silently. The job knew it was running on the backup — it had to, to choose it — and it told nobody.

Make the degraded state say so

The fix turned out to be almost embarrassingly small, and it wasn’t another monitor.

A job should say which source it actually used, in its own output, where whatever reads the output can see it. A run on the backup path gets reported as a run on the backup path. The alert that was really a standing condition now says so, in a form the alerting knows to treat as a signal rather than a fault. The rest are getting the same treatment one at a time, and the question I ask of every dashboard now — including the one I check each morning — is where the number came from, not whether it arrived.

The general version is a rule I’d now apply to anything automated, at home or at work: a system that can quietly substitute one input for another has to announce the substitution. Otherwise you haven’t built resilience. You’ve built a very reliable way of not knowing.

Green is a claim. Ask it where it got its numbers.

■