Seventeen alarms nobody had heard
One of them could not have fired in any state the world could be in. Finding that out took asking a question the green dashboard could not answer.
Seventeen alerting rules watch this project: backups that have stopped succeeding, a zone falling below a playable tick rate, a log shipper that quietly stopped, a disk filling. They had been evaluating for days and every one of them was quiet, which is what you want. It is also what a rule that cannot fire looks like.
Two of them were eventually forced to fire by hand, by planting a stale file on the live host and watching the alert go pending, then firing, then clear. That is worth doing once. It is not worth doing seventeen times: it takes a person, it touches production, and it proves the rule as it was that afternoon rather than as it is after the next edit.
The second one is the reason this post exists. It compares each zone's content version against the highest any zone is serving, and it had been written the obvious way: take the maximum, subtract, alert if the difference is positive. That expression matches nothing. Not when the zones disagree, not when they agree - nothing, in every possible state of the world, because taking a maximum discards the labels that say which zone a number came from, and there is then nothing left to match the two sides on.
So it returned no alerts. And no alerts was the correct answer that day, because the zones were in fact all on the same version. A panel fed by the same expression read zero for a day and was right to. There is no observation you could have made of that rule, short of breaking something on purpose, that would have told you it was broken.
The fix is not cleverer monitoring. It is Prometheus' own test framework, which feeds invented measurements to the real rule evaluator and asserts which alerts come out. All seventeen rules are under test now, and every one is asserted twice: once with data that must make it fire, and once with data that must leave it silent. One direction alone is worthless. A test that only checks firing cannot tell a working rule from one that fires always; a test that only checks silence cannot tell one that fires never, which was exactly the bug.
Then the suite was pointed at itself. The old broken expression went back in, on purpose, to see whether the tests would notice - and the case went red naming the zone it should have caught. A test suite nobody has seen fail is decoration, and this project has been bitten by decoration often enough to stop trusting green on sight.
The same shape turned up twice more the same day, which is why it is worth writing down rather than filing as one bug. A test in the authentication service spent ten seconds failing on a developer machine for the wrong reason: it needed a database, was not marked as needing one, and so neither skipped cleanly nor failed honestly. And a piece of infrastructure was configured to narrow something and, because of how the network underneath it behaved, narrowed nothing at all - while reading, to anybody who looked, exactly like a working one.
All three are the same defect wearing different clothes: a check that cannot distinguish the healthy state from the broken one. It is a more dangerous class of bug than the ordinary kind, because the ordinary kind eventually shows itself. This kind produces exactly the reassuring output it would produce if everything were fine, and it goes on producing it for as long as you leave it alone.
The rule now, written down where the next person will find it: before trusting any check, make it fail on purpose. If you cannot make it fail, it is not a check.
