Turn a failure into ranked hypotheses — and say what would confirm each one.
Debugging with an agent collapses onto the first plausible story, because nothing forces a second. The cost is not the wrong guess, it is the hours spent proving it — measured repeatedly here, where five hypotheses were spent on a service that was genuinely correct, and where a symptom named the wrong component so consistently that "the symptom names the INNOCENT service" had to be written down as a standing rule.
pip install awprism
Hand it a failure and get back ranked candidate causes, each with the one observation that would rule it in or out.
Pairing is composition, never dependency — awprism installs and runs on its own.
Every brick below is live, drawn from that repository's own published manifest.