Skip to content
awprism — /readme
◆ Aither World awprism

awprism

Turn a failure into ranked hypotheses — and say what would confirm each one.

What it does

Instead of trusting

the first explanation that fits

You check

the ranked alternatives, and the observation that separates them

Overview

awprism/README
Debugging with an agent collapses onto the first plausible story, because nothing forces a second. The cost is not the wrong guess, it is the hours spent proving it — measured repeatedly here, where five hypotheses were spent on a service that was genuinely correct, and where a symptom named the wrong component so consistently that "the symptom names the INNOCENT service" had to be written down as a standing rule.

awprism

Docs · Source · pip install awprism · The Aither World

The Aither World is an operating system for agents — a Linux you can hand to one, the runtimes it works in, and the tools it works with. awnix is the Linux underneath it; awprism is one of its 67 bricks — each installs on its own, runs offline, and needs no account.

Start here: Hand it a failure and get back ranked candidate causes, each with the one observation that would rule it in or out.

Turn a failure into ranked hypotheses — and say what would confirm each one.

```bash
pip install awprism
```python
from awprism import Prism

prism = Prism()
diagnosis = prism.diagnose(
    "the database connection timed out",
    context="this happens only on the replica, not primary",
)

for hyp in diagnosis:
    print(f"• {hyp.claim}")
    print(f"  Test: {hyp.falsifier}")
    print()
```bash
awprism diagnose "the API returns 500"
awprism diagnose "service is slow" --context config.txt --markdown
awprism diagnose "network timeout" -k 3 --json
awprism health
awprism --self-test

What this is

When debugging with an agent, the conversation collapses onto the first plausible story. That's not wrong, but it's systematically biased: the loudest error in the logs often names an innocent service, so the first plausible explanation leads away from the root cause.

Prism breaks the first-story bias by generating multiple ranked hypotheses for every failure, each with a falsifier — the single observation that would confirm or rule it out. Instead of picking a theory and defending it, you now know what to check first.

Prism ranks candidates by the evidence at hand, degrades gracefully without an LLM, and works offline.

thing what it does
Prism.diagnose() Takes a symptom + context, returns ranked hypotheses
Hypothesis A candidate cause with a claim, score, and the test that would separate it
Diagnosis A ranked list of hypotheses with .to_dict() / .to_markdown()
StrategyRegistry Reusable diagnostic patterns for timeout, auth, resource, etc.

The bug this package exists to prevent

Debugging produces a narrative. A narrative is a story that sounds good, which is the exact opposite of what you want. The story you pick is the one you'll defend, which is the story you won't question.

Prism forces the question: What if I'm wrong about this? It doesn't answer for you — it just hands you the test that would tell.


Core API

Prism(complete=None, registry=None)

Create a diagnostic engine.

Arguments: - complete — Optional LLM completion callable: (prompt: str) -> str. If omitted, Prism operates in heuristic mode using structural patterns. - registry — Optional StrategyRegistry. If omitted, a default registry with built-in strategies is created.

diagnosis = prism.diagnose(symptom, context="", k=5)

Generate ranked hypotheses for a failure.

Arguments: - symptom — Description of the failure (required, non-empty). - context — Additional background (optional). - k — Maximum number of hypotheses to return (default 5).

Returns: - Diagnosis object with .hypotheses (sorted by score) and .to_dict() / .to_markdown().

Raises: - ValueError if symptom is empty.

Hypothesis

```python
Hypothesis(
    claim="the service is down",
    score=0.7,                                     # 0.0 to 1.0
    falsifier="check the service status endpoint", # the one test
    evidence_for=["timeout occurred"],             # supporting observations
    evidence_against=[],                           # contradicting observations
    rationale="timeouts are common when services are overloaded",
)
  • claim — The hypothesis statement (required).
  • score — Confidence from 0.0 (ruled out) to 1.0 (certain). Always in [0, 1].
  • falsifier — The single observation that would confirm or rule this out (required, non-empty).
  • evidence_for — List of supporting observations (default: []).
  • evidence_against — List of contradicting observations (default: []).
  • rationale — Explanation of the score (default: "").

A hypothesis without a falsifier is an opinion. The package enforces non-empty falsifiers.

Diagnosis

```python
diagnosis.hypotheses       # list of Hypothesis, sorted by score (descending)
diagnosis.symptom          # the original symptom string
diagnosis.context          # the context, if provided

len(diagnosis)             # number of hypotheses
for h in diagnosis:        # iterate hypotheses
    ...
diagnosis[0]               # first hypothesis (highest score)

diagnosis.to_dict()        # serialize to dict (for JSON)
diagnosis.to_markdown()    # render as markdown (for human reading)

Degraded mode (no LLM)

Prism generates hypotheses even without a completion backend, using a registry of reusable diagnostic patterns:

  • timeout — service overload, network throughput, crash+restart
  • auth — wrong credentials, missing permissions, auth system down
  • resource — disk full, out of memory, quota hit, memory leak
  • unknown — fallback for anything else

Each pattern comes with ranked hypotheses backed by structural reasoning. You can add your own with registry.register(DiagnosticStrategy(...)).


CLI

awprism diagnose SYMPTOM [OPTIONS]

Analyze a failure.

Options: - --context FILE — Path to a file with additional context. - -k N — Max hypotheses to return (default 5). - --json — Output as JSON. - --markdown — Output as Markdown.

Examples:

```bash
awprism diagnose "API returns 500"
awprism diagnose "database timed out" --context logs.txt --markdown
awprism diagnose "network unreachable" -k 3 --json

awprism health

Run the self-test to verify Prism is working.

awprism --self-test

Prove Prism holds its core contracts offline (no network, no LLM).

awprism mcp -- for a coding agent

Serves diagnose to any MCP client over stdio. Stdlib only, starts in under a second, read-only.

```json
{"mcpServers": {"awprism": {"command": "awprism", "args": ["mcp"]}}}
tool does
diagnose(failure_text, context_paths?, k?) ranked causes, each with confirm_by: the one observation that would confirm or rule it out
strategies() which failure classes are recognised

Python tracebacks and pytest output are parsed: exception class, message and innermost file:line. A chained traceback (raise ... from) is diagnosed at its FIRST exception, not the wrapper printed last. context_paths are read (tail 64 KiB each, max 8 files), never written; a path that could not be read is listed under context_skipped.


What this does NOT do

It does not execute checks. Prism tells you what to measure, not how to measure it or whether the measurement passes. You decide.

It does not fix anything. It diagnoses and ranks; fixing is your call.

It does not replace a monitoring system. Alerts tell you something broke; Prism tells you what might have broken it.

It does not substitute for domain knowledge. Prism works with what you tell it. If you know the system well, use that — Prism augments it, not replaces it.


Adding diagnostic strategies

A DiagnosticStrategy is a pattern that applies to a class of failures and generates relevant hypotheses.

```python
from awprism import Prism, DiagnosticStrategy, Hypothesis

def my_pattern_applies(symptom: str) -> bool:
    return "crash" in symptom.lower()

def my_hypotheses(symptom: str, context: str) -> list[Hypothesis]:
    return [
        Hypothesis(
            claim="the process ran out of memory",
            score=0.8,
            falsifier="check memory usage at the crash time",
        ),
        Hypothesis(
            claim="a signal was sent to terminate the process",
            score=0.5,
            falsifier="check system logs for SIGKILL or administrative actions",
        ),
    ]

strategy = DiagnosticStrategy(
    name="crash",
    description="Diagnoses process crashes",
    pattern_check=my_pattern_applies,
    generate_hypotheses=my_hypotheses,
)

prism = Prism()
prism.registry.register(strategy)

diagnosis = prism.diagnose("the worker process crashed")

--self-test

Every install can prove Prism still holds its contracts, offline:

```console
$ awprism --self-test
  PASS  Hypothesis requires claim, score, and falsifier
  PASS  Diagnosis sorts by score; Prism generates >= 2 hypotheses
  PASS  Every hypothesis has a non-empty falsifier
  PASS  Registry includes timeout, auth, resource, and unknown strategies
SELF-TEST: awprism ok

The key assertions: - A Hypothesis without a falsifier is rejected. - Prism.diagnose() always generates at least 2 hypotheses (or none). - Every hypothesis has a non-empty, non-whitespace falsifier. - The StrategyRegistry includes built-in patterns for common failure classes.


The aw family

Standalone tools that share one idea: replace something you would otherwise have to trust with something you can check.

Each installs on its own, works offline, and needs no account.

instead of trusting you check
awdk a framework's idea of how your agents should run one loop you can read, pointed at a backend you already pay for
awskills that an agent knows your procedure the procedure written down, versioned, and loadable by any agent
awpack that the pack you want shipped inside somebody's SDK, under whatever licence that SDK happens to carry the pack as its own versioned artifact, with its own licence, that any agent runtime can install
awm that memory stayed in its lane tenant:user:project scopes, so a write cannot cross a boundary
awdesk that the agent is somewhere behind a browser tab a tray icon, a face on your desktop, and the decision card that pops when it needs you
awnode a vendor's cloud with every prompt a local gateway routing to backends you chose
awgraph that grep found everything an AST + tree-sitter call graph an agent can traverse
awgit that no one else is editing this file a lease, refused at commit time if you do not hold it
awdelphi one agent's confident take on a decision the round trace, the anonymity, and who dissents
awclassify a filename, a folder, or whoever last touched it doc_type, visibility, audience and topics, with the evidence lines that decided each
awdecide a hosted classifier's probability that never learns whether it was right the decision, its probability, and the calibration curve from your own resolved outcomes
awtoll that your tooling is saving you context the measured token cost of each tool call, and what the alternative cost
awseal that the artifact came from who you think an Ed25519 seal — the key that verifies is not the key that forges
awshare that the download is intact content-addressed bundles, verified on fetch
awsuite that an agent holding your mailbox will not send on its own every send, draft, upload and create returns a dry-run until confirm is true
awnest that there is a person on the other end a verdict with evidence, where "we could not tell" is not "yes"
awrena a leaderboard someone can edit, and votes nobody counted a scored duel with both answers kept, and a result bound to them
awnboard a share link anyone who sees it can use an invitation addressed to one person, for one gate, revocable
awnix that the box is what you left it as an immutable image you built, with atomic rollback
awrecover that the restore worked a restore that fully lands or does not land at all
awstorage a du you ran last month, and a peers file that says 3 TB free an inventory snapshot per node with a diff since the last one, and each tree classified re-fetchable or not
awrelay a SaaS in the middle of your agents findings, alerts and coordination over your own transport
awask that anyone read the paragraph where you asked the ask itself, with a button that steers the session that raised it
awmail a mailbox somebody else can read mail your agents send and receive over your own server
awswarm that a model either fits your GPU or it doesn't run at all a placement plan and an acquisition-probability estimate before you spend on a run
awfind one vendor's idea of the web results from whichever providers you configured
awbrowse that the page said what you were told the render, the DOM and the requests it made
awvoice that a cloud vendor may hold your audio a transcript and a wav from a service you host
awvision a filename and a caption somebody wrote what a model actually reports about the pixels
awscreen a selector that was true when the page was written the elements actually rendered, by what they look like
awbeads that a layout your users built survives the next deploy the arrangement as data you can read back, diff, and hand to another surface
awbonsai that inference always means a request left the machine a WebGPU model answering on the tab's own GPU, with a consent record logged before it ever loaded
gawbbonet the model to keep a 300-message campaign coherent by itself campaign facts recalled from scoped memory you can list and edit
aitherkvcache a vendor's quantisation defaults sub-byte KV cache kernels you can benchmark yourself
awrtifact a hand-rolled split script and a hand-edited worker manifest byte-verified parts in a release, served with Range + CORS, sizes asserted by a live gate
AitherZero a pile of scripts nobody has numbered numbered, discoverable automation with declarative playbooks
AitherConnect what a page tells your browser to do a federated search and desktop bridge you host
awreason a confident paragraph the phases it went through, and every tool call it made to get there
awrecurse that everything you pasted in was actually read which slices it opened, and what it concluded from each
awprism (you are here) the first explanation that fits the ranked alternatives, and the observation that separates them
awrepl what the agent believes the value is the value, printed from the live session
awreport that the report you pasted carried no token in it a redacted report, and the duplicate it merged into instead of filing twice
awresearch a summary of pages nobody opened every claim against the source it came from
awfocus twelve terminal tabs and a bad memory one command that names every session, finds any transcript, and opens or steers the one you want
awgym that a world model learned anything from the games it saw transitions captured from real play, fed back, and the retrodiction score falling on grids it never saw
awpredict a model because it trained without erroring its prediction against a self-updating lookup, on the rows that are actually novel
awevolve that your optimisation loop is finding anything every version it kept, the score that version earned, and the edit that produced it
awsh that you already know the name of the command what it decided your line meant, before it acts on it
awmine that a session's lesson survived the session a row per outcome, a candidate per lesson, and the transcript line each one came from
awrise that a scheduled agent ran at all, and ran exactly once a durable record of every wake -- fired, skipped, overlapped or timed out -- each with its reason
awkno that the docs site is up, or that you remember the family the whole ecosystem in your terminal, with no network at all
awwall that a service only talks to the hosts you think it talks to an explicit egress allowlist, where a denial names the rule that denied it
awembed a general-purpose embedder that has never seen your code a held-out split of whole directories, scored teacher vs student vs int8
awtax a closed tax app's sealed file you can never read again a plain, provider-neutral schema of every figure, with the page it came from
awsettings that you will remember to re-approve the same thing on every box you work from one profile, unioned rather than overwritten, with the credentials left behind
awavatar a cloud 3D vendor's opaque task id a manifest with a sha256, a licence and a rig-audit verdict per file

awnix is the ground floor — A Linux you can hand to an agent — immutable base, capabilities included.

The Aitherium ecosystem

Every repository here is public. Each publishes an aither-manifest.json beside its page, so any surface can read every sibling's — the network is browsable from any node in it.

repo what it is pages
awdk Build AI agent fleets — 3 lines, any backend, local or cloud docs
awskills Portable agent skills — self-contained procedures an agent loads on demand docs
awpack First-party agent packs — the ones we build, versioned and installable on their own docs
awm A portable, scoped agent memory docs
awdesk Aither World Desk -- the desktop body of AitherOS Online: tray, avatars, decision cards, the Living Desktop as an overlay docs
awnode A lightweight local gateway — bridges your apps to the AI backends you chose docs
awrun A priority-aware queue and dispatcher for agentic runs and ad-hoc CI builds. It also judges whether the runner pool is big enough for the queue it is draining, and can ask a host to grow it -- reserving capacity is zero-sum, so a saturated pool needs more of it, not a different share of it docs
awgraph A semantic code graph for agents — AST + tree-sitter, call graphs docs
awgit Semantic version control on top of git — edit-ops and leases docs
awdelphi Anonymous multi-round expert panels — a converged answer with a trace docs
awclassify Classify any document -- what it is, who may read it, who it is for, what it is about —
awdecide One typed-decision contract -- choice / score / bool with a probability -- over a ladder of backends you already run (rules, tiny local models, an LLM's logprobs), fail-closed, with a Brier ledger that resolves every decision against its outcome —
awtoll What every tool call costs you in context, measured from your own transcripts docs
awseal Sign an artifact so a stranger can verify it docs
awshare Publish an artifact and fetch it back verified docs
awsuite Your Google Workspace as agent tools, and no write happens without a yes —
awdit An append-only audit trail whose gaps are DETECTABLE docs
awbac Role-based access control that fails closed and explains itself docs
awiam Who is this caller? A directory and session store that fails honestly docs
awtunnel Reach a service that has no public address docs
awnest Prove there is a human before you let them into the nest docs
awrena Put two agents head to head and get a verdict you can check docs
awnboard A front gate you can put in front of anything, and hand someone the key to docs
awnix A Linux you can hand to an agent — immutable base, capabilities included docs
awrecover Labelled snapshots with an all-or-nothing restore docs
awstorage Every drive on every node, indexed, classified and diffed -- so you can see what you own before you delete it docs
awrelay Portable agent messaging — findings, alerts, coordination docs
awask Your agent asks you a question — and acts on your answer docs
awmail Give an agent an email address — send, and actually receive docs
awnet The agentic web — agents host a mesh, and agents join one docs
awswarm Run one model too big for any single GPU across a pool of small ones —
awfind A portable search client — query, results, ranking docs
awbrowse A portable browser client — navigate, console, network, DOM, screenshot docs
awvoice Hear and speak — transcribe audio, synthesize a voice docs
awvision See an image — describe it, ask it a question, compare two docs
awscreen See this machine — what is on screen, and where to click it docs
awkit Render an agent panel from a tool result — one component, any React app —
awbeads A spatial canvas for a page — arrange things, connect them, and keep the arrangement —
awbonsai Run a real model in the visitor's own browser — no server round trip, no upload —
awknowledge How to run a coding agent so the result survives — the laws, with evidence docs
awbrain Your history as a wiki of linked markdown — claims pinned to the evidence —
gawbbonet GobboNet campaigns with a real agent brain — scoped memory, graph recall docs
aitherkvcache Near-optimal KV cache quantization for LLM inference — sub-byte compression docs
awrtifact Deliberately chunk artifacts into GitHub release assets — the productized aitherkvcache mirror lane docs
AitherZero PowerShell 7+ automation framework — numbered, self-describing scripts docs
AitherConnect Browser extension — federated AI search, page context, and the Living OS overlay docs
awreason A portable reasoning client — sessions, phases, thoughts, and the chain that produced the answer docs
awrecurse Answer a question over a context far larger than the window — recursively, with the trace kept docs
awprism (you are here) Turn a failure into ranked hypotheses — and say what would confirm each one docs
awrepl A REPL an agent can actually use — state that survives between turns docs
awreport File a bug report that has already scrubbed your secrets and collapsed the duplicate —
awresearch Ask a research question, get a cited report you can check docs
awfocus See, search and steer every Claude session from one command docs
awgym An ARC training gym — a game a world model can watch, and six roles that play through it docs
awpredict Predict what your environment does next, and how surprised you were docs
awevolve Point an agent at a file and a command that scores it, and let it improve —
awsh Your terminal answers you -- type a question where a command would go docs
awmine Mine what your agents did -- outcomes, lessons and procedures out of the transcripts they left behind —
awrise Wake an agent on a schedule, let it do one thing, and put it back to sleep docs
awkno The man page for the Aither World — every brick, stack and law, offline docs
awwall Say what a workload may reach, and watch everything else fail closed docs
awrouter OpenRouter for your own fleet: pick a model backend by cost/latency/ capability, fail over, fit the context window, stream. Standalone, OpenAI-compatible, no Aither-specifics required to be valuable —
awembed Train an embedding model that knows your corpus, and prove it beats the big one docs
awtax Turn any tax PDF -- returns, W-2, 1099, statements, even scans -- into structured data you can check docs
awflow A deterministic workflow runtime — chain agent calls with journal replay and budget control docs
awsettings Your agent's permissions and config, following you to the next machine docs
awavatar One character spec in, a rigged, animated, multi-style avatar pack out docs

Built on llama.cpp · vLLM · ComfyUI · CentOS Stream · Podman · Docker · LanceDB · WireGuard · FFmpeg · Blender + Rigify · headroom · SANA · Hunyuan3D · repowise · Playwright · Chromium · Next.js · React.

License

Apache 2.0. See LICENSE file.

The Aitherium Ecosystem

Portable tools you adopt one at a time. Each one works alone.

AitherConnectAitherConnectBrowser extension — federated AI search, page context, and the Living OS overlay. AitherZeroAitherZeroPowerShell 7+ automation framework — numbered, self-describing scripts. aitherkvcacheaitherkvcacheNear-optimal KV cache quantization for LLM inference — sub-byte compression. awaskawaskYour agent asks you a question — and acts on your answer. awavatarawavatarOne character spec in, a rigged, animated, multi-style avatar pack out. awbacawbacRole-based access control that fails closed and explains itself. awbrowseawbrowseA portable browser client — navigate, console, network, DOM, screenshot. awdelphiawdelphiAnonymous multi-round expert panels — a converged answer with a trace. awdeskawdeskAither World Desk -- the desktop body of AitherOS Online: tray, avatars, decision cards, the Living Desktop as an overlay. awditawditAn append-only audit trail whose gaps are DETECTABLE. awdkawdkBuild AI agent fleets — 3 lines, any backend, local or cloud. awembedawembedTrain an embedding model that knows your corpus, and prove it beats the big one. awfindawfindA portable search client — query, results, ranking. awflowawflowA deterministic workflow runtime — chain agent calls with journal replay and budget control. awfocusawfocusSee, search and steer every Claude session from one command. awgitawgitSemantic version control on top of git — edit-ops and leases. awgraphawgraphA semantic code graph for agents — AST + tree-sitter, call graphs. awgymawgymAn ARC training gym — a game a world model can watch, and six roles that play through it. awiamawiamWho is this caller? A directory and session store that fails honestly. awknoawknoThe man page for the Aither World — every brick, stack and law, offline. awknowledgeawknowledgeHow to run a coding agent so the result survives — the laws, with evidence. awmawmA portable, scoped agent memory. awmailawmailGive an agent an email address — send, and actually receive. awnboardawnboardA front gate you can put in front of anything, and hand someone the key to. awnestawnestProve there is a human before you let them into the nest. awnetawnetThe agentic web — agents host a mesh, and agents join one. awnixawnixA Linux you can hand to an agent — immutable base, capabilities included. awnodeawnodeA lightweight local gateway — bridges your apps to the AI backends you chose. awpackawpackFirst-party agent packs — the ones we build, versioned and installable on their own. awpredictawpredictPredict what your environment does next, and how surprised you were. awprismawprismTurn a failure into ranked hypotheses — and say what would confirm each one. awreasonawreasonA portable reasoning client — sessions, phases, thoughts, and the chain that produced the answer. awrecoverawrecoverLabelled snapshots with an all-or-nothing restore. awrecurseawrecurseAnswer a question over a context far larger than the window — recursively, with the trace kept. awrelayawrelayPortable agent messaging — findings, alerts, coordination. awrenaawrenaPut two agents head to head and get a verdict you can check. awreplawreplA REPL an agent can actually use — state that survives between turns. awresearchawresearchAsk a research question, get a cited report you can check. awriseawriseWake an agent on a schedule, let it do one thing, and put it back to sleep. awrtifactawrtifactDeliberately chunk artifacts into GitHub release assets — the productized aitherkvcache mirror lane. awrunawrunA priority-aware queue and dispatcher for agentic runs and ad-hoc CI builds. It also judges whether the runner pool is big enough for the queue it is draining, and can ask a host to grow it -- reserving capacity is zero-sum, so a saturated pool needs more of it, not a different share of it. awscreenawscreenSee this machine — what is on screen, and where to click it. awsealawsealSign an artifact so a stranger can verify it. awsettingsawsettingsYour agent's permissions and config, following you to the next machine. awshawshYour terminal answers you -- type a question where a command would go. awshareawsharePublish an artifact and fetch it back verified. awskillsawskillsPortable agent skills — self-contained procedures an agent loads on demand. awstorageawstorageEvery drive on every node, indexed, classified and diffed -- so you can see what you own before you delete it. awtaxawtaxTurn any tax PDF -- returns, W-2, 1099, statements, even scans -- into structured data you can check. awtollawtollWhat every tool call costs you in context, measured from your own transcripts. awtunnelawtunnelReach a service that has no public address. awvisionawvisionSee an image — describe it, ask it a question, compare two. awvoiceawvoiceHear and speak — transcribe audio, synthesize a voice. awwallawwallSay what a workload may reach, and watch everything else fail closed. gawbbonetgawbbonetGobboNet campaigns with a real agent brain — scoped memory, graph recall.