A code knowledge graph for agents. It indexes a repository into symbols — functions, methods, classes, their calls and callers — and answers a natural-language task with the handful of chunks that task actually needs, instead of a grep dump the model has to read.
pip install awgraph
| k | awgraph recall | tokens | grep recall | tokens | cheaper |
|---|---|---|---|---|---|
| 10 | 0.803 | 1,311 | 0.924 | 351,427 | 268× |
| 25 | 0.939 | 3,132 | 0.985 | 504,640 | 161× |
| 50 | 0.939 | 6,158 | 1.000 | 668,299 | 109× |
| 100 | 0.939 | 12,059 | 1.000 | 735,727 | 61× |
| 200 | 0.939 | 23,386 | 1.000 | 735,727 | 31× |
| 400 | 1.000 | 45,269 | 1.000 | 735,727 | 16× |
awgraph reaches the same ceiling as exhaustive grep — recall
1.000 — for 16× less context, and costs 16–268×
fewer tokens at every budget in between. k is swept for
both retrievers, because k bounds grep's output too;
sweeping it for only one arm manufactures a win, and an earlier version of
this page did exactly that.
Read the shape honestly. grep is the better finder
at any matched k — it hits 1.000 at k=50 while awgraph is
still at 0.939. The argument is not that grep is worse; it is that at full
recall grep needs 668k tokens per task, which does not fit in most context
windows at all. Note also that awgraph's tokens are previews
(signature + docstring + a body excerpt), so an agent that then reads its top
hits in full pays more than the figure shown; grep's number is whole files,
which is what an agent really has to read.
Two results recorded because they did not work. Embedding coverage was not the gap: going from 33.3% of chunks carrying vectors to 100% moved recall@10 from 0.800 to 0.803, refuting the earlier claim that partial coverage understated the number. And a naive graph-then-grep fusion scored 0.894 at 348,389 tokens — worse recall than grep and nearly its full cost, because a fallback keyed on result count fires when the graph is confidently wrong and stays quiet when it is confidently right.
Caveats, because a benchmark without them is marketing: n=33
real commits, one repository, Python only, and k is a budget a
caller chooses rather than something the tool tunes for itself. Every number
on this page is reproducible from
codegraph_task_context_ab.py --embed --sweep 10,25,50,100,200,400.
Ablated on the same 33 tasks and the same index, with the semantic half switched off:
| k | keyword only | with embeddings | gain |
|---|---|---|---|
| 10 | 0.682 | 0.803 | +0.121 |
| 25 | 0.818 | 0.939 | +0.121 |
| 50 | 0.909 | 0.939 | +0.030 |
Yes at small k, and less so as the budget grows — which is
the regime that matters, since the whole point is a small k.
Embedding on CPU is the slow part of setup; this is what it buys.
The task is a real commit message with the answer filenames stripped out of
it; the truth is the set of files that commit actually modified. Stripping
matters — leave the filenames in and the task becomes
grep <filename>, which the baseline wins by construction
while measuring nothing.
| step | 2,400 chunks | 43,730 chunks |
|---|---|---|
| parse + index | 49.8s | 75.5s |
| embed (CPU) | — | ~97 min |
Indexing is close to size-insensitive — 27× the files for 1.5× the time, because parsing runs across workers. Embedding is the part that hurts on CPU, so it is optional, cached and incremental. Without any embedding backend, queries fall back to keyword scoring and still work — that fallback is silent by design and dangerous by nature, so check vector coverage rather than assuming it.
pip install awgraph awgraph index . # parse + persist an index for this repo awgraph query "retry with exponential backoff" awgraph callers send_request # who calls this awgraph calls send_request # what does this call awgraph stats # what is in the index, incl. embedding coverage awgraph selftest # prove the install works
query prints path:line [type] name plus the
signature, so a result pastes straight into an editor; --json
on any read command gives machine-readable output. Exit codes distinguish
0 success, 1 a real negative answer, and
2 could not run — so a caller can tell “nothing
matched” from “there is no index yet”. The index is cached
outside your repository, keyed by a digest of its absolute path.
import asyncio
from awgraph import CodeGraph
async def main():
graph = CodeGraph(root_path="/abs/path/to/repo", auto_index=False)
await graph.index_codebase("/abs/path/to/repo") # absolute path required
for chunk in await graph.hybrid_query("retry with exponential backoff", max_results=5):
print(chunk.name, chunk.source_path, chunk.start_line)
asyncio.run(main())
The query does not need to contain the symbol name. Asking for “backoff policy for flaky calls” against a class documented as “Backoff policy for flaky calls” returns it by meaning, not by string match.
Three packages, three different questions about the same repository: awgit knows what changed and who is editing it; awgraph knows what the code is and what depends on what; awrelay knows who found what and who still needs to hear it. awdk is the agent runtime that consumes all three.
None of the three requires the others. Used together, an agent can find a symptom with awgraph, check whether it is an in-flight edit with awgit, and tell the agent already working that file with awrelay — three questions a solo grep-and-guess loop cannot ask at all.
The Aither World is an operating system for agents — a Linux you can hand to one, the runtimes it works in, and the tools it works with. awgraph is one of its 66 bricks — each installs on its own, runs offline, and needs no account. All 66 →