awgraph is a code intelligence engine: it reads a repository, extracts symbols (functions, methods, classes, modules), embeds them into a vector space, and serves a hybrid keyword+semantic retrieval system. An agent can ask “what functions call this symbol?” or “show me the impact of changing this module” without grepping the whole codebase.
The index stores ~1.5k tokens of context per symbol (function/method/class body and immediate context), making retrieval precise enough for code change analysis without blowing the agent’s context window.
awgraph does not exist alone. It is one of three complementary packages that together form an agent’s code understanding layer:
Purpose: Know what changed and who is editing it.
awgit rewrites git’s model to be machine-readable:
insert, delete, modify, rename, not raw diffsExample: You’re adding a new permission check. awgit tells you:
Purpose: Know what the code is and what depends on what.
awgraph indexes:
Example: Same permission check. awgraph tells you:
Purpose: Use awgit and awgraph together to make decisions.
The adk is an agent framework that:
Scenario: An agent is asked to add a new permission level to an auth system and update all code paths that call the auth check.
Agent: "I need to add permission level 'sponsor' to the auth module"
awgit response:
- Module last touched 4 days ago by agent-ci (lease expired, can edit)
- 12 commits touch it this quarter
- No concurrent locks
awgraph response:
- auth.check_permission is called 47 times across the codebase
- 31 calls are in service routes (production paths)
- 16 calls are in test code
- 8 of the 31 are in admin-only code (may not need update)
- 3 are in legacy deprecated endpoints (skip these)
- 20 need explicit review before updating (new calls, edge cases)
Agent stages in order:
1. Commit 1: Add "sponsor" to the enum in auth module (small, isolated)
2. Commit 2: Update 20 call sites to handle the new level
3. Commit 3: Add tests for the new permission level
awgit:
- Stacks all 3 commits into one PR (logically coherent)
- Each commit is semantic (no merge conflicts in the oplog)
- Lease prevents concurrent edits to auth.py
Agent queries awgraph:
- "What else calls the 20 functions I just modified?"
- Result: 3 high-level router functions, each already tested
- Conclusion: No secondary impact, PR is safe to merge
Grep is simple and inclusive: search for the function name, get every occurrence. But it has problems:
Semantic search (embeddings) finds meaning, not syntax:
Measured performance on real commits (15 commits, scope ~2400 code chunks, 800 embedded at 33.3% coverage):
| Method | Recall@10 | Avg Context |
|---|---|---|
| grep keyword search | 0.933 | 350k tokens |
| graph semantic+keyword | 0.867 | ~1.5k tokens |
| grep + graph (pick best per query) | 1.000 | ~50k tokens |
Why grep alone still wins on pure recall: The graph indexes only function/method/class bodies. Module-level code, comments, and string literals are invisible to it. On a task where the answer lives in a string literal or a module-level constant, grep finds it and the graph does not.
Why combining them wins: They are complementary.
The caveat: 33.3% embedding coverage means the graph only indexes one-third of the codebase’s symbols. The remaining two-thirds fall back to grep. This is by design — embedding is expensive. Real-world usage tunes the coverage threshold based on latency/cost trade-offs.
pip install awgraph
from awgraph import CodeGraph
# Index a repository
graph = CodeGraph.from_directory("./my-repo")
graph.save("./my-repo.graph")
# Load an index
graph = CodeGraph.load("./my-repo.graph")
# Find who calls a function
callers = graph.callers("auth.check_permission")
for caller in callers:
print(f"{caller.name} at {caller.file}:{caller.line}")
# Find impact of a change
impact = graph.impact_analysis("auth.check_permission")
for affected in impact.affected_symbols:
print(f" {affected.name} will break if signature changes")
# Semantic search
results = graph.hybrid_query(
"permission check in auth system",
k=10
)
for result in results:
print(f" {result.name}: {result.score:.2f}")
from awgraph import CodeGraph
from your_agent_framework import Agent
# Index the repo once
graph = CodeGraph.from_directory("./repo")
graph.save("./repo.graph")
# In your agent loop
class CodeUnderstandingAgent(Agent):
def __init__(self):
self.graph = CodeGraph.load("./repo.graph")
def plan_change(self, task: str) -> list[str]:
"""Return files to read for this task."""
# Semantic search for relevant code
search_results = self.graph.hybrid_query(task, k=20)
# Get impact of changes
impact = self.graph.impact_analysis(search_results[0].name)
# Collect files: search results + impacted code
files = set()
for result in search_results:
files.add(result.file)
for symbol in impact.affected_symbols:
files.add(symbol.file)
return sorted(files)
awdk already carries the integration point — agent.set_code_graph(cg)
registers code_search and code_context as agent tools. awgraph satisfies the
interface it expects (query, get_context_for_chunk, get_full_body) with no
adapter, so this is the whole wiring:
from awgraph import CodeGraph
cg = CodeGraph(root_path="/abs/path/to/repo", auto_index=False)
await cg.index_codebase("/abs/path/to/repo")
agent.set_code_graph(cg) # registers code_search + code_context
Measured end to end with no monorepo on sys.path — the standalone package
only:
TOOLS_REGISTERED 2 ['code_context', 'code_search']
code_search("backoff policy for flaky calls")
-> {"results": [{"name": "RetryPolicy", "type": "class", ...}]}
Note the query contains none of the words in the class name. It matched on the docstring.
These are documented designs, not shipped code — stated plainly so nobody reads them as available:
RetryPolicy.next_delay rather than that
17 lines moved. Feed that symbol to awgraph’s impact query and you get the
blast radius of the change instead of a diff. The join key is the symbol name;
nothing else needs to agree.GitNexus (source-available, noncommercial license) is the closest alternative. It also builds a code graph and serves retrieval. GitNexus’s recommended mode is native_augment — grepping code, then enriching results with graph context. That is exactly what our benchmark found to be optimal: grep for recall, graph for precision and context reduction. awgraph achieves the same conclusion with a smaller, more focused API surface.
The benchmarks measure different things:
The two are related but not directly comparable. A 100% recall retriever may still fail a task if the agent does not know what to do with the results.
awgraph consists of:
The index is read-only after build (no incremental updates). For code that changes frequently, rebuild the index on a schedule or during CI.
Scope: Only indexes function/method/class definitions. Module-level code, config files, and comments are not indexed. Use grep for those.
Language support: Python first; Go, Rust, JavaScript/TypeScript via tree-sitter. Language coverage depends on parser quality.
Embedding cost: Building an index for a 1M-line codebase takes time and disk. Pre-built indexes for popular open-source projects may be available.
Dynamic dispatch: The graph is static. It cannot resolve runtime polymorphism or string-based imports. For dynamic code, augment with grep or runtime analysis.
Vector quality: Embeddings are useful for semantic search, but they are not perfect. Grep is still better for specific literal matches.
See the main README for the full API.
CodeGraph.from_directory(path) — Index a directorygraph.hybrid_query(text, k=10) — Semantic + keyword searchgraph.callers(symbol_name) — Find all callers of a symbolgraph.callees(symbol_name) — Find all functions a symbol callsgraph.impact_analysis(symbol_name) — Estimate blast radius of a changegraph.save(path), graph.load(path) — Persist/restore indexawgraph is maintained as part of the Aitherium platform. Issues and pull requests are welcome.
MIT. See LICENSE in the package root.