Zero-install, deterministic repo-indexing engine — one vendorable
file, a library, a CLI, and 26 MCP tools.
Walks a repo, extracts symbols and imports across 15 languages,
resolves cross-file references into a typed link-graph, and
renders byte-identical artifacts every time it runs. Vendor the
single engine.mjs, install it from npm or Homebrew, or
run it from a container — nothing else to install, ever.
docker run --rm -v "$PWD":/work \
ghcr.io/maxgfr/codeindex scan --repo /work
Same zero-dependency engine.mjs bundle behind all
three — CLI, library import, and Docker image.
Key measured numbers
Fig. 01 — measured, reproducible
Benchmarks
Measured against universal-ctags,
Serena (LSP over MCP) and Graphify
with a reproducible harness
(scripts/bench/) — median of 5 runs, one warmup
discarded. Every number on this page comes from the same
site/benchmarks.json committed to this repo; nothing
here is rounded or typed in by hand. Full methodology and every
scenario:
BENCHMARKS.md.
Cold index — code-du-travail
2,823 files, every tool measured on the same repo. Lower is faster.
Warm reindex — against codeindex's own cold build
No competitor here exposes an incremental reindex, so the baseline is
the cold build of the same repo: what a rerun costs once the cache is
warm, and what one touched file costs. Lower is faster.
Axis by axis
One row per thing an agent-tooling team weighs. A filled tile marks the
leader on that axis — hover any cell for the measured number. Where a
rival wins it stays at the top of the table, unedited: ctags takes both
cold-index rows because a flat tags pass is a smaller job
than a typed cross-file graph. The rows below it are the ones that job
buys.
codeindex leads a competitor leads tie / not this tool's job— = n/a (not measurable)
Fig. 02 — capabilities
Features
Everything below ships in the same zero-dependency bundle — no
feature is an optional add-on package.
Keyless search + deterministic fuzzy
BM25 lexical ranking over symbol names, path segments and
headings. A term with zero document frequency falls back to a
byte-deterministic character-trigram match (Dice ≥ 0.6) — typo
tolerance with no embeddings. Disable with
--no-fuzzy.
Deterministic semantic tier
An opt-in static token → vector lookup table — no
neural forward pass. Encode → mean-pool → L2-normalize →
round-half-to-even int8 quantize; ranking is a pure integer dot
product, byte-identical everywhere. No model on disk, no
behavior change.
Rich embedding tier — v2.11.0
Point at a containerized HTTP endpoint
(ghcr.io/maxgfr/codeindex-embed, all-MiniLM-L6-v2)
via CODEINDEX_EMBED_ENDPOINT. Wins over the static
tier by explicit intent; degrades to lexical on exit 0 if
unreachable — never a silent fallback to the wrong model.
Call graph & caller index
Typed import / call / use
/ doc-link / mention edges at file and
module level. Per-symbol callers and references resolve in-proc
in sub-millisecond time on an already-loaded warm scan.
Graph analytics
Louvain communities, PageRank & betweenness centrality,
dead-code detection, complexity × churn hotspots, change
coupling, surprise-edge detection, and a tests→code map — all
derived from the same link-graph.
SCIP export
index.scip via a hand-rolled zero-dependency
protobuf encoder — validated against the official
scip CLI's stats/lint.
Multi-ecosystem workspaces
npm/yarn, pnpm, lerna, nx, cargo, go.work, maven, gradle, uv and
Composer — detected and named without running any of their own
tooling.
Symbolic edits over MCP
replace_symbol_body, insert_after_symbol, insert_before_symbol — precise edits addressed by
symbol name, not by line ranges that drift.
mermaid & repo map
Render a mermaid dependency diagram or a plain-text repo map
directly from the graph — no separate diagramming step.
Byte-deterministic outputs
Two cold builds of an unchanged repo produce identical
graph.json/symbols.json — confirmed on
7 of 7 measured repos this session. Graphify's graph.json
differed on 5 of the 6 measurable repos; determinism is a value we
hold across every repo, not just where it's easy.
Fig. 03 — usage
Quickstart
Five ways in — pick the one matching your integration.
Install it as a regular dependency and import the pure functions.
Multi-arch (linux/amd64, linux/arm64) —
nothing inside but node and the bundle, no
npm install.
docker run --rm -v "$PWD":/work ghcr.io/maxgfr/codeindex scan --repo /work
docker run --rm -v "$PWD":/work ghcr.io/maxgfr/codeindex index --repo /work --out /work/.codeindex
docker run -i --rm -v "$PWD":/work ghcr.io/maxgfr/codeindex mcp
Pin by digest in CI or anywhere reproducibility matters:
ghcr.io/maxgfr/codeindex@sha256:…. The rich
semantic-search tier ships separately as
ghcr.io/maxgfr/codeindex-embed.
Consumers who don't want an npm dependency vendor
scripts/engine.mjs + engine.d.mts at a
pinned release tag instead — the model downstream consumer
tools use, so each stays a standalone, single-file install.
The axis-by-axis scoreboard sits with the
benchmarks, where its numbers come from. This
is the shape of each tool: what it models, what it covers, what it
costs to install.
Tool comparison
codeindex vs. universal-ctags, Serena and Graphify.
Tool
Model
Language coverage
Install
Worth knowing
codeindex
Static byte-stable artifacts
(graph.json/symbols.json), rebuilt
each run with an incremental cache
A single-file incremental reindex is slower than a live
database (see Benchmarks)
universal-ctags
A flat tags file
Broad, generic parser rules
Binary install
Fastest cold index in this report — definitions only, no
call graph, no import resolution, no cross-file links
Serena
LSP client orchestrating per-language language servers
Any language with an LSP server
114 MB venv + per-language language servers
Type-aware references no static tool claims — bought with the
slowest activation and per-call latency measured here, and no
on-disk artifact to reproduce
Graphify
tree-sitter knowledge graph (graph.json);
code-only mode is keyless, the docs/PDF tier needs an LLM and is
excluded here
36 via tree-sitter
140 MB Python venv (uv tool)
~1 ms per lookup, but graph.json is not
byte-identical across rebuilds on 5 of the 6 measurable repos,
and file nodes are basename-keyed (no file-overview)