The question
You are about to change a function signature. Before you do, you want to know what breaks. The instinctive move is to grep for the function name. That gives you occurrences, not dependencies. It misses re-exports, aliased imports, dynamic imports, and every consumer that imports the module rather than the symbol.
This post is about the structure that actually answers the question: a source graph. Not a full compiler — a bounded, honest model of "this file references that file," good enough to scope a change and to tell you when it does not know.
Why grep is not enough
Consider a small TypeScript module with a barrel file:
// src/payments/index.ts
export { charge } from './charge.js';
export type { ChargeResult } from './types.js';
// src/api/checkout.ts
import { charge } from '../payments/index.js';A grep for charge finds both files and also every unrelated use of the word. It does not tell you that changing charge.ts breaks checkout.ts through the barrel. A source graph does, because it records edges between files after resolution — not between strings.
Dynamic imports make this worse. import(`./plugins/${name}.js`) is not a single edge; it is a pattern that may resolve to many files or none. The honest response is to record a dynamic boundary and say so, not to guess.
The pipeline
Building the graph is four bounded steps. Each step has a job and a limit.
1. Inventory
Walk the repository under the ignore rules — .gitignore, .npmignore, default noise. Produce a file list with paths and kinds (source, config, test, asset). Nothing is analyzed yet. This step answers "what exists," not "what it means."
2. Workspaces and project boundaries
A monorepo is not one project. package.json files, pnpm-workspace.yaml, go.work, Cargo.toml, and pyproject.toml mark ownership boundaries. Imports that cross a boundary are still edges, but they are labeled. A change that fans out across five packages is a different risk from one that stays inside a single package.
3. Resolution
For each source file, extract its references and resolve each one to a concrete file — or record why it could not be resolved. This is the step that differs per language. It is also the step where most tools silently give up.
4. Impact
Given a changed set, walk the graph to its reverse-reachable set. That set is the blast radius. Report it with the edges that got you there, so a human can check the walk.
Changed: src/payments/charge.ts
Blast radius (3 files, 2 boundaries):
src/payments/index.ts via export { charge } from './charge.js'
src/api/checkout.ts via import { charge } from '../payments/index.js'
packages/billing/src/charge.ts via cross-package: @app/payments
Unresolved: 1 dynamic boundary in src/plugins/loader.tsWhat each language gets wrong
Resolution is not one algorithm. Each ecosystem has a rule that will bite you if you assume the others.
- TypeScript / JavaScript: path aliases in tsconfig, extension-less imports, .js suffixes pointing at .ts files, barrel index files, and export * from chains. CJS require and ESM import coexist in the same repo.
- Python: implicit namespace packages, src/ layout versus flat layout, importlib.import_module with a computed name, and relative imports whose meaning changes with the package root.
- Go: module paths that do not match directory layout, internal/ visibility rules, and build tags that include or exclude a file per GOOS.
- Java: package names decoupled from directories just enough to be surprising, wildcard imports that pull in a namespace without naming a class, and annotation processors that generate sources the graph never sees.
- Rust: mod declarations that are not file paths in the way you expect, cfg-gated modules, and crate-relative use statements.
A tool that claims "multi-language" and uses one regex for all five is not lying about the regex. It is lying about the answer.
Missing-target detection
The inverse of impact analysis is often more valuable. Every resolved reference implies a file should exist. When it does not, you have found a bug the type-checker might also find — or might not, if the file is generated, vendored, or conditionally loaded.
Finding: import target missing (source-integrity/missing-target)
Location: src/api/checkout.ts:3:24
Reference: ../payments/charge.js
Resolved to: src/payments/charge.ts — not present in the inventory
Confidence: highThis is the check you want after a merge that deleted files, after a rename, and after a partial revert. It is also the check that catches a broken barrel: the barrel exports from a path that no longer exists, and nothing else in the repo mentions that path.
Stable fingerprints and baselines
A graph that produces findings is only useful in CI if it can tell new problems from old ones. That requires a fingerprint that survives line-number drift but changes when the problem changes.
A workable fingerprint is a hash over the rule identity and the structural location — module, exported symbol, reference text — not over the line number. Then you store a baseline of accepted findings, and CI fails only on findings whose fingerprint is not in the baseline.
# Accept today's known findings.
codebase-doctor audit . --json > baseline.json
# Later, after an external repair, verify.
codebase-doctor verify . --baseline baseline.json
# On every PR, fail only on new findings.
codebase-doctor audit . --changed --baseline baseline.json --fail-on highWithout this, teams turn the check off the first week because it fails on forty known issues. With it, the check stays on and actually catches regressions.
What the graph cannot see
The value of the tool is bounded by its honesty about limits. These are the ones that matter:
- Reflection and runtime wiring. A dependency-injection container that resolves a type by name is not a static edge. The graph should say "dynamic boundary" and stop.
- Generated sources. Code emitted at build time is not in the inventory unless you point the tool at the output directory. Migrations that generate ORM clients are a common blind spot.
- Cross-language edges. A TypeScript file that shells out to a Python script is not a source edge. Record the process boundary or accept the gap.
- Semantic impact. Knowing that A imports B does not mean a change to B is safe for A. A signature-preserving refactor and a behavior change look identical to the graph.
Any tool that reports a clean blast radius without listing its unresolved boundaries is reporting confidence it has not earned. Coverage has to be visible in the report — what was analyzed, what was skipped, and why.
Where this lives
This is the model behind codebase-doctor: a bounded inventory, per-language resolution, a reverse-reachable impact set, missing-target detection, and fingerprints that make the output usable in CI. The report carries its own coverage and limitations next to the findings.
If you want the same result without the tool, you can build a graph for one language in an afternoon and be wrong in exactly the ways listed above. The tool exists because the wrong answers are expensive and quiet.
codebase-doctor is open source, runs offline by default, and never writes to the repository it audits.