Guide · Architecture Diagrams

How to Draw an Architecture Diagram (Worked Example: Ollama)

Six steps from a repo you've never opened to a diagram you can defend, done by hand on one of 2026's most-installed AI tools.

11 min read

Ollama's repository has 1,320 files in it. A useful architecture diagram of Ollama has thirteen boxes.

So drawing it isn't really a drawing problem. The job is deciding which 1,307 files don't get a box, and being right about it. Pick a diagramming tool first and you'll get a beautiful picture of the wrong thing.

Here's the diagram nearly everyone draws first:

flowchart LR A[CLI] --> B[Server] --> C[Model]

Accurate, and useless. It also describes every local LLM tool shipped this year.

Nothing on it is false, which is exactly the problem. It can't be wrong because it doesn't say anything, and a diagram that can't be wrong can't help anyone either.

The fix is a better order of operations. What follows is the sequence Revibe runs when it generates an architecture diagram from a repo: classify the project, filter the files, group them into capabilities, trace the flows, then draw. Here it's done by hand on Ollama, so you can repeat it on your own codebase with a terminal and a text editor.


1. Figure out what kind of thing it is

Before any boxes, answer a blunter question: what is this? A web app, a library, a CLI, a daemon? Each one gets drawn differently, and the README won't tell you reliably because READMEs are written to sell.

Manifests are more honest. Revibe's analyzer reads dependency files (package.json, go.mod, Cargo.toml, requirements.txt) before it reads anything else, and only falls back to a model reading the README when those are inconclusive. Try it on Ollama:

$ head -3 go.mod
module github.com/ollama/ollama

go 1.26.0

$ grep -E "gin-gonic|cobra" go.mod
	github.com/gin-gonic/gin v1.10.0
	github.com/spf13/cobra v1.7.0

$ cat LLAMA_CPP_VERSION
b11081
Three commands, and the shape is already visible.

Gin is an HTTP server framework and Cobra is a CLI framework, so Ollama is a Go program that is both a server and a command-line tool. Then there's that last file, which pins a specific build of llama.cpp, a C++ project. Two languages plus a pinned C++ dependency means there's a hard boundary somewhere between Go and C++, and that boundary is going to be the most interesting line on the diagram. You know this before reading a single function.


2. Throw away almost every file

Knowing the type raises the next question: which files actually matter? Start by deleting things from consideration. Tests, vendored code, generated files, build output, docs and fixtures never earn a box. In Ollama that's 350 test files gone immediately.

Then rank what's left. When Revibe analyzed Ollama for the Ollama architecture diagram in the gallery, this step took 1,168 candidate files down to 39. The ones that rank high tend to be the same four kinds:

  • Entry points like main.go and cmd/, where execution starts.
  • Contracts like api/types.go, which define the shape of every request and response.
  • Routers like server/routes.go, which work as the table of contents for the whole server.
  • Things everything imports, which you find by counting rather than guessing.

That last one is a single command in any Go repo:

$ grep -rhoE '"github.com/ollama/ollama/[a-z0-9/_]+"' \
    --include='*.go' --exclude='*_test.go' . \
  | sort | uniq -c | sort -rn | head -4

 107 "github.com/ollama/ollama/api"
  61 "github.com/ollama/ollama/mlx"
  56 "github.com/ollama/ollama/types/model"
  41 "github.com/ollama/ollama/envconfig"
Import counts: a crude but honest map of what the code leans on.

Read these carefully, because raw counts mislead in two directions. api at 107 is the real thing: it's the contract, and the diagram's gateway should be shaped around it. mlx at 61 is mostly the mlxrunner/ package importing its own internals, which tells you that package is cohesive but not that it's central. envconfig at 41 is configuration, which is load-bearing but not diagram-worthy for the same reason logging isn't: everything touches it, so drawing it tells the reader nothing.

Filenames mislead too. Revibe's first-pass ranking put template/zephyr.json in Ollama's top 39 on the strength of its name alone. It's a prompt template for one model, which makes it data rather than structure. That's why Revibe blends a model's judgment of each file with an import-graph score instead of trusting either one alone.


3. Find the seams, not the folders

Thirty-nine files is still thirty-nine boxes, so they need grouping. The obvious move is one box per top-level folder, and it's the wrong one. Folders describe how code is stored. An architecture diagram describes how it runs.

Look for seams instead: the places where a request crosses a real boundary such as a network port, a process, a disk or someone else's service. Ollama has three, and each one is a single grep away.

Seam one is port 11434. The ollama CLI doesn't run models itself. Every command in cmd/cmd.go starts with api.ClientFromEnvironment() and then makes HTTP calls to 127.0.0.1:11434, the default set in envconfig/config.go. ollama run is an HTTP client wearing a trench coat. That means the CLI, the desktop app and your own code are all peers on the diagram, since they're all just clients of one server.

Seam two is a process boundary. Remember the Go/C++ line from step 1? Here it is, in the header comment of llm/llama_server.go: the file "wraps the llama-server binary as a subprocess." Ollama doesn't link llama.cpp into its own process. It starts llama.cpp's server as a separate program, patched by the files in llama/compat/, and talks to it over HTTP on a private localhost port. A crash in the C++ engine kills that child process and leaves the Go server standing, which is a design decision worth a box of its own.

Seam three is disk. Models live under a manifests directory plus a pile of content-addressed blobs named sha256-… (see manifest/). Two models that share a layer share the file. Storage is a cylinder on any diagram, and this one has an interesting property worth labelling.

With the seams found, name each group by what it does. Revibe's generator prompt is blunt about this: organize around business capabilities like "Booking Engine" or "Payment Processing," not generic tiers like "Frontend" and "Backend." For Ollama that gives you a gateway, a scheduler, the inference runtimes and a model library. The names alone already tell a reader more than the three-box diagram did.


4. Trace one real request to find the arrows

The boxes are settled, so which arrows go between them? Don't infer arrows from imports, because an import says code could call something, not that it does. Follow one real request instead.

Pick a request someone actually sends in 2026. A common one is pointing Claude Code, or any app built on the Anthropic SDK, at a local Ollama so the model runs on your own GPU. That sends a POST /v1/messages to port 11434. Find it in server/routes.go, and something interesting shows up three times:

// server/routes.go (simplified: logging and cloud wrappers removed)
r.POST("/api/chat",
    s.ChatHandler)
r.POST("/v1/chat/completions",
    middleware.ChatMiddleware(lookupThinking), s.ChatHandler)
r.POST("/v1/messages",
    middleware.AnthropicMessagesMiddleware(lookupThinking), s.ChatHandler)
Three API dialects and one handler.

Ollama's native API, the OpenAI-compatible API and the Anthropic-compatible API all end at the same function. The middleware translates each dialect into Ollama's own request shape on the way in, and translates the streamed response back on the way out. This is the adapter pattern at the edge, and it changes the drawing. The right diagram has one handler with three doors in front of it, not three parallel API stacks. The folder view would never have shown you that.

Keep following. ChatHandler asks the scheduler for a runner, and the scheduler in server/sched.go makes the one decision that matters for performance: is this model already loaded? If not, is there enough free GPU memory (checked through discover/) to load it, or does an idle model need to be evicted first? Once it has a runner, the handler streams tokens from it over that private localhost port.

sequenceDiagram participant A as Anthropic SDK / Claude Code participant M as Anthropic middleware participant H as ChatHandler participant S as Scheduler participant R as llama-server (child process) participant D as Model blobs on disk A->>M: POST /v1/messages (port 11434) M->>H: same request, rewritten as an Ollama chat H->>S: getRunner(model) alt model already loaded S-->>H: existing runner else not loaded S->>S: enough free VRAM? evict an idle model if not S->>R: start llama-server with the model path R->>D: load weights S-->>H: new runner end H->>R: HTTP 127.0.0.1 /completion loop each token R-->>H: token H-->>A: streamed event, in Anthropic format end

The request crosses HTTP twice: once into Ollama, once into its own child process.

One trace has told you almost every arrow. A second trace, ollama pull, adds the rest: /api/pull goes to server/download.go, which fetches each blob in 16 parallel parts and writes it into the same store the runners read from.


5. Draw it, with rules

Now, and only now, open the drawing tool. These rules are the constraints Revibe's diagram generator is held to, and they work just as well by hand:

  • One question per diagram. This one answers "what happens to a prompt?" Anything that doesn't serve that question goes on a different diagram.
  • Fifteen nodes, maximum. Past that, readers stop following arrows and start scanning for their own team's box.
  • Name layers by capability. "Scheduling," not "server/."
  • Label every arrow with a verb or a protocol. An unlabelled arrow is a guess the reader has to make for you.
  • Storage is a cylinder, and optional paths are dotted. Readers already know this vocabulary, so use it.
  • Put the path in the box. A reader who doubts a box should be able to open the file and check.

Applied to everything from steps 1 through 4, you get this:

flowchart TD subgraph L1["1. Clients"] CLI["ollama CLI
cmd/cmd.go"] APP["Desktop app
app/"] SDK["OpenAI / Anthropic SDKs,
coding agents"] end subgraph L2["2. API gateway: one port, three dialects"] NATIVE["Native API
/api/chat, /api/generate"] COMPAT["Compat middleware
/v1/chat/completions, /v1/messages"] HANDLER["Chat and Generate handlers
server/routes.go"] end subgraph L3["3. Scheduling"] SCHED["Scheduler
server/sched.go"] DISC["GPU discovery
discover/"] end subgraph L4["4. Inference runtimes (child processes)"] LLAMA["llama-server + Ollama patches
llm/llama_server.go, llama/compat/"] MLX["MLX runner
mlxrunner/"] end subgraph L5["5. Model library"] PULL["Registry pull
server/download.go"] STORE[("Manifests + sha256 blobs
manifest/")] end CLOUD["Ollama Cloud
server/cloud_proxy.go"] CLI -- "HTTP :11434" --> NATIVE APP -- "HTTP :11434" --> NATIVE SDK -- "HTTP :11434" --> COMPAT COMPAT -- "translates request" --> HANDLER NATIVE --> HANDLER HANDLER -- "getRunner()" --> SCHED SCHED -- "checks free VRAM" --> DISC SCHED -- "GGUF models: start or reuse" --> LLAMA SCHED -. "MLX models" .-> MLX HANDLER -- "HTTP /completion, streams tokens" --> LLAMA LLAMA -- "loads weights" --> STORE NATIVE -- "/api/pull" --> PULL PULL -- "16 parallel parts" --> STORE HANDLER -. "cloud models" .-> CLOUD

Ollama's architecture: 13 boxes, every one traceable to a path. Drawn from the September 26, 2026 commit.

Compare it to the three-box version at the top. This one makes claims a reader can act on. The CLI has no special powers. Adding a new API dialect means writing middleware, not a new handler. If inference crashes, look at the child process, and if the second model is slow to start, look at the scheduler's eviction logic. And every one of those claims can be checked, because each box names where it lives.

Copy the Mermaid source for this diagram
flowchart TD
    subgraph L1["1. Clients"]
        CLI["ollama CLI<br/>cmd/cmd.go"]
        APP["Desktop app<br/>app/"]
        SDK["OpenAI / Anthropic SDKs,<br/>coding agents"]
    end
    subgraph L2["2. API gateway: one port, three dialects"]
        NATIVE["Native API<br/>/api/chat, /api/generate"]
        COMPAT["Compat middleware<br/>/v1/chat/completions, /v1/messages"]
        HANDLER["Chat and Generate handlers<br/>server/routes.go"]
    end
    subgraph L3["3. Scheduling"]
        SCHED["Scheduler<br/>server/sched.go"]
        DISC["GPU discovery<br/>discover/"]
    end
    subgraph L4["4. Inference runtimes (child processes)"]
        LLAMA["llama-server + Ollama patches<br/>llm/llama_server.go, llama/compat/"]
        MLX["MLX runner<br/>mlxrunner/"]
    end
    subgraph L5["5. Model library"]
        PULL["Registry pull<br/>server/download.go"]
        STORE[("Manifests + sha256 blobs<br/>manifest/")]
    end
    CLOUD["Ollama Cloud<br/>server/cloud_proxy.go"]

    CLI -- "HTTP :11434" --> NATIVE
    APP -- "HTTP :11434" --> NATIVE
    SDK -- "HTTP :11434" --> COMPAT
    COMPAT -- "translates request" --> HANDLER
    NATIVE --> HANDLER
    HANDLER -- "getRunner()" --> SCHED
    SCHED -- "checks free VRAM" --> DISC
    SCHED -- "GGUF models: start or reuse" --> LLAMA
    SCHED -. "MLX models" .-> MLX
    HANDLER -- "HTTP /completion, streams tokens" --> LLAMA
    LLAMA -- "loads weights" --> STORE
    NATIVE -- "/api/pull" --> PULL
    PULL -- "16 parallel parts" --> STORE
    HANDLER -. "cloud models" .-> CLOUD

Quote your labels. Mermaid breaks on unquoted labels containing ( ) : /, and file paths are full of them. A[Scheduler (server/sched.go)] fails to render, while A["Scheduler (server/sched.go)"] works. Revibe's own generator learned this the hard way on Rust paths like api::core::ciphers.


6. Check it against the code, then date it

How do you know the diagram is right? Run the citation test. For every box, name the directory it lives in, and for every arrow, name the function call or protocol that makes it real. If you can't cite it, it's a guess, so either go and find the citation or delete the element.

Then put a date on it, because architecture diagrams have a half-life, and Ollama's is short. Revibe's gallery analysis of Ollama ran on July 17. On the September 26 commit, 15 of the 39 files it ranked as most important no longer exist. cmd/runner/ is gone, the experimental x/mlxrunner/ graduated to a top-level mlxrunner/, x/imagegen/ has vanished, and the Go tokenizer/ package went with it. The boxes mostly survived and the paths inside them didn't, which is exactly why the paths belong in the boxes: a stale path is how you find out the diagram is stale.

An undated architecture diagram is a claim about a codebase that no longer exists. Put the commit hash or the date in the caption, like the one above, and redraw when the paths stop resolving.


The short version

Classify from the manifests, filter hard, group by seams rather than folders, get your arrows from real traces, draw with rules, and date the result. By hand, on a repo Ollama's size, that's an afternoon of grep and reading. It's also the exact pipeline Revibe runs, and the free architecture diagram generator does it for any GitHub repo in about a minute, with the file paths in the boxes so you can run the citation test yourself.

An architecture diagram is a list of boundaries, drawn. Find the boundaries first, and the drawing is the easy part.

See it done on a real repo

The gallery has full architecture breakdowns of open-source projects, each with its layered diagram, key modules and traced user flows. Start with Ollama, then compare how the same method draws opencode, an AI coding agent, and nanochat, a full LLM training stack.