
Two founders wrapped someone else's hard infrastructure in a friendly interface once before, sold it to Docker, and made a career of it. This time they kept the company - and are trying to keep the trick, too.
Pairs with the Moat Anatomy Canvas — a ready-to-use strategy tool, filled for Ollama. Get it — included with a subscription, or $1.99 →
Type one line - ollama run llama3 - and about ninety seconds later, a language model that would have needed a research team and a rack of GPUs three years ago is answering questions on a laptop. No cloud account, no API key, no bill. That trick just earned its maker a $65 million Series B on top of 8.9 million developers running it every month, a spot inside 85% of the Fortune 500, and 176,000 stars on GitHub.12 The natural assumption is that Ollama built the trick. It didn't.
The engine that actually does the work - compressing a multi-billion-parameter model down to something a laptop's chip can run - is llama.cpp, a separate open-source project a Bulgarian engineer named Georgi Gerganov started in early 2023, weeks after Meta's LLaMA weights leaked online.9 Ollama is a wrapper: a one-line installer, a curated model registry, and a friendly command sitting on top of an inference engine it didn't invent. For the better part of a year after Ollama's own July 2023 launch, its documentation never said so. A developer had to open a GitHub issue in April 2024 just to ask for a mention, writing that the project was 'heavily dependent on llama.cpp' with 'no mention of that in the readme.'6 Ollama's co-founder eventually added one line, roughly halfway down the README, under a header called 'Supported backends': 'llama.cpp project founded by Georgi Gerganov.'5 A separate GitHub issue, open since March 2024, still tracks a related complaint - that Ollama's release binaries didn't include the copyright notice llama.cpp's MIT license requires.7
Ollama's moat was never the engine. It's the on-ramp - and the two people who built it are running the exact play that already made them money once, wrapping a different piece of hard infrastructure somebody else had built.
It didn't build the engine. It built the front door - and this isn't the first door these two have sold.
The founders have run this exact play before: the same interface trick, a different engine underneath
Jeff Morgan and Michael Chiang aren't new to this. Before Ollama, they built Kitematic, a graphical interface that made Docker - genuinely hard, unix-permissions-and-namespaces infrastructure - installable and runnable on a Mac in minutes. Docker Inc. announced its acquisition of Kitematic on March 12, 2015, on undisclosed terms, and folded it into what eventually became Docker Desktop.34 Docker didn't buy Kitematic for a container engine - Kitematic didn't have one; it drove Docker's own engine underneath. It bought the front door. Ollama is the same acquisition thesis, run by the same two people, minus the acquirer: instead of selling the front door to the company that owns the engine, they built the front door for an engine nobody owns, and this time they're trying to keep the toll themselves.
| Kitematic → Docker Desktop (2015) | llama.cpp → Ollama (2023-) | |
|---|---|---|
| The hard infrastructure | Docker Engine - containers, namespaces, cgroups | llama.cpp - the C++ inference engine, quantization, GPU offload |
| Who built the hard part | Docker, Inc. | Georgi Gerganov and the llama.cpp community |
| What Morgan & Chiang built | A one-click GUI to install and run it | A one-line installer, model registry, and API |
| How the engine got credited | Folded straight into Docker's own brand | One README line, added after a public request |
| How it made money | Docker paid to own the interface | Ollama sells a cloud subscription on top of it |
What Ollama actually owns: distribution, not the difficult part
None of this makes Ollama a fraud. Distribution is real value - arguably the only kind of value that compounds in open source. The company built a genuinely good curated model registry, a one-line install that behaves the same on a Mac, a Windows box, or a Linux server, and - since 2024 - an OpenAI-compatible API that lets developers point existing code at it without a rewrite.11 That's the product 8.9 million developers show up for every month, and it's what Ollama is now monetizing directly: Ollama Cloud, a subscription layered on the same free local tool, priced at $0, $20 a month, and $100 a month, billed by GPU-time rather than a token count, for models 'too big to run on your own computer.'8 It's the Kitematic trade again - get millions of people hooked on a free, frictionless interface, then sell the next layer up before anyone else builds the same front door.
Ask three questions. Who actually built the hard technical part, and would they let a rival use it too? What happens to your product if that outside project forks or stalls under you? And are you charging for the interface, or for the difficulty? Ollama passes the first two - llama.cpp is openly licensed and actively maintained, and Ollama has started building its own engine work for cases like multimodal models. It's still charging for the interface, priced by GPU-time rather than by how hard the problem is. That's honest pricing. It just means the moat is only as durable as the habit.
The honest objection: interfaces are real moats too: and Ollama is quietly paying down the debt that started this
The fair objection is that 'just a wrapper' undersells what a good interface is worth. Docker's own engine has been open and forkable since 2013; that never stopped Docker Desktop from becoming the default way people actually use it, because habit and polish are their own kind of lock-in. GitHub proved the same thing against raw git. Ollama is also less of a pure llama.cpp skin than its 2024 critics suggest - the company has been building pieces of its own inference engine, particularly for multimodal models that llama.cpp's architecture doesn't handle cleanly, rather than riding someone else's code forever.9 A company that starts by borrowing infrastructure and ends up owning more of it isn't running a con. It's walking the standard path from wrapper to platform - the one Docker walked, the one AWS walked on top of Linux, and the one Ollama is a few years into.
What an on-ramp can't lock down: the interface Ollama is betting on isn't proprietary either
The crack is that the interface itself is now a commodity. The OpenAI-compatible API - the exact habit Ollama's 8.9 million developers have built - isn't something Ollama owns. LM Studio, vLLM, and even llama.cpp's own official server all expose the identical routes, so a developer can repoint existing code at any of them with a one-line change.10 Kitematic's moat, such as it was, ended the day Docker decided to buy it instead of competing with it. Ollama doesn't have an acquirer waiting this time. It has a cloud giant, or Hugging Face, or Meta, any of which could bundle the same one-line local install into a tool millions of developers already have open - and erase the on-ramp overnight. The $65 million isn't proof the front door is locked. It's the clock Ollama is racing before someone bigger notices it never was.
Ollama didn't build a better engine. It built a better front door, and sold admission to the same people who already had a key to the building next to it. The last time Jeff Morgan and Michael Chiang made that exact bet, Docker paid to find out they were right. This time there's no buyer lined up - just a very fast clock, and a habit they're hoping 8.9 million developers keep long enough to matter.

Moat Anatomy Canvas
A one-page canvas that dissects a moat instead of asserting it: where the advantage comes from, how much of the market it covers, how long it would take to copy, and what keeps it from eroding. Blank to dissect your own claimed edge; filled as the worked example tracing the structure of the story's defensible advantage. Use it to tell a real moat from a head start.
Included with any subscription, or unlock this tool for $1.99. Get it → · See plans →
Sources
Where this comes from — the filings, records, and reporting behind it.
- 1Ollama raised a $65 million Series B led by Theory Ventures, on top of a $15 million Series A led by Benchmark ($88 million total raised); it is used by over 8.9 million developers every month, sits inside 85% of the Fortune 500, and has 176,000 GitHub stars and roughly 17,000 forks, run by a team of 14 employees.
- 2Corroborates Ollama's $65 million Series B, 8.9 million monthly developers, a 14-person team, and presence in 85% of the Fortune 500.
- 3Docker announced its acquisition of Kitematic, the graphical interface for running Docker on a Mac, on March 12, 2015; financial terms were not disclosed.
- 4Ollama's co-founders, Jeff Morgan and Michael Chiang, previously built Kitematic, the open-source Docker GUI startup that Docker Inc. acquired in 2015.
- 5Ollama's GitHub README credits its inference backend in one line, in a 'Supported backends' section roughly midway through the document: 'llama.cpp project founded by Georgi Gerganov.'Ollama (GitHub), ollama/ollama - README.md ↗ · 2026-07-25
- 6A GitHub issue titled 'No llama.cpp acknowledgement' was opened April 17, 2024, stating Ollama was 'heavily dependent on llama.cpp' with 'no mention of that in the readme.'
- 7A GitHub issue opened March 16, 2024 documented that Ollama's release binaries did not include the copyright notice required by the MIT license of statically linked dependencies including llama.cpp.
- 8Ollama Cloud has three published tiers: Free ($0, light usage, 1 concurrent model), Pro ($20/month or $200/year, 50x more usage than Free, 3 concurrent models), and Max ($100/month, 5x more usage than Pro, 10 concurrent models), billed by GPU-time rather than a fixed token count.Ollama, Ollama Pricing ↗ · 2026-07-25
- 9llama.cpp, created by Bulgarian engineer Georgi Gerganov in early 2023 to run LLaMA models in C/C++ with minimal setup, carries over 120,000 GitHub stars and 21,000 forks; Ollama and other local-AI interfaces are built on top of its ggml engine, though Ollama has since begun building its own engine components for cases such as multimodal models.GitHub (ggml-org/llama.cpp), ggml-org/llama.cpp ↗ · 2026-07-25
- 10Ollama, LM Studio, llama.cpp's own llama-server, and vLLM all expose OpenAI-style API routes, so client code written for the OpenAI SDK can be repointed at any of them with a one-line base-URL change.
- 11Ollama added a built-in OpenAI-compatible Chat Completions API endpoint, letting existing OpenAI-SDK code point at a local Ollama instance by changing only the base URL.Ollama, OpenAI compatibility ↗ · 2024-07-01
More like this — beyond Ollama
New Strategically analyses as they publish: the defining moves in business, checked against the record. No noise, and one click to leave.