The Multi-Agent Local AI Promise
Autonomous multi-agent architectures—built with frameworks like LangGraph, CrewAI, and AutoGen—are transforming software development. Instead of asking one model to do everything, you orchestrate specialized subagents: one reads Jira tickets, one parses codebase ASTs, one writes unit tests, and one conducts security linting. When connected to cloud APIs like Claude 3.5 Sonnet or GPT-4o, these agents execute in parallel effortlessly.
However, when engineering teams attempt to run these multi-agent pipelines completely on-premise or on a local workstation to avoid API bills and data leakage, they encounter an immediate bottleneck: Ollama crashes, prompts time out with 500 Internal Server Error, or system RAM usage multiplies instantaneously.
How Ollama Handles Concurrency: The KV-Slot Architecture
By default, Ollama operates with a single execution slot (OLLAMA_NUM_PARALLEL=1). When your LangGraph workflow dispatches three subagents simultaneously, Ollama places requests 2 and 3 into a FIFO queue. Subagent 2 sits idle waiting for Subagent 1 to finish streaming all tokens.
To speed things up, developers often increase concurrency by setting OLLAMA_NUM_PARALLEL=4 in their environment. But here is the architectural catch: Each parallel slot allocates an independent Key-Value cache in VRAM.
┌─────────────────────────────────────────────────────────────────┐
│ MULTI-AGENT LOCAL DISPATCH: PARALLEL SLOTS vs. VRAM FOOTPRINT │
└─────────────────────────────────────────────────────────────────┘
[ LangGraph / CrewAI Agent Engine ]
├── Subagent 1 (Researcher) ──┐
├── Subagent 2 (Coder) ──┼── Simultaneous HTTP Requests
└── Subagent 3 (Reviewer) ──┘
│
▼
[ Ollama Request Dispatcher ] (OLLAMA_NUM_PARALLEL=3)
├── Slot 0: Static Weights + Dynamic KV Cache A (2.4 GB)
├── Slot 1: Static Weights + Dynamic KV Cache B (2.4 GB)
└── Slot 2: Static Weights + Dynamic KV Cache C (2.4 GB)
│
▼
[ Total VRAM Required ] = Base Model (8GB) + (3 × 2.4GB) = 15.2 GB!
The Multiplication Table: Memory Cost per Concurrent Slot
Consider running qwen2.5-coder:14b with an 8k context window. Let us calculate the true memory cost across multiple concurrency levels:
| OLLAMA_NUM_PARALLEL | Model Weights | Total KV-Cache VRAM | Total Memory Allocated | System Compatibility |
|---|---|---|---|---|
| 1 Slot (Default) | 9.2 GB | 1.8 GB | 11.0 GB | 16GB & 18GB Macs |
| 2 Slots | 9.2 GB | 3.6 GB | 12.8 GB | 18GB & 24GB Macs |
| 4 Slots (Multi-Agent) | 9.2 GB | 7.2 GB | 16.4 GB | 32GB & 36GB Macs |
| 8 Slots (Heavy Team) | 9.2 GB | 14.4 GB | 23.6 GB | 48GB & 64GB Macs |
Configuring Ollama Concurrency via macOS launchd
On macOS, Ollama runs as a background service managed by launchd. To permanently configure concurrency and model retention without crashing your terminal, update your environment configuration:
# Stop the current Ollama application
$ pkill ollama
# Configure environment variables in your launch agent or shell:
$ launchctl setenv OLLAMA_NUM_PARALLEL "3"
$ launchctl setenv OLLAMA_MAX_LOADED_MODELS "1"
$ launchctl setenv OLLAMA_KEEP_ALIVE "10m"
# Or launch directly in terminal to test:
$ OLLAMA_NUM_PARALLEL=3 OLLAMA_MAX_LOADED_MODELS=1 ollama serve
Implementing a Token Bucket Request Broker in Python
If your multi-agent framework spawns 10+ subagents, never allow them to hit Ollama directly. Implement a lightweight local token bucket broker that limits concurrent HTTP connections to your configured slot limit:
import asyncio
import httpx
class LocalAgentBroker:
def __init__(self, max_concurrent=2):
self.semaphore = asyncio.Semaphore(max_concurrent)
self.client = httpx.AsyncClient(base_url="http://localhost:11434")
async def generate(self, model: str, prompt: str):
async with self.semaphore:
# Only max_concurrent requests can hold GPU VRAM simultaneously
response = await self.client.post("/api/generate", json={
"model": model,
"prompt": prompt,
"stream": False
}, timeout=120.0)
return response.json()["response"]