The Multi-Agent Local AI Promise

Autonomous multi-agent architectures—built with frameworks like LangGraph, CrewAI, and AutoGen—are transforming software development. Instead of asking one model to do everything, you orchestrate specialized subagents: one reads Jira tickets, one parses codebase ASTs, one writes unit tests, and one conducts security linting. When connected to cloud APIs like Claude 3.5 Sonnet or GPT-4o, these agents execute in parallel effortlessly.

However, when engineering teams attempt to run these multi-agent pipelines completely on-premise or on a local workstation to avoid API bills and data leakage, they encounter an immediate bottleneck: Ollama crashes, prompts time out with 500 Internal Server Error, or system RAM usage multiplies instantaneously.

How Ollama Handles Concurrency: The KV-Slot Architecture

By default, Ollama operates with a single execution slot (OLLAMA_NUM_PARALLEL=1). When your LangGraph workflow dispatches three subagents simultaneously, Ollama places requests 2 and 3 into a FIFO queue. Subagent 2 sits idle waiting for Subagent 1 to finish streaming all tokens.

To speed things up, developers often increase concurrency by setting OLLAMA_NUM_PARALLEL=4 in their environment. But here is the architectural catch: Each parallel slot allocates an independent Key-Value cache in VRAM.

┌─────────────────────────────────────────────────────────────────┐ │ MULTI-AGENT LOCAL DISPATCH: PARALLEL SLOTS vs. VRAM FOOTPRINT │ └─────────────────────────────────────────────────────────────────┘ [ LangGraph / CrewAI Agent Engine ] ├── Subagent 1 (Researcher) ──┐ ├── Subagent 2 (Coder) ──┼── Simultaneous HTTP Requests └── Subagent 3 (Reviewer) ──┘ │ ▼ [ Ollama Request Dispatcher ] (OLLAMA_NUM_PARALLEL=3) ├── Slot 0: Static Weights + Dynamic KV Cache A (2.4 GB) ├── Slot 1: Static Weights + Dynamic KV Cache B (2.4 GB) └── Slot 2: Static Weights + Dynamic KV Cache C (2.4 GB) │ ▼ [ Total VRAM Required ] = Base Model (8GB) + (3 × 2.4GB) = 15.2 GB!

The Multiplication Table: Memory Cost per Concurrent Slot

Consider running qwen2.5-coder:14b with an 8k context window. Let us calculate the true memory cost across multiple concurrency levels:

OLLAMA_NUM_PARALLEL Model Weights Total KV-Cache VRAM Total Memory Allocated System Compatibility
1 Slot (Default) 9.2 GB 1.8 GB 11.0 GB 16GB & 18GB Macs
2 Slots 9.2 GB 3.6 GB 12.8 GB 18GB & 24GB Macs
4 Slots (Multi-Agent) 9.2 GB 7.2 GB 16.4 GB 32GB & 36GB Macs
8 Slots (Heavy Team) 9.2 GB 14.4 GB 23.6 GB 48GB & 64GB Macs

Configuring Ollama Concurrency via macOS launchd

On macOS, Ollama runs as a background service managed by launchd. To permanently configure concurrency and model retention without crashing your terminal, update your environment configuration:

# Stop the current Ollama application
$ pkill ollama

# Configure environment variables in your launch agent or shell:
$ launchctl setenv OLLAMA_NUM_PARALLEL "3"
$ launchctl setenv OLLAMA_MAX_LOADED_MODELS "1"
$ launchctl setenv OLLAMA_KEEP_ALIVE "10m"

# Or launch directly in terminal to test:
$ OLLAMA_NUM_PARALLEL=3 OLLAMA_MAX_LOADED_MODELS=1 ollama serve

Implementing a Token Bucket Request Broker in Python

If your multi-agent framework spawns 10+ subagents, never allow them to hit Ollama directly. Implement a lightweight local token bucket broker that limits concurrent HTTP connections to your configured slot limit:

import asyncio
import httpx

class LocalAgentBroker:
    def __init__(self, max_concurrent=2):
        self.semaphore = asyncio.Semaphore(max_concurrent)
        self.client = httpx.AsyncClient(base_url="http://localhost:11434")

    async def generate(self, model: str, prompt: str):
        async with self.semaphore:
            # Only max_concurrent requests can hold GPU VRAM simultaneously
            response = await self.client.post("/api/generate", json={
                "model": model,
                "prompt": prompt,
                "stream": False
            }, timeout=120.0)
            return response.json()["response"]
Multi-Agent Governor with ContextWarden: Running heavy agent workflows? ContextWarden actively monitors process memory spikes across parallel slots. When multi-agent concurrency approaches dangerous swap thresholds, ContextWarden dynamically adjusts background compiler priorities to keep your system responsive.