Sovereign Intelligence Systems
33 of 42
Chapter 33 of 42

Local AI Proxy Setup

5 min read · 1,158 words

Local AI Proxy Setup Guide

Run any model — local or cloud — through one OpenAI-compatible URL that every tool in your stack can share. Cursor, your web app, your agents, your scripts: one endpoint, one key, all models.

What You're Building

Your app / Cursor / scripts
│
│  POST /v1/chat/completions
│  Authorization: Bearer <your-master-key>
│  { "model": "my-local-model" }
▼
┌──────────────────────────────────┐
│  LiteLLM Proxy (:4000)          │
│  OpenAI-compatible /v1 API      │
│  Routes by model name:          │
│    "local-model" → Ollama       │
│    "deepseek"    → DeepSeek API │
│    "mimo"        → MiMo API     │
│    "embedder"    → Ollama       │
└────────┬─────────────┬──────────┘
│             │
▼             ▼
Ollama (:11434)   Cloud APIs
(your GPU)        (DeepSeek, MiMo, etc.)

Prerequisites

  • <Text tone="narrative"> A Mac with Apple Silicon (M1/M2/M3/M4) or a Linux box with a GPU </Text>
  • <Text tone="narrative"> Python 3.11+ </Text>
  • <Text tone="narrative"> Node.js 20+ (for Tailscale CLI and optional tooling) </Text>
  • <Text tone="narrative"> A Tailscale account (free tier works) </Text>

Step 1: Install Ollama

# macOS
brew install ollama

# Linux
curl -fsSL https://ollama.com/install.sh | sh

Pull the models you want to serve locally:

ollama pull qwen3:8b            # good starter (8B, fast)
ollama pull bge-m3              # embeddings (1024-dim, multilingual)
ollama pull qwen3.6:27b-q8_0   # larger reasoning model (needs 32GB+ RAM)

Verify Ollama is running:

curl http://127.0.0.1:11434/api/tags | python3 -m json.tool

Step 2: Install LiteLLM

Create a dedicated directory and Python venv:

mkdir -p ~/ai-proxy && cd ~/ai-proxy
python3 -m venv .venv
source .venv/bin/activate
pip install -U pip litellm

Step 3: Write Your Config

Create ~/ai-proxy/litellm.config.yaml:

# LiteLLM config — edit model_name entries to match what you want
# to call from Cursor / your apps.

model_list:
# --- LOCAL MODELS (Ollama) ---

# Your main local chat model
- model_name: local-chat
litellm_params:
model: ollama_chat/qwen3:8b
api_base: http://127.0.0.1:11434
keep_alive: "24h"
max_tokens: 4096
num_ctx: 8192

# Embeddings
- model_name: local-embed
litellm_params:
model: ollama/bge-m3:latest
api_base: http://127.0.0.1:11434
keep_alive: "5m"

# --- CLOUD MODELS ---

# DeepSeek V4 Pro (~$0.44/M input, $0.87/M output)
- model_name: deepseek-v4-pro
litellm_params:
model: deepseek/deepseek-v4-pro
api_key: os.environ/DEEPSEEK_API_KEY
max_tokens: 16384

# DeepSeek V4 Flash (~$0.04/M input, $0.09/M output)
- model_name: deepseek-v4-flash
litellm_params:
model: deepseek/deepseek-v4-flash
api_key: os.environ/DEEPSEEK_API_KEY
max_tokens: 8192

# MiMo V2.5 — Xiaomi 310B MoE, 1M context ($0.14/M in, $0.28/M out)
- model_name: mimo-v2.5
litellm_params:
model: openai/mimo-v2.5
api_base: https://api.xiaomimimo.com/v1
api_key: os.environ/MIMO_API_KEY
max_tokens: 16384

# MiniMax M3 — 1M context ($0.14/M in, $0.28/M out)
- model_name: minimax-m3
litellm_params:
model: openai/MiniMax-M3
api_base: https://api.minimax.io/v1
api_key: os.environ/MINIMAX_API_KEY
max_tokens: 16384

litellm_settings:
drop_params: true

general_settings:
master_key: os.environ/LITELLM_MASTER_KEY

Step 4: Set Your Secrets

Create ~/.ai-proxy-keys (chmod 600):

cat > ~/.ai-proxy-keys << 'EOF'
export LITELLM_MASTER_KEY="sk-your-chosen-password-here"
export DEEPSEEK_API_KEY="sk-..."
export MIMO_API_KEY="sk-..."
export MINIMAX_API_KEY="eyJ..."
EOF
chmod 600 ~/.ai-proxy-keys

Get your API keys from:

  • <Text tone="narrative"> : https\://platform.deepseek.com/api\_keys </Text>
  • <Text tone="narrative"> : https\://platform.xiaomimimo.com </Text>
  • <Text tone="narrative"> : https\://platform.minimaxi.com </Text>

The LITELLM_MASTER_KEY is any string you choose — it's the password your clients use to talk to the proxy. Make it strong; it gates access to all your models.


Step 5: Start the Proxy

cd ~/ai-proxy
source .venv/bin/activate
source ~/.ai-proxy-keys

.venv/bin/litellm --config litellm.config.yaml --port 4000

Test it:

# List available models
curl -s http://127.0.0.1:4000/v1/models \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" | python3 -m json.tool

# Chat with your local model
curl -s http://127.0.0.1:4000/v1/chat/completions \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "local-chat",
"messages": [{"role": "user", "content": "Say hello in three words."}],
"max_tokens": 20
}' | python3 -m json.tool

Step 6: Expose It With Tailscale Funnel

This gives you a stable HTTPS URL reachable from anywhere — your phone, Vercel, another machine.

# Install Tailscale (if not already)
# macOS: brew install tailscale  OR  download from tailscale.com
# Linux: curl -fsSL https://tailscale.com/install.sh | sh

# Log in
tailscale up

# Expose your proxy to the internet via Funnel
tailscale funnel 4000

Tailscale prints your public URL, something like:

https://your-machine.tail-network.ts.net

This persists across reboots. Now your proxy is reachable from anywhere as:

https://your-machine.tail-network.ts.net/v1

Test from another machine:

curl -s https://your-machine.tail-network.ts.net/v1/models \
-H "Authorization: Bearer sk-your-chosen-password-here"

Step 7: Connect Cursor IDE

In Cursor:

  • <Text tone="narrative"> → → scroll to </Text>
  • <Text tone="narrative"> Paste your LITELLM_MASTER_KEY value </Text>
  • <Text tone="narrative"> Set to: </Text>
https://your-machine.tail-network.ts.net/v1

(or http://127.0.0.1:4000/v1 if Cursor runs on the same machine)

  • <Text tone="narrative"> Under , add your custom models: </Text>
  • <Text tone="narrative"> local-chat </Text>
  • <Text tone="narrative"> deepseek-v4-pro </Text>
  • <Text tone="narrative"> deepseek-v4-flash </Text>
  • <Text tone="narrative"> mimo-v2.5 </Text>
  • <Text tone="narrative"> minimax-m3 </Text>

Now when you select any of those models in Cursor's model picker, it routes through your proxy to the right backend.


Step 8: Connect Your Web App (Vercel AI SDK)

import { createOpenAI } from '@ai-sdk/openai';
import { streamText } from 'ai';

const llm = createOpenAI({
baseURL: process.env.LOCAL_LLM_BASE_URL || 'http://127.0.0.1:4000/v1',
apiKey: process.env.LOCAL_LLM_API_KEY,  // your LITELLM_MASTER_KEY
compatibility: 'compatible',
});

const result = streamText({
model: llm.chat('local-chat'),  // or 'deepseek-v4-pro', 'mimo-v2.5', etc.
messages: [{ role: 'user', content: 'Hello!' }],
});

On Vercel, set these environment variables:

  • <Text tone="narrative"> LOCAL_LLM_BASE_URL = https://your-machine.tail-network.ts.net/v1 </Text>
  • <Text tone="narrative"> LOCAL_LLM_API_KEY = your LITELLM_MASTER_KEY </Text>

Step 9: Make It Survive Reboots (macOS)

Create ~/Library/LaunchAgents/com.local.litellm.plist:

<?xml version="1.0" encoding="UTF-8"?>
<plist version="1.0">
<dict>
<key>Label</key>
<string>com.local.litellm</string>
<key>ProgramArguments</key>
<array>
<string>/bin/bash</string>
<string>-lc</string>
<string>source ~/.ai-proxy-keys && ~/ai-proxy/.venv/bin/litellm --config ~/ai-proxy/litellm.config.yaml --port 4000</string>
</array>
<key>WorkingDirectory</key>
<string>/Users/YOU/ai-proxy</string>
<key>RunAtLoad</key>
<true/>
<key>KeepAlive</key>
<true/>
<key>StandardOutPath</key>
<string>/tmp/litellm.out.log</string>
<key>StandardErrorPath</key>
<string>/tmp/litellm.err.log</string>
</dict>
</plist>

Load it:

# Replace YOU with your username in the plist first, then:
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.local.litellm.plist

For Linux, use a systemd unit instead.


Quick Reference

WhatURLAuth
Local (same machine)`http://127.0.0.1:4000/v1`Bearer `<LITELLM_MASTER_KEY>`
Remote (Tailscale Funnel)`https://your-machine.tail-network.ts.net/v1`Same
List models`GET /v1/models`Same
Chat`POST /v1/chat/completions`Same
Embeddings`POST /v1/embeddings`Same
Health check`GET /health/liveness`None

Adding More Models Later

Just add a new entry to litellm.config.yaml and restart:

- model_name: my-new-model
litellm_params:
model: openai/some-model-name     # "openai/" prefix for any OpenAI-compatible API
api_base: https://api.provider.com/v1
api_key: os.environ/PROVIDER_API_KEY
max_tokens: 16384

For local Ollama models, use ollama_chat/ (chat) or ollama/ (embeddings) prefix:

- model_name: my-local-llama
litellm_params:
model: ollama_chat/llama3.3:70b
api_base: http://127.0.0.1:11434
keep_alive: "1h"

No client code changes needed — just use the new model_name in your requests.


Troubleshooting

SymptomCauseFix
401 UnauthorizedWrong key in clientUse `LITELLM_MASTER_KEY`, not the provider key
Model not foundTypo in model nameMust match `model_name` in YAML exactly
Ollama timeoutModel not pulledRun `ollama pull <model>` first
Cloud 401/502Provider key missingCheck `~/.ai-proxy-keys` has the right export
Funnel unreachableTailscale not running`tailscale up` then `tailscale funnel 4000`
Port already in useStale process`lsof -ti :4000 | xargs kill -9` then restart

Why This Architecture

  • <Text tone="narrative"> Your app code doesn't know or care whether it's hitting a 7B model on your GPU or DeepSeek's cloud. Switch models by changing a string. </Text>
  • <Text tone="narrative"> Add a new provider by adding 5 lines of YAML. Drop one by removing them. </Text>
  • <Text tone="narrative"> Route cheap tasks to local models ($0), expensive reasoning to cloud ($0.001/turn). Your proxy, your rules. </Text>
  • <Text tone="narrative"> Local models never leave your machine. Cloud calls are opt-in per model name. </Text>
  • <Text tone="narrative"> Any tool that speaks the OpenAI API spec (Cursor, Vercel AI SDK, LangChain, OpenAI Python client, curl) works unchanged. </Text>