Local AI Proxy Setup
5 min read · 1,158 words
Local AI Proxy Setup Guide
Run any model — local or cloud — through one OpenAI-compatible URL that every tool in your stack can share. Cursor, your web app, your agents, your scripts: one endpoint, one key, all models.
What You're Building
Your app / Cursor / scripts
│
│ POST /v1/chat/completions
│ Authorization: Bearer <your-master-key>
│ { "model": "my-local-model" }
▼
┌──────────────────────────────────┐
│ LiteLLM Proxy (:4000) │
│ OpenAI-compatible /v1 API │
│ Routes by model name: │
│ "local-model" → Ollama │
│ "deepseek" → DeepSeek API │
│ "mimo" → MiMo API │
│ "embedder" → Ollama │
└────────┬─────────────┬──────────┘
│ │
▼ ▼
Ollama (:11434) Cloud APIs
(your GPU) (DeepSeek, MiMo, etc.)Prerequisites
- <Text tone="narrative"> A Mac with Apple Silicon (M1/M2/M3/M4) or a Linux box with a GPU </Text>
- <Text tone="narrative"> Python 3.11+ </Text>
- <Text tone="narrative"> Node.js 20+ (for Tailscale CLI and optional tooling) </Text>
- <Text tone="narrative"> A Tailscale account (free tier works) </Text>
Step 1: Install Ollama
# macOS
brew install ollama
# Linux
curl -fsSL https://ollama.com/install.sh | shPull the models you want to serve locally:
ollama pull qwen3:8b # good starter (8B, fast)
ollama pull bge-m3 # embeddings (1024-dim, multilingual)
ollama pull qwen3.6:27b-q8_0 # larger reasoning model (needs 32GB+ RAM)Verify Ollama is running:
curl http://127.0.0.1:11434/api/tags | python3 -m json.toolStep 2: Install LiteLLM
Create a dedicated directory and Python venv:
mkdir -p ~/ai-proxy && cd ~/ai-proxy
python3 -m venv .venv
source .venv/bin/activate
pip install -U pip litellmStep 3: Write Your Config
Create ~/ai-proxy/litellm.config.yaml:
# LiteLLM config — edit model_name entries to match what you want
# to call from Cursor / your apps.
model_list:
# --- LOCAL MODELS (Ollama) ---
# Your main local chat model
- model_name: local-chat
litellm_params:
model: ollama_chat/qwen3:8b
api_base: http://127.0.0.1:11434
keep_alive: "24h"
max_tokens: 4096
num_ctx: 8192
# Embeddings
- model_name: local-embed
litellm_params:
model: ollama/bge-m3:latest
api_base: http://127.0.0.1:11434
keep_alive: "5m"
# --- CLOUD MODELS ---
# DeepSeek V4 Pro (~$0.44/M input, $0.87/M output)
- model_name: deepseek-v4-pro
litellm_params:
model: deepseek/deepseek-v4-pro
api_key: os.environ/DEEPSEEK_API_KEY
max_tokens: 16384
# DeepSeek V4 Flash (~$0.04/M input, $0.09/M output)
- model_name: deepseek-v4-flash
litellm_params:
model: deepseek/deepseek-v4-flash
api_key: os.environ/DEEPSEEK_API_KEY
max_tokens: 8192
# MiMo V2.5 — Xiaomi 310B MoE, 1M context ($0.14/M in, $0.28/M out)
- model_name: mimo-v2.5
litellm_params:
model: openai/mimo-v2.5
api_base: https://api.xiaomimimo.com/v1
api_key: os.environ/MIMO_API_KEY
max_tokens: 16384
# MiniMax M3 — 1M context ($0.14/M in, $0.28/M out)
- model_name: minimax-m3
litellm_params:
model: openai/MiniMax-M3
api_base: https://api.minimax.io/v1
api_key: os.environ/MINIMAX_API_KEY
max_tokens: 16384
litellm_settings:
drop_params: true
general_settings:
master_key: os.environ/LITELLM_MASTER_KEYStep 4: Set Your Secrets
Create ~/.ai-proxy-keys (chmod 600):
cat > ~/.ai-proxy-keys << 'EOF'
export LITELLM_MASTER_KEY="sk-your-chosen-password-here"
export DEEPSEEK_API_KEY="sk-..."
export MIMO_API_KEY="sk-..."
export MINIMAX_API_KEY="eyJ..."
EOF
chmod 600 ~/.ai-proxy-keysGet your API keys from:
- <Text tone="narrative"> : https\://platform.deepseek.com/api\_keys </Text>
- <Text tone="narrative"> : https\://platform.xiaomimimo.com </Text>
- <Text tone="narrative"> : https\://platform.minimaxi.com </Text>
The LITELLM_MASTER_KEY is any string you choose — it's the password your clients use to talk to the proxy. Make it strong; it gates access to all your models.
Step 5: Start the Proxy
cd ~/ai-proxy
source .venv/bin/activate
source ~/.ai-proxy-keys
.venv/bin/litellm --config litellm.config.yaml --port 4000Test it:
# List available models
curl -s http://127.0.0.1:4000/v1/models \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" | python3 -m json.tool
# Chat with your local model
curl -s http://127.0.0.1:4000/v1/chat/completions \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "local-chat",
"messages": [{"role": "user", "content": "Say hello in three words."}],
"max_tokens": 20
}' | python3 -m json.toolStep 6: Expose It With Tailscale Funnel
This gives you a stable HTTPS URL reachable from anywhere — your phone, Vercel, another machine.
# Install Tailscale (if not already)
# macOS: brew install tailscale OR download from tailscale.com
# Linux: curl -fsSL https://tailscale.com/install.sh | sh
# Log in
tailscale up
# Expose your proxy to the internet via Funnel
tailscale funnel 4000Tailscale prints your public URL, something like:
https://your-machine.tail-network.ts.netThis persists across reboots. Now your proxy is reachable from anywhere as:
https://your-machine.tail-network.ts.net/v1Test from another machine:
curl -s https://your-machine.tail-network.ts.net/v1/models \
-H "Authorization: Bearer sk-your-chosen-password-here"Step 7: Connect Cursor IDE
In Cursor:
- <Text tone="narrative"> → → scroll to </Text>
- <Text tone="narrative"> Paste your LITELLM_MASTER_KEY value </Text>
- <Text tone="narrative"> Set to: </Text>
https://your-machine.tail-network.ts.net/v1(or http://127.0.0.1:4000/v1 if Cursor runs on the same machine)
- <Text tone="narrative"> Under , add your custom models: </Text>
- <Text tone="narrative"> local-chat </Text>
- <Text tone="narrative"> deepseek-v4-pro </Text>
- <Text tone="narrative"> deepseek-v4-flash </Text>
- <Text tone="narrative"> mimo-v2.5 </Text>
- <Text tone="narrative"> minimax-m3 </Text>
Now when you select any of those models in Cursor's model picker, it routes through your proxy to the right backend.
Step 8: Connect Your Web App (Vercel AI SDK)
import { createOpenAI } from '@ai-sdk/openai';
import { streamText } from 'ai';
const llm = createOpenAI({
baseURL: process.env.LOCAL_LLM_BASE_URL || 'http://127.0.0.1:4000/v1',
apiKey: process.env.LOCAL_LLM_API_KEY, // your LITELLM_MASTER_KEY
compatibility: 'compatible',
});
const result = streamText({
model: llm.chat('local-chat'), // or 'deepseek-v4-pro', 'mimo-v2.5', etc.
messages: [{ role: 'user', content: 'Hello!' }],
});On Vercel, set these environment variables:
- <Text tone="narrative"> LOCAL_LLM_BASE_URL = https://your-machine.tail-network.ts.net/v1 </Text>
- <Text tone="narrative"> LOCAL_LLM_API_KEY = your LITELLM_MASTER_KEY </Text>
Step 9: Make It Survive Reboots (macOS)
Create ~/Library/LaunchAgents/com.local.litellm.plist:
<?xml version="1.0" encoding="UTF-8"?>
<plist version="1.0">
<dict>
<key>Label</key>
<string>com.local.litellm</string>
<key>ProgramArguments</key>
<array>
<string>/bin/bash</string>
<string>-lc</string>
<string>source ~/.ai-proxy-keys && ~/ai-proxy/.venv/bin/litellm --config ~/ai-proxy/litellm.config.yaml --port 4000</string>
</array>
<key>WorkingDirectory</key>
<string>/Users/YOU/ai-proxy</string>
<key>RunAtLoad</key>
<true/>
<key>KeepAlive</key>
<true/>
<key>StandardOutPath</key>
<string>/tmp/litellm.out.log</string>
<key>StandardErrorPath</key>
<string>/tmp/litellm.err.log</string>
</dict>
</plist>Load it:
# Replace YOU with your username in the plist first, then:
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.local.litellm.plistFor Linux, use a systemd unit instead.
Quick Reference
| What | URL | Auth |
|---|---|---|
| Local (same machine) | `http://127.0.0.1:4000/v1` | Bearer `<LITELLM_MASTER_KEY>` |
| Remote (Tailscale Funnel) | `https://your-machine.tail-network.ts.net/v1` | Same |
| List models | `GET /v1/models` | Same |
| Chat | `POST /v1/chat/completions` | Same |
| Embeddings | `POST /v1/embeddings` | Same |
| Health check | `GET /health/liveness` | None |
Adding More Models Later
Just add a new entry to litellm.config.yaml and restart:
- model_name: my-new-model
litellm_params:
model: openai/some-model-name # "openai/" prefix for any OpenAI-compatible API
api_base: https://api.provider.com/v1
api_key: os.environ/PROVIDER_API_KEY
max_tokens: 16384For local Ollama models, use ollama_chat/ (chat) or ollama/ (embeddings) prefix:
- model_name: my-local-llama
litellm_params:
model: ollama_chat/llama3.3:70b
api_base: http://127.0.0.1:11434
keep_alive: "1h"No client code changes needed — just use the new model_name in your requests.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| 401 Unauthorized | Wrong key in client | Use `LITELLM_MASTER_KEY`, not the provider key |
| Model not found | Typo in model name | Must match `model_name` in YAML exactly |
| Ollama timeout | Model not pulled | Run `ollama pull <model>` first |
| Cloud 401/502 | Provider key missing | Check `~/.ai-proxy-keys` has the right export |
| Funnel unreachable | Tailscale not running | `tailscale up` then `tailscale funnel 4000` |
| Port already in use | Stale process | `lsof -ti :4000 | xargs kill -9` then restart |
Why This Architecture
- <Text tone="narrative"> Your app code doesn't know or care whether it's hitting a 7B model on your GPU or DeepSeek's cloud. Switch models by changing a string. </Text>
- <Text tone="narrative"> Add a new provider by adding 5 lines of YAML. Drop one by removing them. </Text>
- <Text tone="narrative"> Route cheap tasks to local models ($0), expensive reasoning to cloud ($0.001/turn). Your proxy, your rules. </Text>
- <Text tone="narrative"> Local models never leave your machine. Cloud calls are opt-in per model name. </Text>
- <Text tone="narrative"> Any tool that speaks the OpenAI API spec (Cursor, Vercel AI SDK, LangChain, OpenAI Python client, curl) works unchanged. </Text>
