Local models vs hosted models
Neither is better everywhere. Here's what changes when the model runs on your own hardware:
| Local model | Hosted model (Claude, GPT, Gemini) | |
|---|---|---|
| Privacy | Prompts and files never leave your machine | Sent to the provider's servers |
| Running cost | No per-message bill, just electricity | Pay per token, or use a plan's limits |
| Upfront cost | RAM, GPU and disk you may need to buy | Works on a small VPS or Raspberry Pi |
| Agent quality | Small models are more likely to fumble long, tool-heavy tasks | Strongest on complex, multi-step work |
| Speed | Depends on your hardware. CPU-only can take minutes per reply. | Fast and consistent |
| Offline | Works without internet | Needs a connection |
| Safety filters | None. You set the limits. | Provider-side filters included |
Many people use a hybrid setup: a hosted model for hard tasks with a local model as backup, or local first with a hosted safety net. The builder below writes either one.
What can your machine run?
Based on the minimum memory in OpenClaw's managed llama.cpp recommendations. These are floors, not guarantees of fit or speed.
Model sizes and memory needs
OpenClaw's managed llama.cpp setup picks from these models based on your memory, GPU and free disk. Each uses a 65,536-token context and supports tool calls.
- Qwen3.8 27B is the first pick when memory allows. Muse Glimmer can fit a 24 GiB NVIDIA card where Qwen3.8 doesn't.
- CPU-only machines top out at Qwen3.5 9B, and the first reply can take several minutes.
- Add about 0.3 GB for the default local embedding model (EmbeddingGemma), plus room for the runtime. Two NVIDIA cards aren't added together.
Managed backends: Metal on Apple silicon, CUDA 12.4 on Windows x64 with a supported NVIDIA GPU (driver 551.78+), and CPU elsewhere, including Linux. For CUDA on Linux, run your own server and choose Existing llama-server. See also OpenClaw system requirements.
Pick a local model backend
OpenClaw picks a model for your hardware, downloads and verifies it, and runs the server. Best first choice.
CLI workflow, big model library and a background service. Can mix in Ollama Cloud models.
Desktop app with a model browser. Supports the Responses API, which keeps reasoning out of replies.
High-throughput serving on your own GPU box through an OpenAI-compatible endpoint.
Put any OpenAI-style /v1/chat/completions server in front of OpenClaw.
Pulls models from OCI registries and can route small prompts locally, big ones to a hosted model.
Set up a local model
Install the plugin, run onboarding and choose Managed local server. Setup shows the host, model, download size and backend before downloading anything. It then checks that the model can answer and use a tool before making it your default.
openclaw plugins install @openclaw/llama-cpp-provider
openclaw onboard
# choose: Managed local serverEach check has a 90-second limit. If it fails, your previous default model stays selected. Already run llama-server yourself? Choose Existing llama-server instead.
Install from ollama.com/download, pull a model, then run onboarding and choose Ollama → Local only:
ollama pull gemma4
openclaw onboard
# choose: Ollama → Local only
openclaw models list --provider ollamaManual alternative: export OLLAMA_API_KEY="ollama-local" (any value works for a local host), then openclaw models set ollama/gemma4. Automatic discovery only offers models already loaded in memory that support tools and at least 16K of context.
OpenClaw talks to Ollama's native API at http://127.0.0.1:11434. The OpenAI-compatible /v1 URL breaks tool calling, so tool calls come out as plain JSON text.
Install LM Studio, load a model, and start the server. Then run onboarding and choose LM Studio:
lms server start --port 1234
openclaw onboard
# choose: LM Studio
openclaw models set lmstudio/qwen/qwen3.5-9bLM Studio model keys look like author/model. Find yours with curl http://localhost:1234/api/v1/models. Download the largest build your hardware can hold, and keep it loaded to avoid slow cold starts.
Start vLLM's OpenAI-compatible server, set any key, then pick the model:
vllm serve <model-id>
export VLLM_API_KEY="vllm-local"
openclaw models list --provider vllm
openclaw models set vllm/<model-id>vLLM usually runs at http://127.0.0.1:8000/v1. SGLang, MLX (mlx_lm.server) and LiteLLM work the same way as a custom provider with api: "openai-completions".
Hybrid setup builder
Mix local and hosted models
Fallbacks switch models on provider errors, one turn at a time. Copy the result into ~/.openclaw/openclaw.json and run openclaw gateway restart.
The hosted model still needs its own key or sign-in. See the Claude, OpenAI or Gemini guide.
Test before you trust it
A model that answers "hello" may still fail a real agent turn. Test in this order:
- Does the model respond? No tools, no agent context:
Terminal
openclaw infer model run --local --model <provider/model> --prompt "Reply with exactly: pong" --json - Does the gateway route to it?
Terminal
openclaw infer model run --gateway --model <provider/model> --prompt "Reply with exactly: pong" --json - Do real tasks work? Give it an actual job from your first workflow. If tool calls come out malformed, check the context size and server logs before anything else.
For Ollama, LM Studio and managed local servers, OpenClaw loads tool descriptions only when needed, so small models aren't swamped. If a model still struggles, lean mode removes optional tools such as browser, image, video and speech:
{
agents: {
defaults: {
experimental: { localModelLean: true },
},
},
}Speed and context tuning
Ollama setup uses a 32,768-token context, or less if the model's window is smaller. Bigger isn't free: it costs memory and slows the first token.
Large models load slowly. Raise the provider's timeoutSeconds and keep the model loaded with keep_alive.
Set input: ["text", "image"] on a local vision model so image attachments reach it.
{
models: {
providers: {
ollama: {
timeoutSeconds: 300,
models: [
{
id: "gemma4:26b",
name: "gemma4:26b",
params: { keep_alive: "15m" },
},
],
},
},
},
}Local models have no safety filter
Hosted providers screen requests. A local model doesn't, and smaller models are easier to trick with prompt injection hidden in web pages, emails or files.
More: OpenClaw security guide.
Fix common local model problems
Tool calls appear as JSON textOn Ollama, remove /v1 from the URL and use api: "ollama". On other servers, fix the chat template and tool parser first.
Ollama not detectedRun ollama serve and check curl http://localhost:11434/api/tags. Set OLLAMA_API_KEY="ollama-local" for discovery.
First reply times outThe model is loading. Raise the provider's timeoutSeconds to 300 and set keep_alive.
WSL2 keeps rebootingThe Ollama systemd service reloads a GPU model at boot and pins memory. Stop it from autostarting on WSL2 setups.
Context errors or out of memoryLower the model's contextTokens, cap maxTokens, or choose a smaller model.
"content expected a string"Strict server. Add compat.requiresStringContent: true to that model entry.
ECONNRESET mid-replyThe model server was likely killed for memory. Check its log at that timestamp, then use a smaller model or context.
LM Studio "hangs"The model was unloaded. Reload it, and turn on JIT loading in the desktop app.
More fixes: troubleshooting guide.
Local model questions
Can OpenClaw run completely offline with a local model?
Yes. With a local backend such as llama.cpp, Ollama, LM Studio or vLLM, model requests stay on your machine or network. Channels like WhatsApp or Telegram still need internet to deliver messages.
How much RAM do I need for OpenClaw with a local model?
OpenClaw's managed llama.cpp recommendations start at 8 GiB for Qwen3.5 4B and 16 GiB for Qwen3.5 9B. Gemma 4 12B needs 24 GiB with GPU acceleration, and the 27B to 30B models need 32 GiB with a GPU. These are minimums, not guarantees.
What is the easiest way to set up a local model?
Install the llama.cpp plugin with openclaw plugins install @openclaw/llama-cpp-provider, then run openclaw onboard and choose Managed local server. OpenClaw recommends a model for your hardware, downloads it and verifies a real tool call.
Are local models as good as Claude or GPT for OpenClaw?
Usually not for long, multi-step, tool-heavy tasks, especially on smaller models. They're private and free to run. Many people use a hybrid setup with a hosted model as backup, or the other way round.
Why does Ollama print tool calls as text?
You're probably using the OpenAI-compatible /v1 URL. OpenClaw needs Ollama's native API, so set the base URL to http://127.0.0.1:11434 without /v1.
Is a local model safer?
It's more private, but it has no provider-side safety filters and smaller models are easier to mislead. Keep tool permissions narrow and exec approvals on.