OpenClaw With Local Models: Setup and Trade-Offs

Running OpenClaw on a model on your own machine means no per-message bill and no prompts leaving your computer. It also means you supply the hardware, and small models struggle more with long, multi-step agent tasks. This guide helps you decide, then walks through setup with llama.cpp, Ollama, LM Studio or vLLM.

Quick answer

Easiest: openclaw plugins install @openclaw/llama-cpp-provider, then openclaw onboard and choose Managed local server. OpenClaw checks your hardware, recommends a model and tests a real tool call. Prefer Ollama? Run ollama pull gemma4, then openclaw onboard and choose Ollama → Local only.

The trade-offs

Local models vs hosted models

Neither is better everywhere. Here's what changes when the model runs on your own hardware:

Local modelHosted model (Claude, GPT, Gemini)
PrivacyPrompts and files never leave your machineSent to the provider's servers
Running costNo per-message bill, just electricityPay per token, or use a plan's limits
Upfront costRAM, GPU and disk you may need to buyWorks on a small VPS or Raspberry Pi
Agent qualitySmall models are more likely to fumble long, tool-heavy tasksStrongest on complex, multi-step work
SpeedDepends on your hardware. CPU-only can take minutes per reply.Fast and consistent
OfflineWorks without internetNeeds a connection
Safety filtersNone. You set the limits.Provider-side filters included
The middle ground

Many people use a hybrid setup: a hosted model for hard tasks with a local model as backup, or local first with a hosted safety net. The builder below writes either one.

Hardware check

What can your machine run?

Memory (RAM or unified memory)
Graphics acceleration
What matters most?
Recommendation
    See setup steps

    Based on the minimum memory in OpenClaw's managed llama.cpp recommendations. These are floors, not guarantees of fit or speed.

    Hardware

    Model sizes and memory needs

    OpenClaw's managed llama.cpp setup picks from these models based on your memory, GPU and free disk. Each uses a 65,536-token context and supports tool calls.

    Qwen3.5 4B~2.7 GB download
    8 GiB
    Qwen3.5 9B~5.7 GB download
    16 GiB
    Gemma 4 12B IT~7.1 GB · GPU needed
    24 GiB
    Muse Glimmer 30B~16.8 GB · GPU needed
    32 GiB
    Qwen3.8 27B~16.5 GB · GPU needed
    32 GiB
    • Qwen3.8 27B is the first pick when memory allows. Muse Glimmer can fit a 24 GiB NVIDIA card where Qwen3.8 doesn't.
    • CPU-only machines top out at Qwen3.5 9B, and the first reply can take several minutes.
    • Add about 0.3 GB for the default local embedding model (EmbeddingGemma), plus room for the runtime. Two NVIDIA cards aren't added together.

    Managed backends: Metal on Apple silicon, CUDA 12.4 on Windows x64 with a supported NVIDIA GPU (driver 551.78+), and CPU elsewhere, including Linux. For CUDA on Linux, run your own server and choose Existing llama-server. See also OpenClaw system requirements.

    Backends

    Pick a local model backend

    llama-cpp/…llama.cpp (managed)

    OpenClaw picks a model for your hardware, downloads and verifies it, and runs the server. Best first choice.

    ollama/…Ollama

    CLI workflow, big model library and a background service. Can mix in Ollama Cloud models.

    lmstudio/…LM Studio

    Desktop app with a model browser. Supports the Responses API, which keeps reasoning out of replies.

    vllm/…vLLM, SGLang, MLX

    High-throughput serving on your own GPU box through an OpenAI-compatible endpoint.

    customLiteLLM or a proxy

    Put any OpenAI-style /v1/chat/completions server in front of OpenClaw.

    llmman/…llmman

    Pulls models from OCI registries and can route small prompts locally, big ones to a hosted model.

    Setup

    Set up a local model

    Install the plugin, run onboarding and choose Managed local server. Setup shows the host, model, download size and backend before downloading anything. It then checks that the model can answer and use a tool before making it your default.

    Terminal
    openclaw plugins install @openclaw/llama-cpp-provider
    openclaw onboard
    # choose: Managed local server

    Each check has a 90-second limit. If it fails, your previous default model stays selected. Already run llama-server yourself? Choose Existing llama-server instead.

    Make it yours

    Hybrid setup builder

    Mix local and hosted models

    Fallbacks switch models on provider errors, one turn at a time. Copy the result into ~/.openclaw/openclaw.json and run openclaw gateway restart.

    ~/.openclaw/openclaw.json (JSON5)

    The hosted model still needs its own key or sign-in. See the Claude, OpenAI or Gemini guide.

    Verify

    Test before you trust it

    A model that answers "hello" may still fail a real agent turn. Test in this order:

    1. Does the model respond? No tools, no agent context:
      Terminal
      openclaw infer model run --local --model <provider/model> --prompt "Reply with exactly: pong" --json
    2. Does the gateway route to it?
      Terminal
      openclaw infer model run --gateway --model <provider/model> --prompt "Reply with exactly: pong" --json
    3. Do real tasks work? Give it an actual job from your first workflow. If tool calls come out malformed, check the context size and server logs before anything else.
    Tool Search is on automatically

    For Ollama, LM Studio and managed local servers, OpenClaw loads tool descriptions only when needed, so small models aren't swamped. If a model still struggles, lean mode removes optional tools such as browser, image, video and speech:

    JSON5 (troubleshooting only)
    {
      agents: {
        defaults: {
          experimental: { localModelLean: true },
        },
      },
    }
    Tuning

    Speed and context tuning

    Context size

    Ollama setup uses a 32,768-token context, or less if the model's window is smaller. Bigger isn't free: it costs memory and slows the first token.

    Cold starts

    Large models load slowly. Raise the provider's timeoutSeconds and keep the model loaded with keep_alive.

    Vision

    Set input: ["text", "image"] on a local vision model so image attachments reach it.

    JSON5: slow first load on Ollama
    {
      models: {
        providers: {
          ollama: {
            timeoutSeconds: 300,
            models: [
              {
                id: "gemma4:26b",
                name: "gemma4:26b",
                params: { keep_alive: "15m" },
              },
            ],
          },
        },
      },
    }
    Safety

    Local models have no safety filter

    Hosted providers screen requests. A local model doesn't, and smaller models are easier to trick with prompt injection hidden in web pages, emails or files.

    More: OpenClaw security guide.

    Troubleshooting

    Fix common local model problems

    Tool calls appear as JSON text

    On Ollama, remove /v1 from the URL and use api: "ollama". On other servers, fix the chat template and tool parser first.

    Ollama not detected

    Run ollama serve and check curl http://localhost:11434/api/tags. Set OLLAMA_API_KEY="ollama-local" for discovery.

    First reply times out

    The model is loading. Raise the provider's timeoutSeconds to 300 and set keep_alive.

    WSL2 keeps rebooting

    The Ollama systemd service reloads a GPU model at boot and pins memory. Stop it from autostarting on WSL2 setups.

    Context errors or out of memory

    Lower the model's contextTokens, cap maxTokens, or choose a smaller model.

    "content expected a string"

    Strict server. Add compat.requiresStringContent: true to that model entry.

    ECONNRESET mid-reply

    The model server was likely killed for memory. Check its log at that timestamp, then use a smaller model or context.

    LM Studio "hangs"

    The model was unloaded. Reload it, and turn on JIT loading in the desktop app.

    More fixes: troubleshooting guide.

    FAQ

    Local model questions

    Can OpenClaw run completely offline with a local model?

    Yes. With a local backend such as llama.cpp, Ollama, LM Studio or vLLM, model requests stay on your machine or network. Channels like WhatsApp or Telegram still need internet to deliver messages.

    How much RAM do I need for OpenClaw with a local model?

    OpenClaw's managed llama.cpp recommendations start at 8 GiB for Qwen3.5 4B and 16 GiB for Qwen3.5 9B. Gemma 4 12B needs 24 GiB with GPU acceleration, and the 27B to 30B models need 32 GiB with a GPU. These are minimums, not guarantees.

    What is the easiest way to set up a local model?

    Install the llama.cpp plugin with openclaw plugins install @openclaw/llama-cpp-provider, then run openclaw onboard and choose Managed local server. OpenClaw recommends a model for your hardware, downloads it and verifies a real tool call.

    Are local models as good as Claude or GPT for OpenClaw?

    Usually not for long, multi-step, tool-heavy tasks, especially on smaller models. They're private and free to run. Many people use a hybrid setup with a hosted model as backup, or the other way round.

    Why does Ollama print tool calls as text?

    You're probably using the OpenAI-compatible /v1 URL. OpenClaw needs Ollama's native API, so set the base URL to http://127.0.0.1:11434 without /v1.

    Is a local model safer?

    It's more private, but it has no provider-side safety filters and smaller models are easier to mislead. Keep tool permissions narrow and exec approvals on.

    Related guides