Skip to Content
Model providersYour own servers

Your own servers

Besides OpenRouter, you can add your own providers: a model server on your computer or your network — LM Studio, Ollama, Unsloth Studio, llama.cpp, vLLM — or any remote API that speaks the OpenAI or the Anthropic format. Agents then run on a model from that server’s own list.

1. Add the server

Settings → Providers → Your providers → Add provider

Pick the Server you run: LM Studio, Ollama, Unsloth Studio, llama.cpp and vLLM fill in their usual address and port. Other or remote is for anything else.

Check the address

Name, Protocol (http or https), Host (localhost, an IP address or api.example.com — no http:// and no port), Port. Path only if the server serves its API under a sub-path such as /api. API key only if the server asks for one.

Test connection

NeuroSquad reads the server’s model list and finds out by itself which APIs it speaks, then shows the result: Speaks: OpenAI API / Anthropic API, Works with: the agent CLIs that can use it, and how many models it has. Embedding models are hidden.

Save

A provider can only be saved after a successful test. Change a field and you test again.

You don’t choose the API — the test does. A server with no model loaded yet can still be saved: load a model, then Refresh the model list.

The key, if you give one, is encrypted by your operating system and handed only to the agents that run on this provider — never written into their command line or files. A provider without a key never falls back to your own Claude or OpenAI login: NeuroSquad gives the CLI a placeholder key instead.

The list shows each provider as Local (this computer or your home/office network) or Remote, with its address, APIs, model count and whether a key is saved.

2. Put an agent on it

When you create an agent (New agent… in the add menu), or later from the card’s ⋯ menu → Provider and model, choose your server under Your providers and pick a Model from its list. The change takes effect the next time the agent starts. The card’s chip shows the provider and model.

Which agent CLIs can use it

Agent CLINeeds on the server
Claude CodeThe Anthropic API (/v1/messages)
Codex CLIThe OpenAI API — through NeuroSquad’s translator when the server has no Responses API (below)
Qwen CodeThe OpenAI API
Gemini CLIThe OpenAI API, through a small translator NeuroSquad runs on your computer
OpenCode, Kilo Code, Hermes Agent, Kimi Code, Pi, omp, Crush, Factory Droid, Cline CLI, aider, GitHub Copilot CLI, GooseEither API
Cursor CLI, Amp, AuggieCan’t — their models run only on their vendors’ servers

Many local servers speak both APIs on the same port, so every CLI in the table can use them. If you remove a provider, its agents run on their CLI’s own login until you pick another one; the card says so.

Codex CLI on servers without the Responses API

Codex CLI speaks only OpenAI’s Responses API, while most local servers — LM Studio, llama.cpp, vLLM, Ollama — and many remote ones offer only Chat Completions. On such a server NeuroSquad runs a small translator on your computer between Codex and the server, so Codex works there as well. You don’t set anything up: the provider just lists Codex CLI under Works with.

  • Your server’s key stays inside NeuroSquad; Codex only gets a per-card pass for the translator.
  • Token counts are the server’s own numbers, passed on exactly.
  • Codex’s built-in web search is not available this way, and images in tool results reach the model as a placeholder.

A server that does have the Responses API is used directly.

Context window

The agent CLIs don’t know how much context a model on your server has: Claude Code assumes 200K tokens, and OpenCode never compacts the conversation without a number. A long session would then outgrow the server’s window and end in an error.

When the server’s model list says what context window it serves a model with (llama.cpp’s n_ctx, vLLM’s max_model_len, or a context_length field, as LM Studio and others give), NeuroSquad passes it on:

  • Claude Code — the window, and output capped to 32K tokens or a quarter of the window, whichever is smaller.
  • Codex CLI — the window, with automatic compaction at 85% of it.
  • OpenCode — the window and the same output cap, so it compacts in time.

If your server doesn’t report a window, start it with the context size you want the agents to work in, or keep sessions shorter and use Compact.

Costs

Models on your own server have no list price. Their tokens are counted in Usage & costs where the agent CLI logs them, and the cost shows as no price — never as $0. A local server that sends no cache counts shows them as not reported rather than 0.

Small local models may not drive every agent CLI’s tools reliably — Claude Code, in particular, is made for Anthropic’s models. If an agent stalls or misuses its tools, try a larger model or a CLI built for open models, such as OpenCode or Qwen Code.