# Your own servers

> Run agents on models from your own server — LM Studio, Ollama, Unsloth Studio, llama.cpp, vLLM — or any remote API that speaks the OpenAI or Anthropic format.

Source: https://docs.neurosquad.ai/en/providers/your-servers

Besides [OpenRouter](https://docs.neurosquad.ai/en/providers/openrouter), you can add **your own providers**: a model server on
your computer or your network — LM Studio, Ollama, Unsloth Studio, llama.cpp, vLLM — or any remote
API that speaks the OpenAI or the Anthropic format. Agents then run on a model from that server's own
list.

## 1. Add the server

**1. Settings → Providers → Your providers → Add provider**

Pick the **Server** you run: LM Studio, Ollama, Unsloth Studio, llama.cpp and vLLM fill in their
usual address and port. **Other or remote** is for anything else.

**2. Check the address**

**Name**, **Protocol** (http or https), **Host** (`localhost`, an IP address or
`api.example.com` — no `http://` and no port), **Port**. **Path** only if the server serves its
API under a sub-path such as `/api`. **API key** only if the server asks for one.

**3. Test connection**

NeuroSquad reads the server's model list and finds out by itself which APIs it speaks, then shows
the result: **Speaks: OpenAI API / Anthropic API**, **Works with:** the agent CLIs that can use
it, and how many models it has. Embedding models are hidden.

**4. Save**

A provider can only be saved after a successful test. Change a field and you test again.

You don't choose the API — the test does. A server with no model loaded yet can still be saved:
load a model, then **Refresh the model list**.

The key, if you give one, is encrypted by your operating system and handed only to the agents that
run on this provider — never written into their command line or files. A provider without a key never
falls back to your own Claude or OpenAI login: NeuroSquad gives the CLI a placeholder key instead.

The list shows each provider as **Local** (this computer or your home/office network) or **Remote**,
with its address, APIs, model count and whether a key is saved.

## 2. Put an agent on it

When you create an agent (**New agent…** in the add menu), or later from the card's ⋯ menu →
**Provider and model**, choose your server under **Your providers** and pick a **Model** from its
list. The change takes effect the next time the agent starts. The card's chip shows the provider and
model.

## Which agent CLIs can use it

| Agent CLI | Needs on the server |
| --- | --- |
| Claude Code | The Anthropic API (`/v1/messages`) |
| Codex CLI | The OpenAI API — through NeuroSquad's translator when the server has no Responses API ([below](#codex)) |
| Qwen Code | The OpenAI API |
| Gemini CLI | The OpenAI API, through a small translator NeuroSquad runs on your computer |
| OpenCode, Kilo Code, Hermes Agent, Kimi Code, Pi, omp, Crush, Factory Droid, Cline CLI, aider, GitHub Copilot CLI, Goose | Either API |
| Cursor CLI, Amp, Auggie | Can't — their models run only on their vendors' servers |

Many local servers speak both APIs on the same port, so every CLI in the table can use them. If you
remove a provider, its agents run on their CLI's own login until you pick another one; the card says
so.

## Codex CLI on servers without the Responses API

Codex CLI speaks only OpenAI's Responses API, while most local servers — LM Studio, llama.cpp, vLLM,
Ollama — and many remote ones offer only Chat Completions. On such a server NeuroSquad runs a small
translator on your computer between Codex and the server, so Codex works there as well. You don't set
anything up: the provider just lists Codex CLI under **Works with**.

- Your server's key stays inside NeuroSquad; Codex only gets a per-card pass for the translator.
- Token counts are the server's own numbers, passed on exactly.
- Codex's built-in web search is not available this way, and images in tool results reach the model
as a placeholder.

A server that does have the Responses API is used directly.

## Context window

The agent CLIs don't know how much context a model on your server has: Claude Code assumes 200K
tokens, and OpenCode never compacts the conversation without a number. A long session would then
outgrow the server's window and end in an error.

When the server's model list says what context window it serves a model with (llama.cpp's `n_ctx`,
vLLM's `max_model_len`, or a `context_length` field, as LM Studio and others give), NeuroSquad passes
it on:

- **Claude Code** — the window, and output capped to 32K tokens or a quarter of the window,
whichever is smaller.
- **Codex CLI** — the window, with automatic compaction at 85% of it.
- **OpenCode** — the window and the same output cap, so it compacts in time.

If your server doesn't report a window, start it with the context size you want the agents to work in,
or keep sessions shorter and use **Compact**.

## Costs

Models on your own server have no list price. Their tokens are counted in [Usage & costs](https://docs.neurosquad.ai/en/usage)
where the agent CLI logs them, and the cost shows as **no price** — never as $0. A local server that
sends no cache counts shows them as **not reported** rather than 0.

> Small local models may not drive every agent CLI's tools reliably — Claude Code, in particular, is
> made for Anthropic's models. If an agent stalls or misuses its tools, try a larger model or a CLI
> built for open models, such as OpenCode or Qwen Code.
