Skip to Content
Terminal CLI (nsq)Your own model servers

Your own model servers

Besides OpenRouter, an agent can run on a server of your own: a local one — llama.cpp, Ollama, LM Studio, vLLM, SGLang, Unsloth Studio — or any remote API that speaks the OpenAI API (chat completions or responses) or the Anthropic Messages API. Per agent, without touching the CLIs’ own configuration or logins.

Add a server

nsq provider add lmstudio --url http://localhost:1234 nsq provider add box --url http://192.168.1.20:8000 --ask-key # asks for the key, not shown echo "$KEY" | nsq provider add work --url https://llm.example.com/v1 --key-stdin

add tests the server before it stores anything: it lists the models (GET /v1/models) and finds which endpoints the server has — OpenAI POST /v1/chat/completions, OpenAI POST /v1/responses, Anthropic POST /v1/messages. You never pick the API: each CLI is offered the server when the endpoint it needs is there, and add prints which ones are:

added lmstudio: localhost:1234 — 12 models, endpoints: chat, responses, messages (9 ms) Claude Code yes Codex yes OpenCode yes (chat completions)

The address may be pasted with or without /v1 (or even a whole …/v1/chat/completions); a server under a sub-path keeps it (https://api.example.com/anthropic). Without a scheme, a local address gets http:// and anything else https://.

nsq provider list # address, endpoints, models, whether a key is stored nsq provider models lmstudio [qwen] # the server's models now (and the context window it serves) nsq provider test lmstudio # test again: new models, an updated server nsq provider test lmstudio --ask-key # …with a new key (--clear-key forgets it) nsq provider remove lmstudio

Run an agent on it

nsq run claude --provider lmstudio --model qwen/qwen3-coder-30b nsq run codex --provider ollama --model qwen3-coder:30b nsq run opencode --provider vllm --model Qwen/Qwen3-Coder-30B-A3B-Instruct nsq set api-fix --provider lmstudio --model qwen/qwen3-coder-30b # move an existing agent nsq set api-fix --provider none --model none # back to the CLI's own login

The model is an id from the server’s own list. A server with one model needs no --model. Moving an agent restarts it on the same session once it is idle, so the conversation is kept.

In the dashboard: c (New agent) → Provider — ← → cycles through own login, OpenRouter and each of your servers that fits the chosen CLI (the ones that do not are listed with the reason) — and F2 lists that server’s models. m on an agent lists its server’s models. P opens the Providers screen: a adds a server (F2 fills in the default addresses below, the key field is masked), t tests one again, Enter shows its models, x removes it.

Which CLI uses which API

CLIUses
Claude CodeAnthropic Messages (/v1/messages) only — a server without it is not offered
CodexOpenAI Responses (/v1/responses); on a server with chat completions only, nsq runs a local gateway that translates Responses ⇄ chat completions (the key stays with nsq, Codex never sees it)
OpenCodeOpenAI chat completions (@ai-sdk/openai-compatible), or the Anthropic Messages API when the server has no chat endpoint (@ai-sdk/anthropic)

A server with only /v1/responses is accepted too; only Codex is offered it.

Servers

What each server serves today, from its own documentation (checked October 2026) — the connection test decides for your version:

ServerDefault addressChat completionsResponsesAnthropic messagesContext window in its model listKey by default
llama.cpp (llama-server)http://127.0.0.1:8080 (:9931 from build b11521, October 2026)yesyesyesmeta.n_ctxno (--api-key)
Ollamahttp://localhost:11434yesyes (0.13.3+)yes (0.14.0+)not listedno
LM Studiohttp://localhost:1234yesyes (0.3.29+)yes (0.4.1+)not listedno (optional API tokens, 0.4.0+)
vLLM (vllm serve)http://localhost:8000yesyes (0.10.0+)yes (0.11.1+)max_model_lenno (--api-key)
SGLanghttp://localhost:30000yesyesyes (0.5.9+)max_model_lenno (--api-key)
Unsloth Studiohttp://localhost:8888yesyesyescontext_lengthyes (sk-unsloth-…)

So with a current version of any of them, all three CLIs run directly. Tested end to end with the real CLIs against llama.cpp (llama-server b11524): Claude Code on /v1/messages, Codex on /v1/responses and through the gateway on /v1/chat/completions, OpenCode on /v1/chat/completions. llama.cpp needs --jinja for tool calls.

Context window. When the server’s model list says how large a context it serves the model with, each CLI is told — Claude Code (CLAUDE_CODE_MAX_CONTEXT_TOKENS), Codex (model_context_window, compacting at 85 %), OpenCode (the model’s limit.context) — so a long session compacts before the server’s window is full instead of failing. Agentic CLIs send long prompts (Claude Code’s alone is well over 10 000 tokens): start the server with a large enough context (llama-server -c 32768, Ollama’s OLLAMA_CONTEXT_LENGTH, LM Studio’s context length setting).

What happens to the key

  • The key is optional (local servers need none). It is asked for without echo (--ask-key) or read from stdin (--key-stdin) — never a command-line argument, which every process on the machine can read; --key <value> is refused.
  • It is kept in the OS keyring (Windows Credential Manager, the macOS Keychain, the Secret Service on Linux). Without a keyring, set NSQ_PROVIDER_KEY_<NAME> (e.g. NSQ_PROVIDER_KEY_LMSTUDIO) where the daemon starts instead.
  • An agent gets it only through its environment at launch — never in argv, logs or nsq’s files (providers.json holds the address, the endpoints and the model list, no secret). Codex on the gateway gets a per-agent credential of the gateway instead of the key.
  • Plain http:// is accepted only for this machine and the local network (with a warning on the network: the prompts, your code and the key travel unencrypted). Anything else needs https://. Redirects are never followed — the key would go wherever they point.
  • No OpenRouter attribution header is ever sent to your servers.
  • If the server is removed, or no longer has the endpoint the CLI needs, the agent does not start — it never falls back to the CLI’s own login (which would send your code to the cloud).
  • Claude Code: the server’s address and model are also set in the agent’s own settings layer, and CLAUDE_CODE_USE_BEDROCK / _VERTEX / _FOUNDRY are switched off — so an ANTHROPIC_BASE_URL, an ANTHROPIC_AUTH_TOKEN or a cloud backend in your environment or ~/.claude/settings.json does not redirect the agent. Other configuration of the CLIs themselves (a Codex profile, OpenCode plugins) is yours and is not inspected.

Cost

nsq does not know what your server charges: requests to it show no price in nsq cost and the dashboard, never $0 — even when the model is named like a priced one. That includes the agent’s earlier requests from before it moved to the server (nsq cannot tell which server answered them).