PenguinHarness
PenguinHarness is an open-source (Apache-2.0), local-first agent platform from the PrismShadow
AI team: a server with a web app, and the penguin chat command on top of it. In NeuroSquad each
Penguin card runs the real penguin chat in its terminal.
Measured on PenguinHarness 0.2.13 and on its main branch.
Install
npm install -g @prismshadow/penguin-cliIt needs Node 24 or newer. Penguin’s own install script from penguin.ooo works too. Settings → Setup in NeuroSquad can run the npm install for you.
Set up your providers and models in Penguin as usual — cards use them. A card can also run on OpenRouter or on your own server.
Add an agent card and choose PenguinHarness.
How the cards work
NeuroSquad runs its own Penguin server for its own data folder, starts it with the first Penguin
card and stops it when you quit. Each card is a Penguin agent of its own (ns_<card>) with its
own hooks, MCP servers and instructions. Your Penguin project config — your models and keys — is
copied in on every launch; your own ~/.penguin is never written.
What works
- Exact status. The card follows the server’s status stream for its session — working and
finished — and a NeuroSquad hook decides which calls need you. NeuroSquad’s own tools,
read_fileand subagents run unasked; anything else waits on Penguin’s own? Approve this tool call? [Y/n], and the toast names it (PenguinHarness needs your permission: write_file: a.txt). - Arrows. NeuroSquad’s MCP server is in the card’s agent. Penguin takes a snapshot of its tools, so every NeuroSquad tool is offered from the start and the arrows decide at call time; a skill or an installed MCP server needs the card’s next start. Installed servers ask before every call.
- Canvas mode. The hook refuses
exec_commandandinput_command— Penguin’s shell, and its only way to the web — even in dangerous mode. See Canvas mode. - Dangerous mode without a restart. While the switch is on, the hook allows every call. See Dangerous mode.
- Stop. Stop, the budget brake and the phone’s Stop use Penguin’s own abort API; the chat stays open.
- Sessions. After a restart the card continues the same conversation (
--resume); a deleted session starts fresh. - Long sessions. Context meter, Compact (
/compact), the rate-limit chip, the prompt queue, the journal and hand-off. - Usage. Every request is read from Penguin’s Trace files, subagents marked, in Usage & costs, Run Stats and Agent Pulse.
- Models. A lead agent can change the model with
agent_set_model, or move the conversation to Claude Code and back. OpenRouter and your own servers go through NeuroSquad’s local relay — no key on disk.
Good to know
- Conversations live in NeuroSquad’s Penguin folder. Your own
penguin chat --resumedoesn’t list them; withPENGUIN_HOMEpointing at that folder and--agent-id ns_<card>, it does. - Models added inside a card last only until its next launch — add them in your own Penguin.
- Plugins. Caveman, Memory, Context7 and Graphify add text to each turn through a hook Penguin runs on every prompt only after 0.2.13; on 0.2.13 their tools still work by arrow. The token saver can’t work: Penguin’s hook can’t rewrite a command.
- Usage columns. Penguin folds cache creation into the input and reasoning into the output, so those columns read “not reported”; on Claude models the price is a little low for it.
- The card’s MCP token sits in its agent’s config (Penguin can’t read it from the environment there). It is loopback-only, per card, and changes every app run.
PenguinHarness cards run on this computer only — not in WSL or SSH workspaces.