Keep it lean.
Changelogs, summaries, small edits → direct capable worker. No director every turn.
Features & comparisons
Remove the hassle. Leverage every subscription to the max, and put the strongest open-source models into the same workflow. Save tokens. Save headaches. Captain routes every prompt across your Claude Code, Codex, Cursor, Grok, API and open-weight models, automatically or with explicit /frontier /team /repeat /workflow /cheap /oss and named-worker chains.
Architecture · How Captain works
Captain is not another proprietary LLM. It is a local coordinator that unifies the coding subscriptions, CLIs, and API providers you already pay for into one coherent system.
Natural prompts + first-class steering. Multi-turn, terminal tools and slash commands live in your local TUI.
jev + heuristics classify in ms. Ranks by quality bar, cost, latency, cooldowns.
Runs across local CLIs, subs and remote endpoints in their native envs with tools.
Handoffs across rate limits. Prunes churn before summarize, captures structured briefs, records local ledger.
Read the full architecture spec →
Match the intelligence to the job
Routine → economical. Hard debugging → deeper or team. Captain chooses for the task.
Changelogs, summaries, small edits → direct capable worker. No director every turn.
Ambiguous/complex → director for stronger worker or team. Steer with /quality / /frontier.
Chains + test gates for impl+review. Agents earn via added checks.
What informs the pick
Routine: min quality bar then highest value (or limited explore). Hard: director uses same signals.
/cheap, /speed, /quality or name a worker to steer.These are routing estimates. They do not guarantee correctness, the fastest completion or a fixed saving. If no routine candidate clears the bar, a configured fallback ladder may be used. Director calls, tests, reviews and retries count toward the work too. Read the routing mechanics →
Records resets, hands off on limits (keep partials). > seq, + parallel, gate: checks; ends in director synth. /wf previews; /repeat bounds. Payoff: no manual restart juggling, repeatable reviews.
Text summaries only (no hidden state). Max 4 stages/4 workers/8 slots.
/grok fix the retry bug gate: go test ./... > /claude review for duplicate charges + /codex review for lost payments
Test gate must pass to count success. Failed check gets one repair. Reviews share workspace (inspect, don't overlap-edit).
Retries + final director add beyond declared. > /codex is refinement. Named or /cheap ok in chains.
Ledger records worker/class/tokens/cost/outcome. Priors+quality steer; prune dups + reattach task on summarize; handoffs carry state.
captain runs/show/watch. Payoff: learns from you; less re-brief.
Estimates. Mechanics →
Who it is for
Captain is not another coding agent. It is the local crew chief for the agents and accounts you already pay for. The question is whether coordinating those workers is your bottleneck.
You already use Claude, Codex, Cursor, Grok + APIs. Waste time on who-does-next, re-brief after quota, hand-stitch reviews.
Routine shouldn't burn frontier. Hard work can. Inspect every pick with why, quotas, budgets.
/oss, /deterministic, /btw, live /captain helm. Keep tools; constrain the picker.
Happy in Cursor / Claude Code / OpenCode alone? Captain is for cross-runtime coordination.
Fusion/Amp ship their own tuned product+UX. Captain crews your workers locally.
Single-user local creds only. Terminal targets macOS/Linux today.
Why Captain is different
Model choice and multi-agent are common. Captain’s edge: coordinating the separate CLIs/accounts you already have — per task, with recovery, inspectable decisions, named pools — without becoming another cloud product.
Claude/Codex/Cursor/Grok + OpenRouter/NIM/HF share one convo. Local CLIs keep logins. No seats or key proxy.
Routine skips director for qualifying worker. Harder escalates. (Not session-locked like some routers.)
/btw steers live; /interrupt keeps partial; /captain swaps helm. Competitors often queue till end.
/oss open only. /deterministic ADI-green. Compose with /team, /repeat, workflows.
captain why / quota / budget + local ledger. Director only picks runnable legs.
> stages, + parallel, gate: shell checks + director synth. Not just chat.
Mask secrets at tool+egress. Sidebar counts. Never paste creds.
Reviewed 18 September 2026 against official sources and Captain’s local implementation (incl. /btw, /interrupt, /oss, /deterministic, live helm, task MCP, why/quota/budget). Capability comparison, not a performance benchmark. No percentage-savings claim. Fit is our assessment.
Scroll horizontally on smaller screens. Prefer the USP list above if you only need the short version.
| Product | Optimizes | Captain’s edge if you keep that tool | Choose it alone when… |
|---|---|---|---|
| Captain Code Local · MIT | Cross-runtime crew: route, recover, chain, inspect. | n/a | You already run several agents and want one terminal that crews them. |
| Devin Fusion | Tuned frontier lead + sidekick with persistent contexts and published evals. | Your other accounts still need a crew chief; Fusion is one possible future worker, not a shipped adapter. | You want Cognition’s managed lead/sidekick product and measured cache loop. |
| Amp | Integrated agent modes, specialists, local/cloud continuation. | Coordinate Amp-adjacent work with Claude/Codex/Cursor accounts you already have (no Amp adapter today). | You want one vendor’s agent UX end to end. |
| Cursor | Editor-native routing, parallel agents, review. | Built-in: /cursor is a worker; combine it with Claude, Codex, Grok and API legs in one workflow. | The IDE is the product; Teams/Enterprise Router is enough. |
| Omnigent | Broad meta-harness: policies, collaboration, sandboxes; session Smart Routing. | Per-task value ranking, observed-limit recovery, compact gated workflows, ADI/oss pools. | You need governance and collaboration breadth more than terse coding workflows. |
| OpenCode | Open agent runtime: providers, agents, TUI. | Foundation: Captain adds native CLI orchestration, value ranking, recovery and the ledger on top. | One agent runtime with provider choice is enough. |
| Claude Code Router | Request gateway: rules, fallbacks, diagnostics. | Whole-worker scheduling, gates and multi-stage synthesis above the gateway. | You only need to route API requests, not crew CLI workers. |
| Superset | Visual multi-agent workspace, worktrees, quota meters. | Automatic per-task pick + model chains; run Captain inside a worktree manually. | You want visual supervision and stronger quota UI first. |
| Aider | Focused architect/editor editing loop. | Broader stages, parallel reviews and cross-account recovery. | Two roles and a tight edit loop are enough. |
| Pi / Jido | Extensible harness / Elixir agent framework. | Task MCP: captain task mcp + captain host … for bounded delegation. | You are building the harness or app platform itself. |
Fusion: tuned lead + sidekick with persistent contexts + published evals. Captain crews your Claude/Codex/Cursor/Grok/API workers (mid-turn /btw, /oss pools, why, recovery, gated workflows). No Fusion adapter today.
Harness integration
MCP connects tools to an agent. Captain ships captain task mcp for bounded task delegation (plan, submit, inspect, events, artifacts, cancel, resume). The same contract is on the local HTTP /v1/task API. Routing and chat also use the OpenAI-compatible brain endpoint.
Reviewed against local MCP + API impl (incl. /btw, ADI, helm). Use captain host … to certify hosts.
Point a compatible MCP client at captain task mcp, or call /v1/task directly. See the CLI cheat sheet for the tool list and token auth note.
Decision leg
Frontier prose decisions cost seconds+tokens on every turn. Captain’s jev (TypeSafe) answers typed triage (class/domain), choices, scores, yes/no with calibrated prob in ~100-500ms at near-zero cost. Never a worker, never generates text or runs tools.
Class+domain asked first. Over confidence bar → routes the turn. Under → falls back to free classifier. Missing key only costs speed.
Turn shape and /btw target asked in parallel, recorded for calibration. captain jev shadow inspects agreement vs actual outcomes.
captain jev, jev classify, jev ask for probes and custom typed questions.
The cheap decisions stop paying frontier prices; confidence makes safe gating possible.
Estimate only. Costed to the turn, shown in why. TYPESAFE_API_KEY enables; CAPTAIN_TRIAGE_JEV=0 / CAPTAIN_JEV_SHADOW=0 disable. Details →
Supported workers
Local CLIs use your existing logins. Remote via OpenCode (OpenRouter, NIM, HF, xAI, Zen). Add more without rebuild; /oss includes HF legs.
| Group | Legs | How you connect |
|---|---|---|
| Local CLIs | /claude, /codex-cli, /cursor | Vendor CLI login · claude -p, codex exec, cursor-agent -p |
| xAI / OpenAI subs | /grok, /grok-max, /codex | SuperGrok or ChatGPT OAuth through OpenCode |
| OpenRouter | /glm, /deepseek, /gemini, /qwen | API key · add more with captain legs add … openrouter/… |
| OpenRouter (cont.) | /minimax, /gpt-oss | API key · /gpt-oss is the /deterministic pick |
| NVIDIA NIM | /kimi | NVIDIA API key · add more with captain legs add … nim/… |
| Hugging Face | /step, /ds4-flash | HF_TOKEN · add more with captain legs add … huggingface/… |
| OpenCode Zen | /free | Rotating zero-cost roster |
| Decision leg | jev | TypeSafe key · answers questions, never takes a task |
captain legs add muse openrouter/meta/muse-spark-1.3 --prior 7.9
captain legs add nemo nim/nvidia/nemotron-mini --prior 6.5
captain doctor
OpenRouter/NIM/HF auto-fill price+context. HF joins /oss. Restart brain after add; force once or let auto.
Every leg with its model, runtime, price, context window and role: the full roster →
Bare prompts route automatically. These controls apply when you want to steer the work yourself.
| Control | Use it for | Example |
|---|---|---|
/frontier | Claude at maximum effort; fallback depends on availability. | /frontier investigate the deadlock |
/team | A director-planned ensemble and one synthesized answer. You can require a member. | /team /frontier review this design |
/cheap, /save | A saving preference for this turn. Not a dollar cap. Use CAPTAIN_MAX_ATTEMPTS / CAPTAIN_MAX_COST (+ optional CAPTAIN_STRICT) and captain budget for limits. | /cheap update the changelog |
/speed, /quality | Prioritize speed or quality in routing. | /quality review the migration |
/repeat | Repeated rounds; N bounds the loop. Check status, stop after a round, or abort. | /repeat 5 fix the next failing test |
/workflow, /wf | Compile English to a preview; run with the returned workflow ID. | /wf grok drafts, claude reviews |
/wf save, /wf run | Keep a named workflow and run it again. | /wf run release-check |
/parallel | A separate task alongside the conversation. | /parallel review the release notes |
/claude, /codex, /cursor, /grok… | Choose a configured worker. /codex-cli is the native Codex CLI; /grok-max is frontier xAI. | /grok draft the migration |
/captain | See the in-app command guide; switch director at runtime with /captain auto|quality|frontier|<leg>. | /captain |
/btw <note> | Steer the worker that is already running (claude and opencode legs fold it mid-turn at the next tool boundary; codex-cli/cursor get it as next turn). Bypasses the queue for supported transports. | /btw remember the budget cap here |
/oss | Open-weight models only. Composes with /repeat, /team, etc. in either order. | /oss /repeat 5 fix the next failing test |
/deterministic | ADI-green serving tuple only; pins OpenRouter provider prefs + temperature 0 for the run. See ADI. | /deterministic emit the spawn table from the level brief |
/interrupt | Stop the in-flight worker and keep its work (handoff note on claude/opencode; marked partial on codex-cli/cursor). Never reroutes. | /interrupt switch approach |
/btw and /interrupt are out-of-band for supported legs. Esc or captain stop from shell. See install guide for alias.
Built on
Go brain + two plugins. Everything else is external (own repo, license). No vendoring. Rows describe actual role.
| Project | What Captain does with it | License |
|---|---|---|
| sst/opencode | Terminal + transport for remote legs. captain init adds two plugins (route to brain, draw sidebar). Workers are separate opencode serve processes. | MIT |
| sst/opentui | Renders sidebar panels (roster, shield, runs, team) and demo chrome. | MIT |
| solidjs/solid | Reactive layer for panels (poll brain, update signals only). | MIT |
| oven-sh/bun | Runs the plugins; doctor reports the pinned packageManager version. | MIT |
| golang/go | Brain: router, scheduler, ledger, budgets, HTTP API. Build rev stamped for evidence. | BSD-3-Clause |
| stretchr/testify | Only direct Go dep. Tests gate every change (go test ./...). | MIT |
| go-yaml/yaml | Indirect (via testify). | MIT / Apache-2.0 |
| SQLite | Best-effort sqlite3 for per-folder prompt history (skipped if absent). | Public domain |
| Model Context Protocol | captain task mcp + /v1/task for bounded delegation by other harnesses. | MIT |
| Project | What Captain does with it | License |
|---|---|---|
| Claude Code | claude binary under your sign-in (never Captain-held key). Version pinned for evidence reports. | Vendor CLI |
| openai/codex | codex binary, same contract + pinning. | Apache-2.0 |
| Cursor CLI | cursor-agent binary, same contract + pinning. | Vendor CLI |
| Source | What Captain does with it | Terms |
|---|---|---|
| OpenRouter | Remote legs without CLI (DeepSeek, Gemini, GLM, gpt-oss, MiniMax, Qwen) via opencode on your key. | Provider terms |
| NVIDIA NIM | Serves Kimi; first-class preset in captain init + proxy. | Provider terms |
| Hugging Face Inference Providers | Serves Step + DeepSeek V4 Flash; preset alongside NIM. Joins /oss. | Provider terms |
| xAI API | Serves Grok and Grok Max. | Provider terms |
| sst/models.dev | Provider/model catalog shape used by captain init. | MIT |
| Artificial Analysis | Cold-start quality priors (coding index). Local ledger overrides. | Provider terms |
| TypeSafe System One | jev decision leg (typed triage, calibrated confidence, no tools). | Provider terms |
| Webster's Second International Dictionary | 1934 word list for folder classification (fallback). | Public domain |
Licenses reviewed 20 Sep 2026; see each repo. Attribution in NOTICE. Provider terms apply to remote calls.
Comparisons reviewed 18 Sep 2026 against official docs and the local implementation. Not exhaustive. No comparative performance or percentage-savings claim.
Gaps & next evidence
Shipped: routing, /btw, ADI, budgets, mcp. Remains: public baseline evidence + drop-in host recipes.
Blinded 12-task pilot: cost/time vs fixed frontier/econ/auto/workflow, with attribution.
Budgets/quotas ship. No guaranteed savings claim.
Turn MCP+helpers into certified Pi/Jido/editor drop-ins.
Use the accounts you have. Give Captain the next task.