Features & comparisons

Captain picks the best-fit model(s) for the task you prompted.

Remove the hassle. Leverage every subscription to the max, and put the strongest open-source models into the same workflow. Save tokens. Save headaches. Captain routes every prompt across your Claude Code, Codex, Cursor, Grok, API and open-weight models, automatically or with explicit /frontier /team /repeat /workflow /cheap /oss and named-worker chains.

Architecture · How Captain works

Four layers.
Local orchestration above all your models.

Captain is not another proprietary LLM. It is a local coordinator that unifies the coding subscriptions, CLIs, and API providers you already pay for into one coherent system.

Layer 01 · TUI & prompt interface Your terminal · no SaaS bridge

Single entry point & direct controls

Natural prompts + first-class steering. Multi-turn, terminal tools and slash commands live in your local TUI.

  • /frontier max effort
  • /team multi-worker synthesis
  • /cheap cost bias
  • /workflow staged pipelines
  • > sequential chain
  • + parallel review
  • /btw mid-turn steering
  • /rename session title
Layer 02 · Local brain & director engine Sub-second triage + value routing

Decide, route & coordinate

jev + heuristics classify in ms. Ranks by quality bar, cost, latency, cooldowns.

  • jev decision leg ~300ms typed triage
  • Value function quality × speed / cost
  • Quota monitoring cooldown tracking
  • Workflow compiler stage execution & gate checks
  • Shadow recorder calibration ledger
Layer 03 · Unified worker legs All your subscriptions & keys

Connected AI coding agents & models

Runs across local CLIs, subs and remote endpoints in their native envs with tools.

  • Vendor CLIs Claude Code · Codex CLI · Cursor
  • Subscriptions SuperGrok · ChatGPT / OpenCode
  • High-speed API Kimi · MiniMax · GLM · Gemini
  • Open weights Hugging Face (Step, DeepSeek V4)
  • Zero-cost Free / Zen fallback rungs
Layer 04 · Local substrate, memory & ledger Local Euclid & resilient recovery

Context compaction, recovery & project memory

Handoffs across rate limits. Prunes churn before summarize, captures structured briefs, records local ledger.

  • Quota handoff seamless switch without context loss
  • Context compaction prune duplicated tool churn
  • Local ledger quality & cost history
  • Euclid memory registers & git provenance
  • Local orchestration providers still process the prompts you send

Read the full architecture spec →

Match the intelligence to the job

A changelog and a deadlock need different kinds of help.

Routine → economical. Hard debugging → deeper or team. Captain chooses for the task.

Routine

Keep it lean.

Changelogs, summaries, small edits → direct capable worker. No director every turn.

Demanding

Deeper reasoning.

Ambiguous/complex → director for stronger worker or team. Steer with /quality / /frontier.

Worth review

Extra perspective.

Chains + test gates for impl+review. Agents earn via added checks.

What informs the pick

Quality first. Then the tradeoffs.

Routine: min quality bar then highest value (or limited explore). Hard: director uses same signals.

Quality for the work
Task complexity and domain, model priors and locally assessed outcomes. Coding and writing can favor different workers.
Performance in your setup
Observed average run duration informs the speed tradeoff. Workers with repeated provider failures rank behind reliable candidates.
Cost and subscription pressure
Estimated API spend or pressure inferred from recent rate limits. Subscription fees remain separate from per-task cost.
Availability and your intent
Connected workers, cooldowns and supported capabilities shape the options. Use /cheap, /speed, /quality or name a worker to steer.

These are routing estimates. They do not guarantee correctness, the fastest completion or a fixed saving. If no routine candidate clears the bar, a configured fallback ladder may be used. Director calls, tests, reviews and retries count toward the work too. Read the routing mechanics →

02 · Recover and compose

Keep moving. Draft, check, challenge.

Records resets, hands off on limits (keep partials). > seq, + parallel, gate: checks; ends in director synth. /wf previews; /repeat bounds. Payoff: no manual restart juggling, repeatable reviews.

Text summaries only (no hidden state). Max 4 stages/4 workers/8 slots.

Write a patch, run its tests, get two reviews

/grok fix the retry bug gate: go test ./... > /claude review for duplicate charges + /codex review for lost payments

Test gate must pass to count success. Failed check gets one repair. Reviews share workspace (inspect, don't overlap-edit).

> next receives outputs+ parallelgate: shell check4 stages / 4 per / 8 total

Retries + final director add beyond declared. > /codex is refinement. Named or /cheap ok in chains.

Read the workflow reference →
03 · Learn and carry context

Local record. Less re-briefing.

Ledger records worker/class/tokens/cost/outcome. Priors+quality steer; prune dups + reattach task on summarize; handoffs carry state.

captain runs/show/watch. Payoff: learns from you; less re-brief.

Estimates. Mechanics →

04 · Local control + shield

MIT. Secrets masked by default.

Go brain local. Add via config, logs, OpenAI-comp endpoint, doctor readiness. Shield masks creds at boundary/egress ([[secret:…]]); counts in sidebar. Env usable in shell.

Payoff: your accounts, your control; fewer leaks.

No Captain fee. macOS/Linux. Provider charges apply. Security / SECRETS.

Who it is for

Best fit when you already run more than one agent.

Captain is not another coding agent. It is the local crew chief for the agents and accounts you already pay for. The question is whether coordinating those workers is your bottleneck.

Best fit

Juggle several agents.

You already use Claude, Codex, Cursor, Grok + APIs. Waste time on who-does-next, re-brief after quota, hand-stitch reviews.

Best fit

Spend intelligence on purpose.

Routine shouldn't burn frontier. Hard work can. Inspect every pick with why, quotas, budgets.

Best fit

Need policy over pool.

/oss, /deterministic, /btw, live /captain helm. Keep tools; constrain the picker.

Not a fit

One agent is enough.

Happy in Cursor / Claude Code / OpenCode alone? Captain is for cross-runtime coordination.

Not a fit

Want managed cloud crew.

Fusion/Amp ship their own tuned product+UX. Captain crews your workers locally.

Not a fit

Need SaaS or Windows.

Single-user local creds only. Terminal targets macOS/Linux today.

Why Captain is different

Shared ideas. A different layer.

Model choice and multi-agent are common. Captain’s edge: coordinating the separate CLIs/accounts you already have — per task, with recovery, inspectable decisions, named pools — without becoming another cloud product.

Your crew, not a new agent

Claude/Codex/Cursor/Grok + OpenRouter/NIM/HF share one convo. Local CLIs keep logins. No seats or key proxy.

Per-task economics

Routine skips director for qualifying worker. Harder escalates. (Not session-locked like some routers.)

Mid-turn control

/btw steers live; /interrupt keeps partial; /captain swaps helm. Competitors often queue till end.

Policy pools

/oss open only. /deterministic ADI-green. Compose with /team, /repeat, workflows.

Inspectable

captain why / quota / budget + local ledger. Director only picks runnable legs.

Real gated workflows

> stages, + parallel, gate: shell checks + director synth. Not just chat.

Shield on wire

Mask secrets at tool+egress. Sidebar counts. Never paste creds.

Reviewed 18 September 2026 against official sources and Captain’s local implementation (incl. /btw, /interrupt, /oss, /deterministic, live helm, task MCP, why/quota/budget). Capability comparison, not a performance benchmark. No percentage-savings claim. Fit is our assessment.

Scroll horizontally on smaller screens. Prefer the USP list above if you only need the short version.

At a glance: what each product optimizes
ProductOptimizesCaptain’s edge if you keep that toolChoose it alone when…
Captain Code
Local · MIT
Cross-runtime crew: route, recover, chain, inspect.n/aYou already run several agents and want one terminal that crews them.
Devin FusionTuned frontier lead + sidekick with persistent contexts and published evals.Your other accounts still need a crew chief; Fusion is one possible future worker, not a shipped adapter.You want Cognition’s managed lead/sidekick product and measured cache loop.
AmpIntegrated agent modes, specialists, local/cloud continuation.Coordinate Amp-adjacent work with Claude/Codex/Cursor accounts you already have (no Amp adapter today).You want one vendor’s agent UX end to end.
CursorEditor-native routing, parallel agents, review.Built-in: /cursor is a worker; combine it with Claude, Codex, Grok and API legs in one workflow.The IDE is the product; Teams/Enterprise Router is enough.
OmnigentBroad meta-harness: policies, collaboration, sandboxes; session Smart Routing.Per-task value ranking, observed-limit recovery, compact gated workflows, ADI/oss pools.You need governance and collaboration breadth more than terse coding workflows.
OpenCodeOpen agent runtime: providers, agents, TUI.Foundation: Captain adds native CLI orchestration, value ranking, recovery and the ledger on top.One agent runtime with provider choice is enough.
Claude Code RouterRequest gateway: rules, fallbacks, diagnostics.Whole-worker scheduling, gates and multi-stage synthesis above the gateway.You only need to route API requests, not crew CLI workers.
SupersetVisual multi-agent workspace, worktrees, quota meters.Automatic per-task pick + model chains; run Captain inside a worktree manually.You want visual supervision and stronger quota UI first.
AiderFocused architect/editor editing loop.Broader stages, parallel reviews and cross-account recovery.Two roles and a tight edit loop are enough.
Pi / JidoExtensible harness / Elixir agent framework.Task MCP: captain task mcp + captain host … for bounded delegation.You are building the harness or app platform itself.

Captain and Fusion optimize at different layers.

Fusion: tuned lead + sidekick with persistent contexts + published evals. Captain crews your Claude/Codex/Cursor/Grok/API workers (mid-turn /btw, /oss pools, why, recovery, gated workflows). No Fusion adapter today.

Fusion release →

Harness integration

Keep your harness. Let Captain crew the task.

MCP connects tools to an agent. Captain ships captain task mcp for bounded task delegation (plan, submit, inspect, events, artifacts, cancel, resume). The same contract is on the local HTTP /v1/task API. Routing and chat also use the OpenAI-compatible brain endpoint.

Reviewed against local MCP + API impl (incl. /btw, ADI, helm). Use captain host … to certify hosts.

Point a compatible MCP client at captain task mcp, or call /v1/task directly. See the CLI cheat sheet for the tool list and token auth note.

Decision leg

Decide in milliseconds. Spend a model call only on the work.

Frontier prose decisions cost seconds+tokens on every turn. Captain’s jev (TypeSafe) answers typed triage (class/domain), choices, scores, yes/no with calibrated prob in ~100-500ms at near-zero cost. Never a worker, never generates text or runs tools.

Triage (acted on)

Class+domain asked first. Over confidence bar → routes the turn. Under → falls back to free classifier. Missing key only costs speed.

Shadow (not acted on)

Turn shape and /btw target asked in parallel, recorded for calibration. captain jev shadow inspects agreement vs actual outcomes.

Inspect

captain jev, jev classify, jev ask for probes and custom typed questions.

The cheap decisions stop paying frontier prices; confidence makes safe gating possible.

Estimate only. Costed to the turn, shown in why. TYPESAFE_API_KEY enables; CAPTAIN_TRIAGE_JEV=0 / CAPTAIN_JEV_SHADOW=0 disable. Details →

Supported workers

Claude, Codex, Cursor, Grok and the models you add.

Local CLIs use your existing logins. Remote via OpenCode (OpenRouter, NIM, HF, xAI, Zen). Add more without rebuild; /oss includes HF legs.

GroupLegsHow you connect
Local CLIs/claude, /codex-cli, /cursorVendor CLI login · claude -p, codex exec, cursor-agent -p
xAI / OpenAI subs/grok, /grok-max, /codexSuperGrok or ChatGPT OAuth through OpenCode
OpenRouter/glm, /deepseek, /gemini, /qwenAPI key · add more with captain legs add … openrouter/…
OpenRouter (cont.)/minimax, /gpt-ossAPI key · /gpt-oss is the /deterministic pick
NVIDIA NIM/kimiNVIDIA API key · add more with captain legs add … nim/…
Hugging Face/step, /ds4-flashHF_TOKEN · add more with captain legs add … huggingface/…
OpenCode Zen/freeRotating zero-cost roster
Decision legjevTypeSafe key · answers questions, never takes a task

Register a hosted model

captain legs add muse openrouter/meta/muse-spark-1.3 --prior 7.9 captain legs add nemo nim/nvidia/nemotron-mini --prior 6.5 captain doctor

OpenRouter/NIM/HF auto-fill price+context. HF joins /oss. Restart brain after add; force once or let auto.

Adding a leg →

Every leg with its model, runtime, price, context window and role: the full roster →

Pick the level of control you need.

Bare prompts route automatically. These controls apply when you want to steer the work yourself.

ControlUse it forExample
/frontierClaude at maximum effort; fallback depends on availability./frontier investigate the deadlock
/teamA director-planned ensemble and one synthesized answer. You can require a member./team /frontier review this design
/cheap, /saveA saving preference for this turn. Not a dollar cap. Use CAPTAIN_MAX_ATTEMPTS / CAPTAIN_MAX_COST (+ optional CAPTAIN_STRICT) and captain budget for limits./cheap update the changelog
/speed, /qualityPrioritize speed or quality in routing./quality review the migration
/repeatRepeated rounds; N bounds the loop. Check status, stop after a round, or abort./repeat 5 fix the next failing test
/workflow, /wfCompile English to a preview; run with the returned workflow ID./wf grok drafts, claude reviews
/wf save, /wf runKeep a named workflow and run it again./wf run release-check
/parallelA separate task alongside the conversation./parallel review the release notes
/claude, /codex, /cursor, /grokChoose a configured worker. /codex-cli is the native Codex CLI; /grok-max is frontier xAI./grok draft the migration
/captainSee the in-app command guide; switch director at runtime with /captain auto|quality|frontier|<leg>./captain
/btw <note>Steer the worker that is already running (claude and opencode legs fold it mid-turn at the next tool boundary; codex-cli/cursor get it as next turn). Bypasses the queue for supported transports./btw remember the budget cap here
/ossOpen-weight models only. Composes with /repeat, /team, etc. in either order./oss /repeat 5 fix the next failing test
/deterministicADI-green serving tuple only; pins OpenRouter provider prefs + temperature 0 for the run. See ADI./deterministic emit the spawn table from the level brief
/interruptStop the in-flight worker and keep its work (handoff note on claude/opencode; marked partial on codex-cli/cursor). Never reroutes./interrupt switch approach

/btw and /interrupt are out-of-band for supported legs. Esc or captain stop from shell. See install guide for alias.

Built on

Credit where the stack came from.

Go brain + two plugins. Everything else is external (own repo, license). No vendoring. Rows describe actual role.

What Captain is built out of
ProjectWhat Captain does with itLicense
sst/opencodeTerminal + transport for remote legs. captain init adds two plugins (route to brain, draw sidebar). Workers are separate opencode serve processes.MIT
sst/opentuiRenders sidebar panels (roster, shield, runs, team) and demo chrome.MIT
solidjs/solidReactive layer for panels (poll brain, update signals only).MIT
oven-sh/bunRuns the plugins; doctor reports the pinned packageManager version.MIT
golang/goBrain: router, scheduler, ledger, budgets, HTTP API. Build rev stamped for evidence.BSD-3-Clause
stretchr/testifyOnly direct Go dep. Tests gate every change (go test ./...).MIT
go-yaml/yamlIndirect (via testify).MIT / Apache-2.0
SQLiteBest-effort sqlite3 for per-folder prompt history (skipped if absent).Public domain
Model Context Protocolcaptain task mcp + /v1/task for bounded delegation by other harnesses.MIT
The worker CLIs Captain drives
ProjectWhat Captain does with itLicense
Claude Codeclaude binary under your sign-in (never Captain-held key). Version pinned for evidence reports.Vendor CLI
openai/codexcodex binary, same contract + pinning.Apache-2.0
Cursor CLIcursor-agent binary, same contract + pinning.Vendor CLI
Routes, catalogs and data Captain reads
SourceWhat Captain does with itTerms
OpenRouterRemote legs without CLI (DeepSeek, Gemini, GLM, gpt-oss, MiniMax, Qwen) via opencode on your key.Provider terms
NVIDIA NIMServes Kimi; first-class preset in captain init + proxy.Provider terms
Hugging Face Inference ProvidersServes Step + DeepSeek V4 Flash; preset alongside NIM. Joins /oss.Provider terms
xAI APIServes Grok and Grok Max.Provider terms
sst/models.devProvider/model catalog shape used by captain init.MIT
Artificial AnalysisCold-start quality priors (coding index). Local ledger overrides.Provider terms
TypeSafe System Onejev decision leg (typed triage, calibrated confidence, no tools).Provider terms
Webster's Second International Dictionary1934 word list for folder classification (fallback).Public domain

Licenses reviewed 20 Sep 2026; see each repo. Attribution in NOTICE. Provider terms apply to remote calls.

Sources and comparison scope

Comparisons reviewed 18 Sep 2026 against official docs and the local implementation. Not exhaustive. No comparative performance or percentage-savings claim.

Gaps & next evidence

Prove the economics. Package the hosts.

Shipped: routing, /btw, ADI, budgets, mcp. Remains: public baseline evidence + drop-in host recipes.

Baseline

Blinded 12-task pilot: cost/time vs fixed frontier/econ/auto/workflow, with attribution.

Honest limits

Budgets/quotas ship. No guaranteed savings claim.

Host packaging

Turn MCP+helpers into certified Pi/Jido/editor drop-ins.

Roadmap →

Your agents. One captain.

Use the accounts you have. Give Captain the next task.