An installable Agent Skill and project scaffold for routing Orca Agent IDE work across six model-specific roles, plus a four-model idea-debate mode.
| Role | Default model | Best for |
|---|---|---|
architect |
Claude Fable 5.1 | Architecture, planning, high-risk review |
executor |
GPT-6 Astra via Codex | Implementation, debugging, verification, raster images via $imagegen |
thrifty |
Grok 4.6 | Exploration, research, small low-risk changes |
ui |
Gemini 3.7 Flash (Medium) via agy |
User-visible surface drafts, always routed back to architect for approval |
reviewer |
Claude Opus 5 | Final pre-merge gate only — APPROVE/BLOCK, never implements |
fallback |
Gemini 3.7 Flash (Medium) via agy |
Continuity after rate or session limits |
Bootstrap starts the four primaries (architect/executor/thrifty/fallback); ui and
reviewer tabs are created on their first dispatch. The idea-debate mode below adds four more
read-only debater_* seats, one per provider.
The defaults are intentionally opinionated. Launch commands live in scripts/orca-roles-lib.sh (role_meta / role_launch_cmd), not roles.yaml — edit that library if you need different model IDs or CLI flags.
- Orca Agent IDE with Settings → Experimental → Agent orchestration enabled
orca,claude,codex,grok, andagyavailable onPATH- Python 3 and Bash
Check the local runtime before bootstrapping:
.orca/orchestration/scripts/orca-status.sh # after install; exit 1 = something to fixDon't have all four CLIs? Point a role at a model you do have — see per-project overrides.
This repo is a self-contained marketplace (same pattern as Superpowers): root SKILL.md + .claude-plugin/ manifests. Layout stays at repo root so scripts/install-to-project.sh paths keep working.
/plugin marketplace add zeromountain/orca-role-orchestration
/plugin install orca-role-orchestration@orca-role-orchestration
CLI:
claude plugin marketplace add zeromountain/orca-role-orchestration
claude plugin install orca-role-orchestration@orca-role-orchestrationLocal dry-run from a checkout:
claude plugin validate .
claude --plugin-dir "$(pwd)"
# skill namespace: /orca-role-orchestration:…Refresh after new commits (plugin version is SHA-based — no pin field in plugin.json):
/plugin marketplace update orca-role-orchestration
/plugin update orca-role-orchestration@orca-role-orchestration
Claude plugin install loads the skill for Claude Code. For project scaffold and multi-agent paths (~/.agents/skills, Codex/Grok), also use install-skill.sh below (or run install-to-project.sh from the plugin cache / a clone).
Codex discovers this repo via .agents/plugins/marketplace.json and loads the plugin from .codex-plugin/plugin.json (root single skill, skills: "./", empty hooks: {} so no Claude hooks leak in).
codex plugin marketplace add zeromountain/orca-role-orchestration
codex plugin add orca-role-orchestration@orca-role-orchestrationLocal dry-run from a checkout:
codex plugin marketplace add "$(pwd)"
codex plugin add orca-role-orchestration@orca-role-orchestration
codex plugin listRefresh marketplace snapshots after new commits:
codex plugin marketplace upgradeSame note as Claude: plugin install loads the skill into Codex; project scaffold still uses install-to-project.sh from a full skill root (install-skill.sh, clone, or plugin cache).
Same command installs and updates (clone-or-pull + optional multi-agent symlinks):
# from a checkout, or curl raw from GitHub
curl -fsSL https://raw.githubusercontent.com/zeromountain/orca-role-orchestration/main/scripts/install-skill.sh | bash
# or:
./scripts/install-skill.shCanonical path: ~/.agents/skills/orca-role-orchestration. If ~/.claude/skills, ~/.codex/skills, or ~/.grok/skills exist, they get a symlink to that checkout.
Restart or reload your agent so it discovers SKILL.md.
One flagless command — safe to re-run anytime:
~/.agents/skills/orca-role-orchestration/scripts/install-to-project.sh \
--project-root "$(pwd)"| Layer | Path | On re-run |
|---|---|---|
| Managed routing | .orca/orchestration/roles.yaml |
Always refreshed (.bak if changed) |
| Your hints | .orca/orchestration/project_hints.yaml |
Created once; never overwritten |
| Personas | .orca/orchestration/personas/*.md |
Refresh if unmodified; skip if forked |
| Scripts / docs | scripts/, PLAYBOOK.md, … |
Always refreshed |
| Your overrides | .orca/orchestration/roles.local.json |
Created by you; never touched |
| Version stamp | install-manifest.json |
Written every run |
A changed managed file is backed up to .bak, and an existing .bak rotates to
.bak.1, .bak.2, … so an older fork is never destroyed.
…/install-to-project.sh --project-root "$(pwd)" --dry-run # preview, writes nothing
…/install-to-project.sh --project-root "$(pwd)" --reset # overwrite forked personas
…/install-to-project.sh --project-root "$(pwd)" --uninstall # keeps your filesOptional .orca/orchestration/roles.local.json repoints a role without forking a
script — for when you don't have one of the default CLIs:
{
"thrifty": {
"model": "claude-sonnet-5",
"launch_command": "claude --model claude-sonnet-5 --dangerously-skip-permissions"
}
}Fields: title, model, agent, launch_command. Role names stay fixed —
they are wired into routing, DAGs, personas and every command file.
Bootstrap can also run a subset: --roles architect,executor.
Then bootstrap workers (idempotent — re-run to finish a partial bootstrap):
orca repo add --path "$(pwd)" # only if the project is not already in Orca
.orca/orchestration/scripts/orca-bootstrap-roles.sh --worktree "path:$(pwd)"See SKILL.md for routing behavior and templates/PLAYBOOK.md
for the supervised lifecycle. Installation, update policy, and per-project overrides:
references/installation.md. Why each model holds its
role: references/model-roles.md.
tests/repo-lint.sh # plugin manifests, Claude/Codex command pairs, personas
tests/install.sh # installer regressions (managed / user-owned / fork policy)
tests/runtime.sh # runtime scripts against a fake `orca` — no Orca needed
shellcheck scripts/*.sh tests/*.sh tests/fake-orca/orcatests/runtime.sh puts tests/fake-orca/ first on PATH and installs the
scaffold into a tmp project, so each case also proves that what the installer
emits is runnable. tests/fake-orca/orca supports failure injection via
$FAKE_ORCA_STATE/fail/<subcommand> — that is what covers the close/reap
failure paths. CI runs all three suites plus shellcheck on macOS and Ubuntu.
Changes to managed files are recorded in CHANGELOG.md.
Once the skill is loaded, ordinary requests route themselves — the scripts below are what the coordinator runs on your behalf, not something you normally type.
| You ask for | Routed to | Shape |
|---|---|---|
| "Plan this refactor before anyone touches code" | architect |
plan only, no edits |
| "Implement the approved plan and get the build green" | executor |
implement + verify |
| "Map where auth lives" / "prototype this quickly" | thrifty |
read-only survey, cheap changes |
| "Draft the settings screen" | ui → architect |
draft, then approval before implementing |
| "Final check before merge" | reviewer |
APPROVE / BLOCK, never implements |
| "Make a hero image" | clarity gate → executor |
Codex $imagegen only |
| "Opus hit its limit — keep going" | fallback |
continuity on Gemini Flash |
| "Let's sharpen this idea" | four debater_* seats |
idea debate |
Korean triggers work the same way (역할 오케스트레이션, 모델별 역할 분리, 이미지 생성,
아이디어 토론, 니치 찾기).
Slash commands are the explicit form of the same routes — /orca-install, /orca-bootstrap,
/orca-dispatch, /orca-wait, /orca-fallback, /orca-debate, /orca-close, /orca-status
work bare in both Claude Code (v2.1.216+) and Codex; the namespaced Claude form
(/orca-role-orchestration:orca-dispatch, …) always works as a fallback.
For a fixed role chain, .orca/orchestration/or dag plan-exec-review "<goal>" (or the long form
orca-dispatch-dag.sh) wires the whole DAG in one call instead of one dispatch per step — see
Plan → implement → review below.
The standard DAG, dispatched by hand. Each tab auto-closes when its task completes — there is no close step:
D=.orca/orchestration/scripts/orca-dispatch-role.sh
# 1. Opus plans. No file edits in this pass.
"$D" architect --spec "Plan only: add refresh-token rotation to the auth service.
Constraints: follow AGENTS.md; no schema migration in this pass.
Scope: src/auth/**. Done: numbered plan + risk list, zero file edits."
# 2. Astra implements the approved plan and blocks until it reports back.
"$D" executor --wait --spec "Implement the approved plan (rotation + revoke-on-reuse).
Scope: src/auth/**, tests/auth/**. Done: pnpm typecheck && pnpm test:auth both green."
# 3. Opus gates the diff. Review only.
"$D" reviewer --spec "Pre-merge gate on the auth diff. APPROVE or BLOCK with reasons.
Do not implement or rewrite."Cheap work skips the ladder entirely — "$D" thrifty --spec "Read-only: list every call site of issueToken(). No edits."
--wait on dispatch already pins the wait to that dispatch's own task. Waiting separately means
passing the task id dispatch printed, otherwise the first matching message wins — including a
leftover one from an unrelated flow:
.orca/orchestration/scripts/orca-wait-done.sh --task task_abc123 --timeout-ms 900000A timeout or count:0 is a checkpoint, not a failure.
Raster images route to executor through Codex $imagegen. If subject, intended use, or
destination is missing, the coordinator asks before dispatching rather than inventing them:
.orca/orchestration/scripts/orca-dispatch-role.sh executor --spec "
Use Codex \$imagegen skill only
(read \${CODEX_HOME:-\$HOME/.codex}/skills/.system/imagegen/SKILL.md).
Goal: dark-mode hero image for the landing page.
Subject: a single orca breaching over a calm night sea.
Use: web hero, 16:9.
Destination: public/img/hero.png
Avoid: text, logos, brand marks.
Done: final path + mode (built-in|CLI).
"Vector icon sets, repo-native logos, and shapes better done in CSS/SVG are not $imagegen work.
Hand the remaining work to the fallback seat instead of retrying the limited primary:
.orca/orchestration/scripts/orca-fallback-on-limit.sh --from architect \
--spec "Continue: finish the rotation plan.
Done so far: threat model + token schema. Remaining: revoke-on-reuse and rollout steps."Fallback is continuity, not a default quality lane.
Four models argue an idea into a niche direction — propose, critique each other anonymously, then converge:
.orca/orchestration/scripts/orca-debate.sh --topic "your idea or open question"Transcript lands in .orca/orchestration/debates/<slug>/ (gitignored); the decision document goes
to docs/ideas/. Round output is written under a shuffled label from the start (round-1/A.md,
never a model-named file); the label map and per-round manifest live outside the debate directory
entirely, in .orca/orchestration/debate-labels/ and debate-manifests/ — nothing inside the
debate directory itself, in a filename or in file contents, ever names a debater. See the
anonymity guarantee and its limits under Security below. Only one debate runs at a time: starting
a second one — same slug or a different one — while a debate is still live is refused, since
terminals are reused per-role globally and two concurrent debates would otherwise collide in the
same four sessions (a same-slug collision would also reset the live debate's tracked handles,
leaving its tabs unprotected from the new driver's own cleanup).
Orca's documented recipes (onorca.dev/docs/recipes) are written as GUI steps. The scaffold
ships the CLI half of each as one or subcommand (.orca/orchestration/or), and leaves the
steps that have no CLI — annotating a diff, clicking in Design Mode, the Cmd-J palette,
registering an SSH host — to the human, saying so in its output:
.orca/orchestration/or race start "Fix the login bug" # Fable vs Astra vs Grok, one worktree each, same base branch
.orca/orchestration/or race status <race_id>
.orca/orchestration/or race pick <race_id> 2 # keep seat 2 (opens its diff), delete the losers' worktrees
.orca/orchestration/or race finish <race_id> # release the winner's tab; the worktree stays
.orca/orchestration/or review [--race <race_id> 2] # open the diff viewer, then j/k/c + Send to agent in Orca
.orca/orchestration/or ps # every worktree: agents, notes, who needs input
.orca/orchestration/or gc [--close] # merged worktrees — report-only until --close
.orca/orchestration/or note "reproduced" --workspace-status in-progress
.orca/orchestration/or fix http://localhost:3000/page # Design Mode loop: ui tab active, page openSlash commands: /orca-race, /orca-review, /orca-worktrees, /orca-design-fix. A race
seat is an ordinary supervised dispatch in its own worktree; its tab is retained until you
pick, because the diff viewer's Send to agent needs a live agent there. pick and
gc --close delete checkouts and branches — both scripts only ever touch worktrees they
can prove are theirs (race ledger) or merged (Orca's own guard, no --force).
The default launch commands disable or bypass agent permission checks. Use them only in trusted repositories and review the commands before running orca-bootstrap-roles.sh. Remove the bypass flags if you want each provider's normal approval boundaries.
Generated .orca/orchestration/handles.json files are local runtime state and must not be committed.
Idea-debate anonymity is not cryptographic. The achievable guarantee is: nothing instructs a
debater to deanonymize, and no single diff, glob, or file read inside the debate directory
reveals authorship. Debaters run under the same permission-bypass flags as every other role, so
this is enforced by the dispatch spec and persona (templates/roles.yaml's read_only field is
prompt-enforced, not a sandboxing fact) — a debater that ignored its instructions could still read
dispatch-ledger.jsonl, handles.json, terminal-journal.jsonl, or orca terminal list titles,
none of which live inside the debate directory but none of which are hidden from that process
either. The transcript (transcript.md, inside the debate directory) is the one deliberate
exception: it re-attributes each contribution by real short name for the human reader, written
only after the debate has concluded and never referenced by any round spec.
MIT — see LICENSE.