EVAL Engine

Orchestrator: turning a plain request into scoped, dispatched agent work

What it is

Orchestrator is an agent-first repository: it holds no application code, only skills that coding agents load. It follows the skills.sh convention, so the same skills work in Claude Code, Codex, Droid, Pi and other agents that read that format.

The goal is to take a one-line request and turn it into work an agent can pick up. Things you can say to an agent working in this repo:

  • "Create a ticket to add a hero section to the landing page."
  • "Grill me on this plan before we write any tickets."
  • "Break this epic into ordered tickets, then send them to Multica."
  • "Check CHR-8's status and verify its PR actually shipped."

How it works

The flow chains several skills:

  1. grill-me interviews you about the plan until every open decision is resolved. grill-with-docs does the same while checking the plan against the project's domain model and updating CONTEXT.md and ADRs.
  2. plan-ceo-review runs a founder-style review before scoping. It asks whether the problem is framed right and picks a scope mode: expansion, selective expansion, hold scope or reduction.
  3. product-manager writes PRDs and tickets with verifiable acceptance criteria into a temp/ folder for you to review.
  4. multica installs and operates the Multica CLI, dispatches tickets to coding agents and monitors them. It checks each result instead of trusting a "completed" status.

grill-me and grill-with-docs are vendored from mattpocock/skills. The other three are written and maintained in this repo.

Keeping skills honest

Most of the engineering here is about making sure the skills an agent loads are the ones you meant:

  • skills-lock.json pins each skill's source and a content hash.
  • A Husky pre-commit hook and a CI workflow both fail if a hash is stale. The CI check was tested by committing a deliberate drift and reverting it.
  • bunfig.toml sets a supply-chain policy: npm packages must be at least 8 days old, post-install scripts are disabled, and the Socket scanner runs on install.
  • A script mirrors Multica agents, squads and projects into a local .multica/ folder so agents can see who they can delegate to.

AGENTS.md is the single entry point for every agent, and CLAUDE.md is a symlink to it.

Stack

TypeScript scripts on Bun, Husky, GitHub Actions, and the skills CLI for managing skills:

1npx skills list
2npx skills update