cujo
Skip to content
Contents

Self-hosting

The services, the environment, your own GitHub App, and the two settings that fail confusingly.

What you are deploying

One Compose project, so the services share a network. Nothing here is tied to a particular cloud: any container host and any reverse proxy that terminates TLS will do.

ServiceRolePublished?
harnessThe agent harness: sessions, turns, the event log, the subagents and the approval gate, on the pi coding agent SDK. Its state is one SQLite file on a volume.No — internal only. There is no console.
cujoThe service. The only thing GitHub touches, and the harness’s only client.One hostname, for the webhook and the Discord endpoint.
webThe board. Holds no secret and no state.One hostname.
github-mcpThe MCP server the agent calls to post a review. Holds the App’s private key.No — internal only.
sandbox-mcpThe MCP server that provisions the sandbox: a container on the host with its egress filtered by a gateway outside it. The one service that holds the Docker socket.No — internal only.
Cujo itself has no application-level login. The webhook host carries two signature-gated routes; the board’s read API answers only on the internal Compose name, where anything outside /public is 404 rather than 401. /healthz and /readyz are ungated operational endpoints. The harness and both MCP servers answer to nobody outside the Compose network, and have no authentication of their own.

Your own GitHub App

A self-hosted instance does not install this board’s App — it needs its own, so that the private key is yours.

  • Permissions: Contents read, Metadata read, Pull requests write, Checks write, Issues read. Checks write is the cujo/guard check run, the merge lock; every installation approves it once. Issues read is for event delivery only; GitHub releases issue_comment on it and on nothing else, and the settings page will not offer the event until the permission is set.
  • Events: pull_request, issue_comment, pull_request_review_comment, repository.
  • Webhook URL: POST /webhook on the webhook hostname, with a secret you also set as GITHUB_WEBHOOK_SECRET.

Configuration

Required. The services exit at start without these.

GITHUB_APP_ID
GITHUB_APP_PRIVATE_KEY     # the PEM text; literal \n is accepted
GITHUB_WEBHOOK_SECRET
CUJO_MODEL                 # <provider name>/<model name>

The model provider, which the service registers on the harness at boot. There is nowhere else to configure one. The three limits apply to every model in the list: the window the harness compacts against, the output cap, and whether a reasoning effort means anything to the model.

MODEL_PROVIDER_NAME
MODEL_PROVIDER_BASE_URL             # an OpenAI-compatible chat-completions endpoint
MODEL_PROVIDER_API_KEY
MODEL_PROVIDER_MODELS               # <model name>=<provider model id>, comma separated
MODEL_PROVIDER_CONTEXT_WINDOW       # default 128000
MODEL_PROVIDER_MAX_TOKENS           # default 16384
MODEL_PROVIDER_REASONING            # 0 for a model that takes no reasoning effort

Any OpenAI-compatible provider works — the model is chosen at deploy time and nothing in the code names one. Discord is configured with DISCORD_BOT_TOKEN and DISCORD_PUBLIC_KEY; leave them unset and the service runs and notifies nobody. The rest — turn timeout, conversation limits, the public stream cap, the pull request reaction switch — have working defaults.

The diff review has its own five, all optional. The model must be one of MODEL_PROVIDER_MODELS; a cheap one is the point. The budget is billed tokens, with the context counted again on every message, and a run that passes it ends in error with the count on its record.

CUJO_REVIEW_MODE          # sandbox (default) or diff, when a repo declares none
CUJO_DIFF_MODEL           # <provider name>/<model name>; unset means CUJO_MODEL
CUJO_DIFF_BUDGET_TOKENS   # billed tokens per diff run, default 400000
CUJO_DIFF_TIMEOUT_MS      # default 600000
CUJO_DIFF_BYTES           # bytes of diff the model is handed, default 60000

Two settings that fail confusingly

  • CUJO_MODEL_REASONING_EFFORT is clamped to what the model can take: a model with MODEL_PROVIDER_REASONING=0 gets none, and xhigh or max fold to high unless the model maps them. Nothing refuses to start over it; check the harness log for what was actually sent if a review reasons less than you asked.
  • A sampling key your provider rejects — CUJO_MODEL_TEMPERATURE on a model that only accepts one value, say — produces exactly that failure: readiness stays green and every review fails. Both are unset by default, and unset means the key is not sent at all rather than sent with a default. Check what your provider accepts before setting either.

Trying it locally

cp .env.example .env
make up-local

That publishes the board on 3000, the harness on 8790, the service on 8080 and the MCP servers on 8081 and 8082, all on loopback. The service dispatches on Host, which is what every production request carries too:

curl -s -H 'Host: cujo' http://localhost:8080/public/runs

Point the App’s webhook at your machine with any HTTP tunnel, with the Host set to the webhook hostname, then open a pull request on a repository the App is installed on. The deployment uses the base Compose file alone; the local overlay is only for this.

The full component and deployment reference lives in the repository, in docs/architecture.md, and every contract the code follows is in docs/spec.md. Start with the sandbox boundary if you are going to change anything that moves data around.