The runner protocol

How a saggar-worker runner pairs with the service, claims queued jobs, streams logs, pushes artifact blobs and settles the build. Everything here is machine-facing, at the server root (/worker/…, /nodes). Like the registries, these paths are spoken by the saggar-worker binary and never carry the /api/v1 prefix.

Authentication levels

Three credentials ride this surface, one per relationship:

Account token

X-Auth-Token or Authorization: Bearer with an sgt-… token. Pairing and the roster reads. A registry-scoped token is refused.

Runner token

Authorization: Bearer swt-…. The token issued at pairing, matched by hash against the roster. Every authenticated call refreshes the runner’s last_seen.

Lease secret

X-Saggar-Lease: <secret>. The claim’s lease secret, the unguessable capability the claim minted, presented on every job-scoped call. An authenticated runner without it cannot touch a job it does not hold; a foreign lease answers 409. Guessing a job id buys nothing.

Pairing

POST /worker/pair, account auth:

Pair the node with the presenting account. Body:

{
  "node_id": "build-01",
  "arch": "arm64",
  "version": "0.2.0",
  "network": "online",
  "docker_format": true
}

node_id, arch and version are the runner’s identity; network (online | pinned | offline, default online) and docker_format (default true) are the owner’s declared conduct profile, what the owner states the node can do, never probed (see Pair your own runners). Omitted fields pair with the safe defaults.

The pairing is admitted inline; the account token proves the account, so the owner approves by pairing. The answer is {"status": "approved", "token": "swt-…"}, the runner’s bearer token, delivered once; the service stores only its hash. Registry tokens are refused, and attempts are capped per peer address.

The issued token persists under the runner’s data dir (0600) and is reused from there on; --token on saggar-worker (env SAGGAR_WORKER_TOKEN) is the infra-as-code shape: a pre-provisioned runner token that skips pairing entirely.

Roster

Management-auth routes (account tokens):

GET /api/v1/nodes

The runners the subject owns or is shared with: node id, arch, engine version, state (active, retired), owner handle, owned flag, declared profile (network, docker_format), last seen.

GET /api/v1/nodes/:id/shares

The accounts the runner is shared with (owner).

PUT /api/v1/nodes/:id/shares

Grant use to an account: {"handle": "…"} → the sharee list. Sharing grants use: a shared account’s jobs may schedule on the runner, eligibility re-evaluated per claim.

DELETE /api/v1/nodes/:id/shares/:handle

Revoke a share → the sharee list. Future claims stop instantly; a running job finishes its lease. Management stays with the owner.

POST /api/v1/nodes/retire

Retire a runner you own: {"node_id": …}. Its token stops authenticating immediately; the row stays listed retired for the record. Not idempotent: unknown or already-retired answers 404.

Claim and build

Runner-auth routes (Bearer swt-…), at the root:

POST /worker/claim

Pull the oldest queued job for the runner’s arch: {"arch": "arm64"} → the full job, the lease, and the service’s resolved cache policy:

{
  "job": {"id": "j-…", "…": "…"},
  "lease_secret": "…",
  "lease_expires": "…",
  "lease_seconds": 90,
  "cache": "incremental"
}

204 when the queue is empty. The 90 s lease is heartbeat-extended; a lease lost to a dead runner is reaped server-side: requeued once silently, then failed infra-retryable with lease_lost in the failure metadata. The claim’s cache is the cache policy the service resolved for the job; a per-push cache override cannot cross the queue.

When the service runs a GitHub App whose installations cover the job’s github.com repository, the claim response’s git URL arrives rewritten to carry a freshly minted installation token (https://x-access-token:…@github.com/…); a private-repo clone needs no credential on the runner. The token is valid for seconds, appears in this one response only, and is never stored or logged; GitHub expires it after an hour regardless (see Import from GitHub: the GitHub App).

POST /worker/jobs/:id/heartbeat

{"phase": "build"} (advisory) → {"lease_expires": …, "action": "continue"} or {"action": "cancel"}. Cancellation comes back on the heartbeat, never on a second channel. A foreign lease → 409 (drop the job).

POST /worker/jobs/:id/logs

Append an ordered chunk: {"lines": […], "after_seq": 42} → {"seq": 57}. The service stamps its own sequence numbers, and an exact replay of the newest chunk is answered without double-appending (the retry-safety key).

Artifact blobs

Every byte that reaches the artifact store is written by the service, authorized by the job’s lease; a runner holds zero store credentials ([storage] in a runner’s config is ignored, with a loud warning). Blobs move in upload sessions:

POST /worker/jobs/:id/blobs

Open a session: {"digest": "<sha256 hex>", "size": 12345} → {"blob": "bs-…"}. The size is declared before the first byte moves and refused over the 2 GiB cap by name.

PUT /worker/jobs/:id/blobs/:blob

One chunk of raw bytes at the X-Saggar-Offset position → {"offset": …}, the session’s new durable size, so a client that fell behind resumes from there. An exact replay of written bytes is acknowledged without double-writing.

POST /worker/jobs/:id/blobs/:blob/finalize

Verify the spooled bytes against the declared digest → {"digest", "size"}. The service moves the bytes into its store under its own key (dedup when present); the runner never learns the key.

GET /worker/jobs/:id/source

Stream the job’s uploaded source tarball (known length). Git sources answer 404; the runner acquires those itself, as always.

Finish

POST /worker/jobs/:id/finish

Settle the job. The verdict (succeeded | failed | cancelled), failure metadata when failed (phase, exit code, outcome, class as the runner classified it), structured test results, the resolved env_key and content_id (the commit sha, or a content digest), and, on success, artifact descriptors naming the digests uploaded through this job’s sessions:

{
  "status": "succeeded",
  "env_key": "debian:12",
  "content_id": "9f2c…",
  "artifacts": [
    {
      "name": "my-project", "version": "1.0", "format": "deb",
      "arch": "amd64",
      "files": [
        {"filename": "my-project_1.0_amd64.deb",
         "media_type": "application/vnd.debian.binary-package",
         "sha256": "…", "size": 4096}
      ]
    }
  ]
}

The service verifies each named digest against what it actually received before the artifact rows exist, fills the server-side facts (node, retries, log tail, store keys), and answers with the job. A foreign lease → 409; an already-final job → its row (retries are safe).

Runner-side notes

The runner runs one build at a time (claim → build → report). Its data dir holds no database: the run’s working rows live in an in-memory SQLite that dies with the process, and the local spool behind the blob uploads plus the runner-local caches (warm trees, env snapshots, ccache) are swept on the runner’s own timer. A service restart is non-destructive: leases live server-side, so claimed jobs with live leases keep building through the downtime and only what actually lapsed is reaped.