The runner protocol¶
How a saggar-worker runner pairs with the service, claims queued
jobs, streams logs, pushes artifact blobs and settles the build.
Everything here is machine-facing, at the server root (/worker/…,
/nodes). Like the registries, these paths are spoken by the
saggar-worker binary and never carry the /api/v1 prefix.
Authentication levels¶
Three credentials ride this surface, one per relationship:
- Account token
X-Auth-TokenorAuthorization: Bearerwith ansgt-…token. Pairing and the roster reads. A registry-scoped token is refused.- Runner token
Authorization: Bearer swt-…. The token issued at pairing, matched by hash against the roster. Every authenticated call refreshes the runner’slast_seen.- Lease secret
X-Saggar-Lease: <secret>. The claim’s lease secret, the unguessable capability the claim minted, presented on every job-scoped call. An authenticated runner without it cannot touch a job it does not hold; a foreign lease answers409. Guessing a job id buys nothing.
Pairing¶
POST /worker/pair, account auth:Pair the node with the presenting account. Body:
{ "node_id": "build-01", "arch": "arm64", "version": "0.2.0", "network": "online", "docker_format": true }
node_id,archandversionare the runner’s identity;network(online|pinned|offline, defaultonline) anddocker_format(defaulttrue) are the owner’s declared conduct profile, what the owner states the node can do, never probed (see Pair your own runners). Omitted fields pair with the safe defaults.The pairing is admitted inline; the account token proves the account, so the owner approves by pairing. The answer is
{"status": "approved", "token": "swt-…"}, the runner’s bearer token, delivered once; the service stores only its hash. Registry tokens are refused, and attempts are capped per peer address.The issued token persists under the runner’s data dir (
0600) and is reused from there on;--tokenonsaggar-worker(envSAGGAR_WORKER_TOKEN) is the infra-as-code shape: a pre-provisioned runner token that skips pairing entirely.
Roster¶
Management-auth routes (account tokens):
GET /api/v1/nodesThe runners the subject owns or is shared with: node id, arch, engine version, state (
active,retired), owner handle,ownedflag, declared profile (network,docker_format), last seen.GET /api/v1/nodes/:id/sharesThe accounts the runner is shared with (owner).
PUT /api/v1/nodes/:id/sharesGrant use to an account:
{"handle": "…"}→ the sharee list. Sharing grants use: a shared account’s jobs may schedule on the runner, eligibility re-evaluated per claim.DELETE /api/v1/nodes/:id/shares/:handleRevoke a share → the sharee list. Future claims stop instantly; a running job finishes its lease. Management stays with the owner.
POST /api/v1/nodes/retireRetire a runner you own:
{"node_id": …}. Its token stops authenticating immediately; the row stays listedretiredfor the record. Not idempotent: unknown or already-retired answers404.
Claim and build¶
Runner-auth routes (Bearer swt-…), at the root:
POST /worker/claimPull the oldest queued job for the runner’s arch:
{"arch": "arm64"}→ the full job, the lease, and the service’s resolved cache policy:{ "job": {"id": "j-…", "…": "…"}, "lease_secret": "…", "lease_expires": "…", "lease_seconds": 90, "cache": "incremental" }
204when the queue is empty. The 90 s lease is heartbeat-extended; a lease lost to a dead runner is reaped server-side: requeued once silently, then failedinfra-retryablewithlease_lostin the failure metadata. The claim’scacheis the cache policy the service resolved for the job; a per-pushcacheoverride cannot cross the queue.When the service runs a GitHub App whose installations cover the job’s
github.comrepository, the claim response’s git URL arrives rewritten to carry a freshly minted installation token (https://x-access-token:…@github.com/…); a private-repo clone needs no credential on the runner. The token is valid for seconds, appears in this one response only, and is never stored or logged; GitHub expires it after an hour regardless (see Import from GitHub: the GitHub App).POST /worker/jobs/:id/heartbeat{"phase": "build"}(advisory) →{"lease_expires": …, "action": "continue"}or{"action": "cancel"}. Cancellation comes back on the heartbeat, never on a second channel. A foreign lease →409(drop the job).POST /worker/jobs/:id/logsAppend an ordered chunk:
{"lines": […], "after_seq": 42}→{"seq": 57}. The service stamps its own sequence numbers, and an exact replay of the newest chunk is answered without double-appending (the retry-safety key).
Artifact blobs¶
Every byte that reaches the artifact store is written by the service,
authorized by the job’s lease; a runner holds zero store credentials
([storage] in a runner’s config is ignored, with a loud warning).
Blobs move in upload sessions:
POST /worker/jobs/:id/blobsOpen a session:
{"digest": "<sha256 hex>", "size": 12345}→{"blob": "bs-…"}. The size is declared before the first byte moves and refused over the 2 GiB cap by name.PUT /worker/jobs/:id/blobs/:blobOne chunk of raw bytes at the
X-Saggar-Offsetposition →{"offset": …}, the session’s new durable size, so a client that fell behind resumes from there. An exact replay of written bytes is acknowledged without double-writing.POST /worker/jobs/:id/blobs/:blob/finalizeVerify the spooled bytes against the declared digest →
{"digest", "size"}. The service moves the bytes into its store under its own key (dedup when present); the runner never learns the key.GET /worker/jobs/:id/sourceStream the job’s uploaded source tarball (known length). Git sources answer
404; the runner acquires those itself, as always.
Finish¶
POST /worker/jobs/:id/finishSettle the job. The verdict (
succeeded|failed|cancelled), failure metadata when failed (phase, exit code, outcome, class as the runner classified it), structured test results, the resolvedenv_keyandcontent_id(the commit sha, or a content digest), and, on success, artifact descriptors naming the digests uploaded through this job’s sessions:{ "status": "succeeded", "env_key": "debian:12", "content_id": "9f2c…", "artifacts": [ { "name": "my-project", "version": "1.0", "format": "deb", "arch": "amd64", "files": [ {"filename": "my-project_1.0_amd64.deb", "media_type": "application/vnd.debian.binary-package", "sha256": "…", "size": 4096} ] } ] }
The service verifies each named digest against what it actually received before the artifact rows exist, fills the server-side facts (node, retries, log tail, store keys), and answers with the job. A foreign lease →
409; an already-final job → its row (retries are safe).
Runner-side notes¶
The runner runs one build at a time (claim → build → report). Its data dir holds no database: the run’s working rows live in an in-memory SQLite that dies with the process, and the local spool behind the blob uploads plus the runner-local caches (warm trees, env snapshots, ccache) are swept on the runner’s own timer. A service restart is non-destructive: leases live server-side, so claimed jobs with live leases keep building through the downtime and only what actually lapsed is reaped.