.. _reference-worker-api: The runner protocol =================== How a ``saggar-worker`` runner pairs with the service, claims queued jobs, streams logs, pushes artifact blobs and settles the build. Everything here is machine-facing, at the server root (``/worker/…``, ``/nodes``). Like the registries, these paths are spoken by the ``saggar-worker`` binary and never carry the ``/api/v1`` prefix. Authentication levels --------------------- Three credentials ride this surface, one per relationship: Account token ``X-Auth-Token`` or ``Authorization: Bearer`` with an ``sgt-…`` token. Pairing and the roster reads. A registry-scoped token is refused. Runner token ``Authorization: Bearer swt-…``. The token issued at pairing, matched by hash against the roster. Every authenticated call refreshes the runner's ``last_seen``. Lease secret ``X-Saggar-Lease: ``. The claim's lease secret, the unguessable capability the claim minted, presented on every job-scoped call. An authenticated runner without it cannot touch a job it does not hold; a foreign lease answers ``409``. Guessing a job id buys nothing. Pairing ------- ``POST /worker/pair``, account auth: Pair the node with the presenting account. Body: .. code-block:: json { "node_id": "build-01", "arch": "arm64", "version": "0.2.0", "network": "online", "docker_format": true } ``node_id``, ``arch`` and ``version`` are the runner's identity; ``network`` (``online`` | ``pinned`` | ``offline``, default ``online``) and ``docker_format`` (default ``true``) are the owner's declared conduct profile, what the owner states the node can do, never probed (see :doc:`../how-to/runners`). Omitted fields pair with the safe defaults. The pairing is admitted inline; the account token proves the account, so the owner approves by pairing. The answer is ``{"status": "approved", "token": "swt-…"}``, the runner's bearer token, delivered once; the service stores only its hash. Registry tokens are refused, and attempts are capped per peer address. The issued token persists under the runner's data dir (``0600``) and is reused from there on; ``--token`` on ``saggar-worker`` (env ``SAGGAR_WORKER_TOKEN``) is the infra-as-code shape: a pre-provisioned runner token that skips pairing entirely. Roster ------ Management-auth routes (account tokens): ``GET /api/v1/nodes`` The runners the subject owns or is shared with: node id, arch, engine version, state (``active``, ``retired``), owner handle, ``owned`` flag, declared profile (``network``, ``docker_format``), last seen. ``GET /api/v1/nodes/:id/shares`` The accounts the runner is shared with (owner). ``PUT /api/v1/nodes/:id/shares`` Grant use to an account: ``{"handle": "…"}`` → the sharee list. Sharing grants use: a shared account's jobs may schedule on the runner, eligibility re-evaluated per claim. ``DELETE /api/v1/nodes/:id/shares/:handle`` Revoke a share → the sharee list. Future claims stop instantly; a running job finishes its lease. Management stays with the owner. ``POST /api/v1/nodes/retire`` Retire a runner you own: ``{"node_id": …}``. Its token stops authenticating immediately; the row stays listed ``retired`` for the record. Not idempotent: unknown or already-retired answers ``404``. Claim and build --------------- Runner-auth routes (``Bearer swt-…``), at the root: ``POST /worker/claim`` Pull the oldest queued job for the runner's arch: ``{"arch": "arm64"}`` → the full job, the lease, and the service's resolved cache policy: .. code-block:: json { "job": {"id": "j-…", "…": "…"}, "lease_secret": "…", "lease_expires": "…", "lease_seconds": 90, "cache": "incremental" } ``204`` when the queue is empty. The 90 s lease is heartbeat-extended; a lease lost to a dead runner is reaped server-side: requeued once silently, then failed ``infra-retryable`` with ``lease_lost`` in the failure metadata. The claim's ``cache`` is the cache policy the service resolved for the job; a per-push ``cache`` override cannot cross the queue. When the service runs a GitHub App whose installations cover the job's ``github.com`` repository, the claim response's git URL arrives rewritten to carry a freshly minted installation token (``https://x-access-token:…@github.com/…``); a private-repo clone needs no credential on the runner. The token is valid for seconds, appears in this one response only, and is never stored or logged; GitHub expires it after an hour regardless (see :doc:`../how-to/github-import`). ``POST /worker/jobs/:id/heartbeat`` ``{"phase": "build"}`` (advisory) → ``{"lease_expires": …, "action": "continue"}`` or ``{"action": "cancel"}``. Cancellation comes back on the heartbeat, never on a second channel. A foreign lease → ``409`` (drop the job). ``POST /worker/jobs/:id/logs`` Append an ordered chunk: ``{"lines": […], "after_seq": 42}`` → ``{"seq": 57}``. The service stamps its own sequence numbers, and an exact replay of the newest chunk is answered without double-appending (the retry-safety key). Artifact blobs -------------- Every byte that reaches the artifact store is written by the service, authorized by the job's lease; a runner holds zero store credentials (``[storage]`` in a runner's config is ignored, with a loud warning). Blobs move in upload sessions: ``POST /worker/jobs/:id/blobs`` Open a session: ``{"digest": "", "size": 12345}`` → ``{"blob": "bs-…"}``. The size is declared before the first byte moves and refused over the 2 GiB cap by name. ``PUT /worker/jobs/:id/blobs/:blob`` One chunk of raw bytes at the ``X-Saggar-Offset`` position → ``{"offset": …}``, the session's new durable size, so a client that fell behind resumes from there. An exact replay of written bytes is acknowledged without double-writing. ``POST /worker/jobs/:id/blobs/:blob/finalize`` Verify the spooled bytes against the declared digest → ``{"digest", "size"}``. The service moves the bytes into its store under its own key (dedup when present); the runner never learns the key. ``GET /worker/jobs/:id/source`` Stream the job's uploaded source tarball (known length). Git sources answer ``404``; the runner acquires those itself, as always. Finish ------ ``POST /worker/jobs/:id/finish`` Settle the job. The verdict (``succeeded`` | ``failed`` | ``cancelled``), failure metadata when failed (phase, exit code, outcome, class as the runner classified it), structured test results, the resolved ``env_key`` and ``content_id`` (the commit sha, or a content digest), and, on success, artifact descriptors naming the digests uploaded through this job's sessions: .. code-block:: json { "status": "succeeded", "env_key": "debian:12", "content_id": "9f2c…", "artifacts": [ { "name": "my-project", "version": "1.0", "format": "deb", "arch": "amd64", "files": [ {"filename": "my-project_1.0_amd64.deb", "media_type": "application/vnd.debian.binary-package", "sha256": "…", "size": 4096} ] } ] } The service verifies each named digest against what it actually received before the artifact rows exist, fills the server-side facts (node, retries, log tail, store keys), and answers with the job. A foreign lease → ``409``; an already-final job → its row (retries are safe). Runner-side notes ----------------- The runner runs one build at a time (claim → build → report). Its data dir holds no database: the run's working rows live in an in-memory SQLite that dies with the process, and the local spool behind the blob uploads plus the runner-local caches (warm trees, env snapshots, ccache) are swept on the runner's own timer. A service restart is non-destructive: leases live server-side, so claimed jobs with live leases keep building through the downtime and only what actually lapsed is reaped.