Kanea

Documentation

Overview

Kanea runs containers on one machine and gives that machine the things a platform normally needs a cluster for: service discovery, load balancing, network policy, TLS, autoscaling, GitOps, backups. It is one static binary, and the concepts it asks you to learn number about six.

If you have not installed it yet, Installing is the next section: one command for the binary, one for the node. The rest of this page assumes you have a node running and want to know what you are looking at.

Concepts

Kanea borrows from Nomad and Kubernetes and keeps the smallest set of concepts that covers the work. If you know either system, the middle columns tell you what maps to what; if you are coming from a managed cloud service, the ECS and Container Apps table is below.

KaneaNomadKubernetesWhat it is
NodeClient + server agentNode + control planeOne machine running kanea agent. In v1 there is exactly one.
ProjectNamespaceNamespaceA named group of services, and the isolation boundary: network policy, secrets, containerd namespace and DNS all key off it.
ServiceJob + groupDeployment + ServiceA declarative long-running workload with count replicas.
TaskTaskContainerThe container inside a service. Exactly one per service in v1: sidecars are a v1.1 concept.
AllocAllocationPodOne running instance of a service. count = 3 means three allocs.
StorageCSI volumePV / PVCA named volume backend (local, host, nfs, smb or s3) that services mount by name.
Pipeline-Tekton / CI jobA build run producing an image, from a build block or a git push.

The one that surprises people is project. It is not a label: it is a boundary that four subsystems enforce independently. A service in project shop cannot reach a service in analytics unless somebody wrote that down, it cannot read analytics' secrets even by naming them, its containers live in a separate containerd namespace, and its DNS names sit under a separate suffix.

Coming from ECS or Container Apps

The two managed services closest to Kanea in shape are Amazon ECS and Azure Container Apps. (On Azure, Container Apps is the right comparison: AKS is the Kubernetes column above, and Container Instances is a single container with no orchestration around it.) Both map onto Kanea closely enough to be worth a table - but read the ECS column carefully, because one of its terms means the opposite of what it means here:

KaneaAmazon ECSAzure Container Apps
NodeCluster, plus its container instancesEnvironment
Project- (a cluster is the nearest thing)Environment, again
ServiceServiceContainer app
TaskContainer definitionContainer
AllocTaskReplica
Spec fileTask definitionApp YAML / ARM template
Deploy (spec-hash change)New task definition revisionNew revision
StorageVolume: bind mount, EFS, FSxVolume mount: Azure Files, ephemeral
Internal DNSService Connect / Cloud MapBuilt-in environment DNS
expose + the edgeALB target group + listener ruleIngress
scalingApplication Auto ScalingKEDA scale rules
PipelineCodePipeline + CodeBuildACR Tasks / GitHub Actions
SecretsSecrets Manager / SSM, by ARNContainer app secrets / Key Vault
"Task" means the opposite thing in ECS

In ECS a task is a running instance - the unit a service keeps desiredCount of, and the direct equivalent of a Kubernetes pod. In Kanea that is an alloc. A Kanea task is the container inside a service, which ECS calls a container definition and puts inside a task definition.

So "three tasks" means three replicas on ECS and is not a sentence you can say in Kanea, where a service has exactly one task and three allocs. It is the single most reliable way to misread this documentation coming from ECS.

Where the mapping breaks down

  • Project has no ECS equivalent. ECS groups by cluster and otherwise relies on IAM and tags, so there is no boundary that carries network policy, secret scope and DNS at once the way a Kanea project does. On Container Apps the environment is the closest thing, which is why it appears twice in the table above: it is both the hosting boundary and the network boundary, where Kanea splits those into the node and the project.
  • One node, not a region. Both services schedule across availability zones and survive a machine dying; Kanea v1 is a single machine, and if it stops, everything on it stops. That is the trade the whole design makes, and state replication and restore is the answer it offers instead of failover: an encrypted snapshot and change log in S3, restored onto a new box.
  • Nothing is billed per control plane, because there is no control plane to rent. kanead is a process on your machine inside a 256 MiB reservation. The flip side is that patching, capacity and the node's own uptime are yours.
  • Load balancing is in the datapath, not a resource you create. A service gets a virtual IP the moment it exists, and east-west traffic is rewritten at connect() by eBPF rather than crossing a proxy at all. Container Apps gives you ingress built in and is closest here; on ECS the equivalent is a load balancer, target group and listener rules to declare, wire up and pay for. North-south traffic goes through the edge, which is a process on the same box, not a product.
  • The spec is one file for the whole project. An ECS deployment is typically a task definition plus a service definition plus target groups plus listener rules, assembled by Terraform or CDK; a Kanea project is one HCL file that kanea plan validates offline, covering the containers, routing, TLS, volumes, scaling and builds together.

Naming

Project and service names must be DNS-1123 labels: lowercase alphanumeric and -, starting and ending alphanumeric, at most 63 characters. This is checked when the spec is parsed, not when something fails later.

The reason is that names compose into DNS without an escaping step. A service web in project shop is web.shop.kanea internally and web.shop.<base_domain> publicly. A name that needed encoding to become a hostname would be a name that means two different things in two places.

Where the prose goes

Every project and service takes a description: free text, up to 512 characters, shown in the dashboard. That is where the human-readable detail belongs; the name stays a label.

Lifecycle

Nothing in Kanea is applied directly to the runtime. A spec becomes desired state in the Store, and a reconciler converges the world toward it: continuously, not once.

Job spec (HCL) ──parse/validate──▶ Desired state (Store)
                                        │
                                   Reconciler loop
                                        │
                        ┌───────────────┼────────────────┐
                        ▼               ▼                ▼
                    containerd      eBPF datapath    Edge proxy
                   (tasks/images) (policy/LB)       (routes/TLS)
                        │               │                │
                        └───────────────┴────────────────┘
                                        ▼
                            Actual state / events / metrics

Consequences worth knowing up front:

  • Drift is repaired. Delete a container by hand and it comes back. The reconciler is comparing, not remembering what it did.
  • Restart policies are always (default), on-failure with backoff, and never.
  • Updates are rolling by default, health-gated, bounded by max_parallel.
  • Dependencies start first. A depends_on edge (or any ${service.…} reference, which creates one implicitly) means the dependent does not start until its dependency is healthy. If the dependency degrades later, dependents keep running: no cascading stops, just events.
  • Storms are capped. Per-service restart rate limits plus a node-wide circuit breaker that pauses rollouts and scale actions when failure rates spike. A trip emits an event and a notification.

What counts as a deploy

A deploy is a spec-hash mismatch. The reconciler hashes the parts of a service that are baked into a container at creation time; if an alloc's recorded hash differs from the current one, that alloc is replaced under the update policy. If it matches, nothing happens, however many times you run kanea run.

This is why kanea restart is not a separate path into the runtime: it bumps a generation counter that participates in the hash, and the ordinary rolling update does the rest.

Two things about rolling updates

max_parallel bounds allocs that are down, not replacements in flight. Anything already unavailable spends the budget first, so a deploy that starts going wrong stops, instead of walking through every replica.

min_healthy applies only to allocs the current deploy has already replaced, and health means a probe said so. A service with no health_check block never reports healthy for any alloc, which is fine, and is why the update logic asks whether a check is configured before it asks whether it passed.

Secrets

Secrets are never written in a spec. They are referenced as secret:<path> and resolved when an alloc starts:

env = {
  DATABASE_URL = "secret:shop/database-url"
}

The variable then carries the path of a file in a per-alloc tmpfs mount (/run/kanea/secrets/<alloc>/shop/database-url, read-only, owned by the service's user) - the value never sits in an environment block. Software that can only read environment variables opts into the weaker form, secret-env:shop/database-url, which inlines the value (visible via /proc/<pid>/environ); the record keeps the reference either way, and a rotated secret takes effect at the next replacement.

Two rules make that reference safe rather than merely tidy:

  • References are project-scoped at validation time. A service in shop may name secret:shop/… or secret:shared/… and nothing else. A spec reaching for another project's secret fails to parse: it does not fail at runtime, where the failure would be a log line nobody reads.
  • The default injection is a tmpfs file, at /run/kanea/secrets/<alloc>/<name>. Environment variables work and are documented as the weaker option, because they are visible in /proc/<pid>/environ, in runtime inspect APIs, and to every child process.

The API and the MCP server are write-only for secrets. There is no get, not for an operator, not for an agent, at any permission tier.

Your first service

The smallest useful thing needs no file at all:

kanea run --image nginx:1.27-alpine --name web --project demo
kanea ps -p demo
kanea logs -f demo/web

When you want it written down, the same deployment as a spec is three lines plus a wrapper, and from there every block in the job spec reference is additive. Run kanea plan first; it prints the create/change/destroy diff, and it is where every validation rule fires.

Installing

Two commands: one puts the binary on the node, one turns the node into a platform. The first only ever writes /usr/local/bin/kanea; everything else (the runtime, the keys, the units, the accounts) belongs to kanea init, which you can read before you run.

RequirementWhat it must be
Architectureslinux/amd64, linux/arm64. Anything else is refused before a byte is downloaded.
Kernel≥ 5.10, cgroups v2 unified hierarchy.
Init systemsystemd.
ClockNTP-synchronised (certificates and audit timestamps depend on it).

That is the whole list. containerd, runc, rootless buildkitd and the wasmtime shim are not prerequisites: kanea init installs them at versions pinned by SHA-256 inside the binary. The network layer needs no component at all, because the eBPF datapath is compiled into kanea.

The install script

curl -fsSL https://m18h.github.io/kanea/install.sh | sudo bash

It resolves the latest release, downloads the archive for this architecture, verifies it, and installs one binary. Then it stops. It generates no keys, starts no services, writes no units and installs no runtime, so running it on a node that is already serving traffic changes nothing until you act on what it prints.

There are no arguments. The three knobs are environment variables, and sudo does not forward them by default - use sudo VAR=value bash or sudo -E:

VariableDefaultMeaning
KANEA_VERSIONlatestThe release tag to install, e.g. vX.Y.Z. latest is resolved from the redirect GitHub already serves, so no jq and no API rate limit.
KANEA_PREFIX/usr/local/binDirectory the binary is installed into, mode 0755.
KANEA_REPOm18h/kaneaThe owner/repo releases come from, so a fork installs from itself.
KANEA_REQUIRE_SIGNATURE0Set to 1 to make the cosign verification mandatory: a missing cosign binary or an unsigned release becomes fatal instead of a note. The script never downloads cosign for itself.
curl -fsSL https://m18h.github.io/kanea/install.sh | sudo KANEA_VERSION=vX.Y.Z bash

It needs curl, tar and install on the node, and it must run under bash, not sh. Four things are refused up front rather than discovered after a download: a non-Linux kernel, an architecture outside amd64 and arm64, an unreachable release endpoint, and a repository with no published release (which would otherwise resolve to a plausible-looking archive name and a 404 nobody can read).

What is verified

The checksum is mandatory and there is no flag to skip it. The script downloads checksums.txt, takes the line for the archive it fetched, and runs it through sha256sum -c. A mismatch stops with checksum mismatch: do not run this binary.

The Sigstore signature is checked when cosign is on PATH: the keyless signature over checksums.txt, bound to this repository's release workflow through the GitHub OIDC issuer. The asymmetry is deliberate and worth being precise about:

  • cosign present, signature valid - prints Signature verified.
  • cosign present, signature invalid - fatal. The install stops.
  • cosign present, no signature published - a note, then continues on the checksum alone.
  • cosign absent - prints cosign not found; checksum verified but signature not checked and continues. Installing cosign first is the stronger path; verifying by hand is the same check written out.

KANEA_REQUIRE_SIGNATURE=1 turns the last two endings into refusals: a node that cares gets to insist on the signature, and kanea upgrade --require-signature is the same posture for every upgrade after this one.

One side effect, and only one

The script creates the kanea-edge system account if it is missing (useradd --system --no-create-home --shell /usr/sbin/nologin). The ingress process runs as that user rather than root, and a failure here is a warning, never fatal. Nothing else on the node is touched.

Setting up the node

sudo kanea init is the rest of the way: preflight checks, the host components, the master-key ceremony, the systemd units, the first admin account, and a summary of what it built. Run it once, as root. It is safe to re-run - see what a re-run does below.

sudo kanea init
$ sudo kanea init
kanea init: <version>

API/dashboard listen address [127.0.0.1:8600] ("none" for socket-only): 192.168.1.10:8600
Listener certificate already present at /etc/kanea/api.crt; leaving it alone.

Checking this node:
  ok    platform         linux/amd64
  ok    cgroups v2       unified hierarchy with cpu, memory and pids
  ok    kernel           6.12.43-amd64
  ok    clock            synchronised (systemd-timesyncd)
  ok    systemd          running
  ok    data directory   /var/lib/kanea
  ok    edge user        kanea-edge

Installing the host components:
kanea install: <version>

Source: upstream
Prefix: /usr/local/lib/kanea

  up-to-date  containerd 2.3.3     container runtime daemon
  up-to-date  runc 1.5.1           OCI runtime
  up-to-date  wasmtime-shim 0.6.1  wasmtime containerd shim; functions

Wrote /etc/kanea/containerd/config.toml
Wrote /etc/systemd/system/kanea-containerd.service

Starting containerd to pull the image components…
  up-to-date  buildkit v0.32.0-rootless  rootless build daemon (the only build driver)

Wrote /etc/systemd/system/kanea-buildkit.service

Created /var/lib/kanea, /var/log/kanea/allocs and /etc/kanea
Master key already present at /var/lib/kanea/master.key; leaving it alone.
Wrote /etc/systemd/system/kanea.slice
Wrote /etc/systemd/system/kanea-workloads.slice
Wrote /etc/systemd/system/kanead.service
Wrote /etc/systemd/system/kanea-edge.service

Starting kanead:
  waiting for the control plane…
First admin username: <name>
password for <name>:
again:
Created admin account "<name>".
Restarting kanead so the settled listener takes effect…

───────────────────────────────────────────────────────────────
Kanea is running.

  Dashboard      https://192.168.1.10:8600   (log in as "<name>")
  Internal DNS   10.244.0.1:53   (allocs resolve <service>.<project> here)

  Addressing
    node CIDR      10.244.0.0/24      this node's allocs
    cluster CIDR   10.244.0.0/16      routed, masqueraded as internal
    service CIDR   10.201.0.0/16      service VIPs

Deploy something:  kanea run <spec.hcl>
CLI without sudo:  sudo usermod -aG kanea <user>   # root-equivalent; log in again
Configure a backup destination before you need one; see docs/DR_RUNBOOK.md.

Versions and the kernel string are placeholders; the shape is what a run looks like. This one is a re-run on a node that already had a key and a listener certificate, which is why two lines say "leaving it alone" and the components read up-to-date. A first run mints both and installs each component.

What it asks

  1. The listen address, defaulting to 127.0.0.1:8600. Answer none to keep the API on its unix socket only. This is asked only when --listen was not passed and stdin is a terminal, so a scripted install never has a prompt eat one of its lines. A bind stanza in the server config replaces the question entirely.
  2. The master key, typed back. It is printed once and you must retype it exactly; a mismatch writes nothing and stops. There is no flag to accept it unseen.
  3. The first admin's username, unless --admin-user was given.
  4. That account's password, twice.
The key is shown once

Without it, every backup this node ever makes is unreadable and every stored secret is unrecoverable. Have somewhere to record it before you start. Every prompt in a run shares one reader, so a piped install works exactly as a typed one does.

TLS for the dashboard

Beyond loopback the listener is served over TLS, and init provisions it rather than refusing: with no explicit --listen-cert/--listen-key pair it mints a ten-year self-signed pair at /etc/kanea/api.crt and api.key (key mode 0600, with a real IP SAN when the address is an IP) and points the unit at it. It is minted once and never re-minted, because re-minting would flip the fingerprint an operator has already accepted.

Bring your own with the two flags, or hand the whole listener to the bind stanza, which also gets you acme and the node CA. One case is still refused: an unspecified host (0.0.0.0, :8600), because a certificate's SAN needs something to name - bind a specific address, or set bind.api_domain.

What a re-run does

KeptRewritten
The master key, always. An existing one is never regenerated.The four systemd units, every time, at mode 0644.
The api.crt/api.key pair. Half a pair is a hard error, never a guess.The containerd config and component units.
The first admin. If any account exists the step is skipped, never re-prompted.
/etc/kanea/kanea.hcl. Init never writes it, on a first run or a re-run.

That is exactly why a release whose notes mention changed units is a re-run of sudo kanea init followed by systemctl daemon-reload: it is the supported way to regenerate them, and it costs you neither the key nor the accounts. kanead is only restarted when the settled listener actually differs from the running one, so an unchanged re-run restarts nothing.

Useful flags

The full list is in the CLI reference; these are the ones that come up:

FlagDefaultWhy
--listen127.0.0.1:8600Skips the prompt. none keeps the API socket-only.
--admin-userpromptWith a piped password, makes the whole run scriptable.
--containerd externalinstall our ownAdopt the containerd already on the node instead of installing one.
--bundle-Install the host components from an offline bundle.
--reserve256MThe control plane's memory floor. A node that runs builds wants 512M: buildkitd alone holds ~157 MiB.
--node-cidr10.244.0.0/24This node's container subnet. Its .1 is also the internal DNS address.
--no-startfalseWrite everything and stop, without starting kanead or creating an account.
--skip-checksfalseRun the ceremony without the preflight. The checks gate, so this is how you override one you disagree with.

Host components

containerd, runc, rootless buildkitd and the wasmtime shim are installed by Kanea, pinned by version and SHA-256 in a manifest compiled into the binary. Hashes are never fetched, so moving a component is a code change with a review behind it, and the same manifest is the version matrix kanea doctor checks the node against.

kanea install --list         # the pinned versions, straight from the binary
sudo kanea install --dry-run # resolve and verify every artefact, write nothing
sudo kanea install           # normally you never run this; kanea init does

Nothing at a distribution's paths is touched. Kanea's containerd lives under /usr/local/lib/kanea on its own socket at /run/kanea/containerd.sock, which is why a node that ran Docker yesterday runs it tomorrow.

FlagMeaning
--listPrint the pinned matrix and exit. Needs no network and no root.
--dry-runDownload and verify everything, write nothing. The honest pre-flight for an upgrade.
--onlyComma-separated component names. They are reordered into manifest order, because install order is a dependency: containerd and runc first, containerd starts, and only then is the buildkit image pulled through it.
--forceReinstall components already at the pinned version.
--bundleInstall from an offline bundle. This turns network fetching off entirely; it is not a preference.
--containerd externalAdopt an existing containerd: drops containerd from the set and writes a slice drop-in instead.
--archTarget architecture. Meaningful only with --dry-run; installing for a foreign architecture is refused.
--prefix, --conf-dir, --data-dir, --run-dir, --unit-dirWhere binaries, configuration, state, sockets and units go.
--skip-unitsBinaries only. Also skips the image components, since containerd is never started.

Installing by hand

Every release publishes per-architecture archives, an SPDX SBOM beside each one, a kanea_<version>_source.spdx.json for the build's own graph (where the embedded dashboard's npm dependencies are listed), a checksums.txt, and a keyless cosign signature over that file. The SBOMs are listed inside the checksums, so the one signature covers them too:

VERSION=vX.Y.Z; ARCH=amd64   # the tag you want, e.g. from the releases page
BASE=https://github.com/m18h/kanea/releases/download/$VERSION

curl -fLO $BASE/kanea_${VERSION#v}_linux_$ARCH.tar.gz
curl -fL -O $BASE/checksums.txt -O $BASE/checksums.txt.sig -O $BASE/checksums.txt.pem

cosign verify-blob \
  --certificate checksums.txt.pem \
  --signature   checksums.txt.sig \
  --certificate-identity-regexp "https://github.com/m18h/kanea/" \
  --certificate-oidc-issuer "https://token.actions.githubusercontent.com" \
  checksums.txt

sha256sum --ignore-missing -c checksums.txt
tar xzf kanea_${VERSION#v}_linux_$ARCH.tar.gz
sudo install -m 0755 kanea /usr/local/bin/kanea

There is no long-lived signing key to guard. The signature is bound by Sigstore to the release workflow in this repository, and the proof is in a public transparency log. Then continue at sudo kanea init.

Homebrew

brew tap m18h/kanea
brew trust m18h/kanea   # brew ≥ 6 refuses formulae from untrusted third-party taps
brew install kanea

Homebrew is a CLI channel, never a node channel. On macOS you get the authoring half, where kanea plan validates job specs with file-and-line diagnostics and needs no daemon. On Linux the formula installs the same full binary, but a node belongs to the install script above: root-owned at /usr/local/bin, where kanea upgrade owns the swap. A brew-owned binary upgrades with brew upgrade kanea, then sudo kanea upgrade --no-fetch for the restart-and-migrate half.

The container image

docker run --rm ghcr.io/m18h/kanea:latest version

# the shape a pipeline uses: the spec on a mount, the node over the network
docker run --rm -v "$PWD:/workspace:ro" \
  -e KANEA_URL -e KANEA_TOKEN \
  ghcr.io/m18h/kanea:vX.Y.Z run shop.hcl   # pin it; that is what tags are for

linux/amd64 and linux/arm64 in one manifest list, tagged vX.Y.Z and latest; latest moves only for a plain vX.Y.Z, so a prerelease never becomes what everyone pulls by default. Like Homebrew it is a CLI channel, never a node channel: it is for the pipeline that deploys to a node, so every verb that acts on a host, its systemd and its own binary - agent, edge, init, install, doctor, upgrade - is not what it carries.

The image contains the released binary, not a rebuild of it. The release workflow builds the image after it signs, from the same linux_amd64/linux_arm64 archives you can download, after verifying checksums.txt under its signature and each archive against it - the two checks the by-hand install performs, in that order. So the manifest of hashes you can verify yourself describes what is inside the image, with no second build to reconcile.

The digest is signed with the same keyless identity as the checksums, and carries an SPDX attestation beside the signature:

cosign verify ghcr.io/m18h/kanea:vX.Y.Z \
  --certificate-identity-regexp 'https://github.com/m18h/kanea/' \
  --certificate-oidc-issuer 'https://token.actions.githubusercontent.com'
Two things about running it

It runs as an unprivileged user (uid 65532) with /workspace as its working directory, so mount your spec there. And if your network terminates TLS with its own CA, either drop the root into /usr/local/share/ca-certificates and run update-ca-certificates in a derived image, or pass the node's own CA with KANEA_CA_CERT - which replaces the system pool rather than adding to it, so it is the node's CA or the public ones, never both.

Upgrading

sudo kanea upgrade alone is the whole upgrade: it downloads and verifies the release itself before restarting anything. Re-running the install script works too, and is what you do when the binary is what you want replaced first; it changes the binary only, and says so:

re-running the install script
$ curl -fsSL https://m18h.github.io/kanea/install.sh | sudo bash
Upgrading kanea <installed> -> <latest> (linux/amd64)
Verifying the checksum
cosign not found; checksum verified but signature not checked
Installed /usr/local/bin/kanea

The running daemons are still on the old binary. Next:

    sudo kanea upgrade

It backs up, drains and restarts kanea-edge, then restarts kanead, which
runs any state migrations. Running workloads are untouched throughout.

The sequence kanea upgrade performs, in order: resolve and verify the release (sha256 always, cosign when present, fatal if present and failing), install it atomically over its own path, take a pre-upgrade backup, restart kanea-edge, restart kanead, run state migrations, and wait for health. The edge goes first because it drains in seconds and must be current before kanead publishes a new-format route snapshot; kanead goes last because it runs the migration once everything else has settled.

Running allocs are not touched at any point. Already being at the target version means nothing to download, so running it twice is safe by construction.

An admin can run the same flow from the dashboard's Updates page, reached from a pinned sidebar item whose badge marks a newer published release or a pending OS reboot. The page checks for updates and installs one (GET and POST /v1/upgrade, admin-only, audited). The daemon performs the sequence above on itself - pre-upgrade backup, fetch and verify, atomic install - then answers the request, restarts kanea-edge, and exits into systemd's Restart=always; the page reloads once the new version answers health, which is also what swaps in the new embedded dashboard. The check never runs unprompted: the node asks GitHub only while an admin has the dashboard open, and caches the answer for an hour, so a node never phones home on its own. There is no downgrade over the API (--allow-downgrade stays CLI-on-the-node), and outside systemd nothing restarts: the binary is installed and the response says restart_required. KANEA_REQUIRE_SIGNATURE=1 on the unit gives this path the --require-signature posture.

FlagMeaning
--checkReport the running, installed and latest versions, and stop.
--version vX.Y.ZGo to a specific release instead of the latest.
--no-fetchDownload nothing; restart onto whatever binary is already installed. This is the half a package manager leaves for you.
--require-signatureRefuse the release unless its cosign signature verifies: a missing cosign binary, or a release published without one, becomes fatal instead of a note. Never downloads cosign itself - a verifier fetched over the same channel would be a second trust root vouching for the first.
--allow-downgradePermit a target older than the running daemon. Refused otherwise, because the Store's schema is forward-only.
--skip-backupSkip the pre-upgrade backup. The schema migration still takes its own local copy.
--dry-runPrint what would run and stop.
Upgrading never rewrites systemd units

That is deliberate: your units may carry local edits. When release notes say the units changed, re-run sudo kanea init (idempotent - key, accounts and settings are kept) and then sudo systemctl daemon-reload. Host components are pinned by the new binary, so sudo kanea install --dry-run shows whether any need to move and sudo kanea doctor confirms the node agrees with the matrix.

Air-gapped nodes

A node with no egress is a supported installation, not a workaround. Build a bundle where there is a network, carry it across, install from it:

kanea bundle create --arch amd64 -o kanea-bundle.tar.gz   # connected machine
sudo kanea init --bundle kanea-bundle.tar.gz              # air-gapped node

The bundle carries no hashes of its own. Its contents are verified against the ones compiled into the installing node's binary; a bundle that supplied its own would be a bundle that authenticates itself. For image components that means a digest comparison rather than a name lookup, because containerd names image records from the archive's own annotation verbatim.

Releases publish one bundle per architecture, covered by the same signed checksums.txt. This covers Kanea's own components; your workload images still come from a registry the node can reach. On an air-gapped node kanea doctor --offline skips the single network probe, and kanea upgrade refuses to fetch with a pointer to this flow rather than hanging.

Architecture

One binary produces two long-running processes and a CLI. The datapath that networks your services is not one of them: it is a set of eBPF programs compiled into the binary and loaded into the kernel, not a daemon Kanea drives. Everything else on the node (containerd, buildkitd) is software Kanea drives rather than software it contains, but it is no longer software you have to find: kanea init installs each of them at a version pinned by SHA-256 in the binary, under Kanea's own prefix and on Kanea's own sockets, so nothing already on the node changes.

The shape

                    ┌──────────────── kanead (control plane) ────────────────┐
                    │                                                        │
Browser ──HTTPS──▶  │ ┌────────────┐ ┌───────────┐ ┌──────────┐ ┌─────────┐  │
CLI ──────HTTPS──▶  │ │ API server │ │ Dashboard │ │Reconciler│ │Autoscale│  │
Webhooks ────────▶  │ │ REST + WS  │ │ (embedded)│ │          │ │ (eBPF)  │  │
                    │ └─────┬──────┘ └───────────┘ └────┬─────┘ └────┬────┘  │
                    │       └─────────────┬─────────────┴────────────┘       │
                    │            ┌────────┴─────────┐                        │
                    │            │ Store  (bbolt)   │                        │
                    │            └────────┬─────────┘                        │
                    │   ┌─────────┬───────┼────────┬──────────┐              │
                    │   │ Runtime │Network│ GitOps │ Notifier │              │
                    │   │containerd│ eBPF │BuildKit│          │              │
                    │   └────┬────┴───┬───┴────┬───┴──────────┘              │
                    └────────┼────────┼────────┼───────────────┬─────────────┘
                             │        │        │               ▼
                             │        │        │      State replicator ──▶ S3
                             ▼        ▼        ▼

┌── kanea-edge (separate process; reads a projection, never the Store) ──┐
│  L7 routing · TLS termination · middleware · request metrics           │ ◀── :80/:443
└────────────────────────────────────────────────────────────────────────┘

External:  containerd  ·  buildkitd  ·  Linux kernel (eBPF, cgroups v2, netfilter, bpffs)

Two processes, one binary

kanead is the control plane: API, dashboard, reconciler, autoscaler, GitOps, notifications, backups. kanea-edge is the ingress proxy, and it is a separate systemd unit running as a separate unprivileged user.

That split is the single most consequential design decision in the system, and it buys two things:

  • Restarting, upgrading or crashing the control plane does not interrupt public traffic. The edge unit deliberately has no After=kanead.service.
  • The process that terminates untrusted public traffic cannot mutate the platform. It has no Store access and no write path at all.
How the edge learns anything

bbolt takes a lock on the whole database file, so a second process opening it (even read-only) would block until kanead exits. So the edge does not open it. kanead projects what the edge needs into /run/kanea-edge/: routes.json (0644, host → service frontend; nothing secret, the domains are in public DNS) and certs.json (0640, private keys). Two files, two permission sets, so neither has to compromise for the other.

Both are written temp-then-rename(2), so a half-written file is never observable. The projection carries the Store index it was built from, which is how a stale snapshot can be recognised rather than merely suspected.

A missing or stale snapshot is not an outage: the edge keeps serving the last table it loaded for as long as kanead is away, and starts with an empty table rather than refusing to start. "The control plane is down" must never become "the site is down".

The Store

One embedded bbolt database holds every mutation, behind a Store interface with monotonic indexes. Buckets: projects, services, allocs, events, certs, secrets, pipelines, audit, kv.

  • Single writer. All mutations serialise. Reads are bounded and paginated, because a long read transaction blocks the writer.
  • Raft-shaped on purpose. The interface and its index semantics are what a Raft FSM would need, so a clustered implementation can replace it without touching call sites.
  • Metrics and logs never touch it. Time series live in a bounded in-memory ring; logs go to file pipelines with non-blocking drains. This is a hard constraint, not a performance preference: a metrics write that contends with the reconciler's writer would make observability and convergence share a failure mode.
  • Migrations are explicit. The Store does not migrate itself at open. A migration rewrites state in place, and the copy that makes a bad migration survivable needs the database open and the migration not yet started: exactly one window, between Open and Migrate.

Runtime driver: containerd

  • Driven over its socket with the official Go client. One containerd namespace per project (kanea-<project>), which makes image and container isolation free rather than enforced.
  • Responsibilities: image pull (credentials from the secrets store, digest pinning supported), task lifecycle, per-alloc netns setup, cgroup metrics sampling, and stdout/stderr capture.
  • Hardening is not opt-in. Every alloc starts from a baseline capability set; the uid-switching grants PUID-style images need to chown a volume and drop to their configured user, and nothing more (CAP_NET_RAW is deliberately excluded, and binding :80 needs no capability at all - the alloc netns has no privileged-port floor); plus no-new-privileges, the default seccomp profile, and its own PID and IPC namespaces. capabilities = ["none"] drops to nothing; a capability a workload genuinely needs beyond the baseline is named in the spec and drawn from a permitted set; the privilege-equivalent ones are rejected when the spec is parsed. No job spec can lift any of it on its own: there is no privileged field. The one way past these defaults is a host device or socket the operator granted on the node (R17, R18), which a spec requests by name and can never define.
  • The strong posture has a name. hardening = "restricted" on a service requires a non-zero user, runs the task with no capabilities at all, and refuses any grant declared beside it; a service that owns a volume chowns it in a root init step first. task.read_only_rootfs = true mounts the image read-only on top of that. Neither is the default - the baseline exists so stock images start - but kanea plan warns on the weak shapes it replaces: an effective uid of 0, a writable root filesystem, and an image following a moving tag without opting into auto-update.
  • Disk hygiene is part of the driver. Image GC, build-cache caps across both content stores (containerd's and the rootless buildkitd user's) per-service log caps, and watermark alerts at 80% and 90%. One disk holds images, logs, state and volumes; pressure must never surprise the control plane.

Network driver: eBPF datapath

The datapath is Kanea's own: three small eBPF programs, a handful of pinned maps and plain netlink plumbing, all loaded and written by kanead from one object compiled into the binary. There is no network agent, no kvstore and no CNI. The loader is the standalone github.com/cilium/ebpf library: importing github.com/cilium/cilium would pull the Kubernetes client graph, which is the one dependency the project does not have. The programs are compiled ahead of time and committed, so go build needs no clang and the node needs no BTF: they read only UAPI context types, so there is no CO-RE and no vmlinux.h.

IP is identity

kanead allocates every alloc address from the node CIDR and every service VIP from the service CIDR, durably in the Store. Because the platform that hands out addresses is the same one that writes the kernel's maps, the identity map (alloc IP → {project, service}) is written by the allocator itself. There is no identity-allocation protocol, no label race and no settle window. The numeric project and service ids behind the maps are Store-allocated, monotonic and never reused, which is what keeps a pinned map meaningful across a kanead restart.

Attach is deny-closed by construction

Per alloc, in order: netns → identity map write → veth created with the host side down → policy programs attached at tc → addresses and static neighbours → link up → the host /32 route last. The first moment a packet can reach or leave the alloc, policy is already enforcing and identity is already resolved; a skipped step fails closed, because an identity miss is a drop. The deny-by-default guarantee is structural, not temporal: there is no unlabelled window to hold shut with retries, the way the previous design held its reserved:init state closed. Attach has no wait loop and completes in milliseconds.

Load balancing is connect-time

One cgroup connect4 program at the root cgroup, held by a pinned bpf_link, rewrites VIP:port → backend at connect(2); for host processes and containers alike, which is the one property the edge depends on: it dials a VIP with a plain dialer. There is no per-packet NAT and no conntrack entry per flow, and an established connection never consults a map: kanead can restart, or recreate every map, without touching live traffic. Backends update by generation flip: the new set is written under the next generation and one atomic map update commits it, so a concurrent connect() sees a complete old set or a complete new one, never a torn one. A VIP with no backends refuses at connect() rather than black-holing into a timeout. Service ports are TCP-only in v1, refused at plan otherwise.

Policy is SYN-gated map entries

The tc program on each alloc's host-side veth admits a source carrying the host identity (the edge's upstream dials, kanead's DNS replies and probes), a source in the same project, or a source named by an allow_from edge (R14), and drops every cluster-internal source the identity map does not know. A source outside the cluster CIDR carries no identity by construction: it is the internet answering a connection the alloc opened, un-NATed by conntrack on the way back in, and passes (v1.65); nothing unsolicited arrives that way, because the pod CIDR is unroutable from off-node and published ports terminate at the edge. The mirror-image rule guards the other direction: a packet leaving an alloc must carry a cluster source, so a forged external address cannot ride that pass. There is no policy file, no selector language and no translation step that could make a rule silently match nothing; policy is map entries the datapath enforces directly, and rules only ever union, so allow_from can never weaken the project default-deny.

One honest weakening

Enforcement is per connection attempt: TCP that is not a SYN passes, which is what lets cross-project replies flow without a conntrack. That is deliberately weaker than stateful tracking: an in-node ACK probe traverses the filter and is stopped only by the receiving stack's RST, and it is stated in the threat model rather than hidden. The upgrade to an LRU conntrack map is additive.

A second, small egress program is load-bearing for A10: it drops the cloud-metadata range 169.254.0.0/16 in the kernel with a per-alloc drop counter (not asserted in a policy file) along with any service-CIDR destination that escaped connect-time rewrite, and it counts per-endpoint traffic. Masquerade for routed pod traffic is one nftables rule in an owned kanea table, and since v1.65 the rule and net.ipv4.ip_forward are re-asserted every thirty seconds: a firewall reload that flushes the ruleset costs seconds, not a daemon restart. A drop policy another tool installs can still eat that traffic, on the forward hook and on the input hook alike - the second is how an alloc reaches the internal resolver, and a default-deny ufw eats it while the host's own dig keeps working, since every manager accepts lo unconditionally. kanea doctor detects and names both, passing once the ruleset carries an accept for the cluster CIDR, and kanea firewall prints the rules; a missing kanea table and forwarding turned off are findings of their own.

Internal DNS

Allocs resolve through kanead's own resolver, bound to the datapath's host anchor (the node CIDR's .1) and never a wildcard, on UDP and TCP alike. It answers <service>.<project>.kanea names authoritatively and forwards everything else upstream. The upstream list is, in order of precedence: the --dns-upstream flag; the server config's dns stanza (v1.66); the host's own /etc/resolv.conf, read once at startup, taken exactly as listed. That includes systemd-resolved's 127.0.0.53 stub, which is the only nameserver a stock Debian or Ubuntu server has: kanead forwards from the host's own namespace, so the stub is as reachable for it as for any other process, and a workload inherits the node's cache and DNSSEC posture with it (an alloc is never handed these addresses; its resolv.conf names Kanea's resolver). The stanza is how you pin resolvers on a node whose resolv.conf is DHCP's to rewrite:

# /etc/kanea/kanea.hcl
dns {
  upstreams = ["1.1.1.1", "10.0.0.53:5353"]  # a bare address gets :53
}

Entries are validated when the file parses; an empty list is refused by name; a stanza that meant "no upstreams" would silently turn external resolution into SERVFAIL. An explicit --dns-upstream wins over the stanza and the daemon says so in its log, the same precedence every half of the server config follows. The stanza's own reference, and how to apply an edit, is in Node configuration.

A host firewall is what usually breaks this

An alloc's query to the resolver is a new inbound connection to the host on a veth, so it crosses the input hook like any other arrival and a default-deny ufw or firewalld eats it. Nothing about it looks like a firewall problem: the host's own dig @10.244.0.1 keeps answering, because every manager accepts lo unconditionally, and the resolver's logs show a healthy server nobody is reaching. kanea doctor names it and kanea firewall prints the rules.

TCP is served beside UDP (v1.86) because three things set the truncation bit - an internal answer over 512 bytes, an upstream reply that filled the read buffer, and an upstream reply that already carried TC - and each one tells a client to retry over TCP. Serving only UDP made every one of those a dead end. The TCP half is bounded the way the UDP half is: a hard concurrent-connection cap whose refusals are counted, deadlines on every connection, and a TCP query forwarded over TCP, so the retry can carry what the datagram could not.

Datapath state is derived state. Programs, maps and the cgroup link are pinned under /sys/fs/bpf/kanea with a schema stamp; a kanead restart leaves the dataplane untouched, and a stamp mismatch recreates and repopulates the maps inside the first reconcile pass: safe precisely because established flows bypass them. Nothing under the pin root is ever backed up.

The edge

A Go reverse proxy: host-based L7 routing to service frontends, TLS termination with Let's Encrypt certificates, WebSocket and gRPC support, HTTP→HTTPS redirects and security headers.

The middleware chain runs in a fixed order, per service, from the spec's expose block:

Host match → IP allow/deny → rate limit → header transforms → upstream proxy

All of it is validated at kanea plan time and fails closed. An ingress control that silently does nothing at runtime is worse than one that is absent.

Hardening is mandatory rather than tunable: read/header/idle timeouts and header size caps against slowloris, per-route upstream timeouts, bounded connection pools, client-supplied X-Forwarded-* stripped, and an unknown Host answered with 404, which is also the DNS-rebinding defence for the co-hosted API.

Because it already sits in the request path, the edge is the primary source of L7 metrics for exposed services: requests per second and latency percentiles at no extra data-plane cost. The datapath's own map counters cover east-west.

ACME runs in kanead, never in the edge: obtaining a certificate means writing one, and the edge does not write.

Resource isolation

The control plane must survive anything a workload does. Enforcement is cgroups v2, arranged as two sibling slices:

/sys/fs/cgroup
├── kanea.slice                # kanead, kanea-edge (+ containerd and buildkitd, via their own units)
│     memory.min      = system_reserve_memory     # kernel-protected floor (default 256 MiB; build nodes raise it)
│     memory.swap.max = 0                         # the floor is RAM, not swap
│     cpu.weight      = 10000
│     OOMScoreAdjust  = -900                      # the global OOM killer picks workloads first
└── kanea-workloads.slice      # every alloc lives under this one parent
      memory.max      = total RAM − system_reserve_memory
      memory.swap.max = 0
      cpu.weight      = 100
      └── per-alloc: memory.max · cpu.max · pids.max from the spec's resources block
"Memory lock" means guarantee, not mlock

Calling mlockall on a Go control plane is rejected outright: the GC grows the heap unpredictably and RLIMIT_MEMLOCK turns pin overflow into hard allocation failure; the lock itself could crash kanead. The guarantee comes from memory.min (the kernel refuses to reclaim the floor under pressure), the OOM score, and no swap in the slice.

Per-alloc limits are enforced where declared. Omit resources (or a field of it) and the alloc is unbounded: all cores, all allocatable memory, capped only by the workload parent's collective ceiling, which, with the control-plane floor, is the isolation that actually protects the platform. A default pids.max applies to every alloc; resources { pids = 2048 } declares a different value the same way the other limits are declared (the cap is always on, omission means 256). Declared resources are also the admission units: kanea plan renders the workload budget and apply refuses a total above it unless the operator has explicitly enabled oversubscription.

Metrics and autoscaling

Three scrapers feed one bounded in-memory time series: containerd cgroup metrics (CPU, memory), the edge (requests per second, latency percentiles for exposed services), and the datapath's own per-CPU map counters (east-west flows and drops, on by default). Roughly 27 MiB at the 2 000-alloc target, and there is a test that says so.

"No data" is never zero

A missing metric and an idle service lead to opposite decisions, so the distinction is preserved end to end: in the time series, in the evaluator, in the Prometheus exporter and in the dashboard. Each layer has a test asserting it.

The evaluator applies the scaling block with guardrails and a circuit breaker, and the budget is 20 seconds from a sustained breach to a decision: a 15-second averaging window (three samples at the 5-second scrape resolution) plus one evaluation tick. A large spike decides sooner.

The autoscaler is not a second scheduler. It writes one number (the desired count) and the reconciler converges. kanea scale uses the same route, which is why manual and automatic scaling cannot disagree about mechanism.

State replication and restore

  • Change data capture. Store mutations produce change segments; snapshots and segments are replicated to any S3-compatible bucket. The client is hand-written against the REST API (SigV4 and four verbs) and CI exercises it against MinIO in both addressing styles, path-style and virtual-hosted.
  • Archives are chunked AEAD, with keys HKDF-derived from the master key rather than being it. The last chunk is sealed under different additional data, which is the only thing standing between a truncated snapshot and a restore that decrypts cleanly into half a platform.
  • Manifests are unencrypted and hashes are over the ciphertext, so someone holding the bucket and no key can still see what is there and whether it is intact, which is the question you ask before going to find the escrowed key.
  • Restore is staged, never performed in place. It can be requested over the API, and the daemon performs it at the next start, before anything opens the Store. The API has no method that restores; that is the interface, not a check.
  • The replication cursor is derived from the sink, never stored: writing it to the Store would emit a change that needs shipping, which would write it again.
  • The destination can change at runtime; the dashboard's Settings page, or PUT /v1/settings/backup. A new destination is probed with a real test write before anything commits, so a typo is a refusal with the old replication untouched; the old destination receives one final segment ship, the new one an immediate full snapshot. The unit's flags remain the seed: a settings record, once written, wins, and deleting it reverts.

The master key is generated by kanea init, shown once, and must be typed back; it is discarded if that fails. Without it, every archive is unreadable. The DR runbook is worth reading before you need it.

API, dashboard and MCP

The API is REST plus one multiplexed WebSocket, and every route is deny-by-default: including the WebSocket and every MCP route. The exceptions are enumerable: /login, the ACME challenge path, and /healthz - and even the last one keeps its answers short. Unauthenticated, it returns the status and the OIDC issuer (a login screen needs that before it holds a credential) and nothing else; version, PID, store index, listen address and uptime need an identified caller. A bad token still gets the slim 200, so a load balancer forwarding a stale Authorization header cannot flap a health check.

There is exactly one authentication mechanism that is not the §13 one: the git webhook route, which uses a per-project HMAC over the raw body, rejects replays, and is audited. It never deploys from the request: it marks the project, and the sync loop re-reads the source over Kanea's own credential.

The dashboard is a React SPA embedded in the binary with go:embed and served by the daemon. A service page charts CPU, memory, request rate and p95 over a real time axis, seeded on the first frame of the subscription that carries them (since v1.79) so a chart draws its window on arrival rather than growing one point per scrape, and an empty panel distinguishes still arriving from no samples yet, which is also the honest answer for the ten seconds after a restart, during which no rate can exist. It streams logs live (virtualized, filterable, copy, download, and a full-screen view that keeps your place in the buffer across the toggle), shows a restart or deploy as rollout progress (the same spec-hash rule the planner uses, served on the wire since v1.64) and gives an admin a shell into any running alloc, over the same exec websocket the CLI uses: the CSRF token rides the handshake as a subprotocol, since a browser cannot set the header there. Its Settings page shows the node's configuration and lets an admin change what changes at runtime: the backup destination and the notification channels (node-wide defaults and per-project, each with a test button); plus accounts, API tokens and the audit log, one tab each; what belongs to the unit (listen address, subnets, DNS, the port policy) is shown read-only with a note saying so. A pinned Updates page holds the rest of the node's software: check for or install a Kanea release (the upgrade flow, run by the node on itself), and read the OS's pending packages, security count, list freshness and reboot flag beside the manifest-pinned host components - reported, never installed: the page names the apt command to run on the node. The MCP server exposes 24 tools in three tiers (read, mutate, destructive) over stdio and streamable HTTP.

Why MCP tools have no privileges of their own

Every tool reaches the platform by making an HTTP request against the API's own handler. A tool's only verb is "send this request", so it can never be more privileged than the credential its caller presented: nothing in the MCP package may hold a Store, a secrets store, or an auth store. That is what makes "no side channels" structural rather than a rule somebody has to remember.

There are also no secret tools at any tier. The safety requirement is that no tool returns a secret value; the implementation goes further and gives an agent no secrets verb whatsoever, and a test fails if one appears.

Exposing the API and dashboard

The API and the dashboard are one listener: the dashboard is served by kanead on the same address the REST API, the WebSocket and the MCP HTTP transport bind, in front of kanea-edge entirely; neither is a service, and neither takes an expose block. The consequence worth internalising is that the dashboard's URL carries the listener's port - https://<name>:8600 by default, not :443, which belongs to the edge and has no route to the dashboard. There is a worked example in the bind stanza's own section. Exposing them means setting where that listener binds, and there are two ways: the --listen/--listen-cert/--listen-key flags (asked by kanea init and rendered into the unit), or (since v1.61) a bind stanza in the server config, which survives re-runs and makes moving the listener later an edit plus systemctl restart kanead, never a re-init:

# /etc/kanea/kanea.hcl
bind {
  api_addr   = "192.168.1.10:8600"
  api_tls    = "self-signed"       # acme | self-signed | provided | plaintext
  # api_domain = "kanea.home.example"  # required by acme; names a self-signed cert
  # api_cert   = "/etc/kanea/api.pem"  # provided only; always with api_key
  # api_key    = "/etc/kanea/api.key"
}
FieldTypeNotes
api_addrstringThe host:port the API and dashboard bind. Required by every other field in the stanza: TLS with nothing to serve it on is a parse error.
api_tlsstringacme, self-signed, provided or plaintext; the same mode vocabulary services use. Unset resolves at the daemon: a declared pair means provided, loopback means plaintext, and anything beyond loopback refuses there.
api_domainstringRequired by acme (an IP cannot hold an ACME certificate). For self-signed it names the certificate, and it is required when api_addr binds every interface (":8600", 0.0.0.0): otherwise the certificate would have no name at all.
api_cert, api_keystringYour own PEM pair, for provided only. Always together, and refused beside a managed mode or plaintext; a control that cannot act is refused, never silently dropped.

What each mode gives you:

  • self-signed: a certificate from the node's own CA, the one kanea ca show installs on your devices, with a real IP SAN when the address is bare. Renewed automatically. The right default on a home network.
  • acme: a Let's Encrypt certificate for api_domain, issued and renewed by the same account and pass that serve your services. Needs --acme-email configured on the daemon, refused at startup without it.
  • provided: your own api_cert/api_key pair, for a certificate something else manages.
  • plaintext: explicit HTTP. Allowed beyond loopback because you typed it, logged loudly, and it implies the insecure-cookie posture; a Secure cookie over plain HTTP is a login that silently fails.

Precedence: an explicit --listen always beats the stanza, and --listen none keeps the node socket-only regardless of the file. When the stanza is declared and --listen was not passed, kanea init skips the listen question and renders no listen flags into the unit: the file owns the listener. On the flags path, a non-loopback --listen with no explicit --listen-cert/--listen-key pair no longer refuses: init provisions a default 10-year self-signed pair at /etc/kanea/api.crt/api.key (minted once, never re-minted) and points the daemon at it. The one refusal left is an unspecified host (0.0.0.0): a SAN needs something to name, so bind a specific address or set api_domain in the stanza. The stanza's full field list and every refusal it can raise are in Node configuration → bind.

One consequence worth knowing when adopting the stanza on an existing node: an init run from before it existed rendered --listen 127.0.0.1:8600 into the kanead unit, and that flag shadows the file; the stanza appears to do nothing. Remove the flag from the unit's ExecStart, then systemctl daemon-reload and restart; the troubleshooting page has the steps.

Node configuration

/etc/kanea/kanea.hcl is the node's own file: the settings that belong to whoever owns the machine rather than to whoever writes a job spec. A spec can reference what this file permits and can never add to it, which is the whole point of the split - GitOps deploys specs automatically, so anything a spec could declare, anyone who can push to a synced repository could declare.

It does not exist by default, and that default is "off": a node with no file grants nothing, allows no host paths, and takes its listener and resolvers from flags. kanea init creates the /etc/kanea directory but never writes this file, on a first run or a re-run.

The file

Four properties decide how it behaves, and each one surprises somebody:

  • It is read once, at daemon start. One stat, never a poll. There is no reload, no SIGHUP and no watcher, so an edit means a restart - see applying a change. A grant is a decision, not a thing that rotates behind your back.
  • Absent is off, but malformed is fatal. A missing file is every setting's zero value and no error. A file that does not parse, or breaks a stanza's rules, stops the daemon starting. There is deliberately no keep-last-good and no partial load: a policy file half-applied is worse than one refused.
  • It is trust-checked before it is parsed. It must be a regular file (not a symlink, not a directory), owned by root or by the daemon's own uid, and writable only by its owner. Anything else refuses startup. World-readable is fine; this is policy, not a secret.
  • An unknown stanza is a warning; an unknown attribute is an error. See the stanza list for why the asymmetry is intentional.

Creating it, if it is not there:

sudo install -o root -g root -m 0644 /dev/null /etc/kanea/kanea.hcl
sudo $EDITOR /etc/kanea/kanea.hcl
sudo systemctl restart kanead

What may go in it

Seven stanzas are read. Everything else is accepted and warned about by name.

StanzaRepeats?What it decides
bindonceWhere the API and dashboard listen, and how that listener gets its certificate.
storageonceWhich host directories a host volume may mount.
dnsonceWhich resolvers the internal DNS forwards external names to.
imagesonceThe node's default for where a service's images may come from, and the image GC's gc block.
variablesonceNode-wide defaults for spec variables.
devicemanyA named device grant, and the projects that may claim it.
socketmanyA named socket grant, and the projects that may claim it.
Unknown stanza: warning. Unknown attribute: error.

A top-level block this version does not read is loaded, ignored, and named at startup in a server config carries stanzas this version does not read warning - never silently swallowed. That is what lets the file carry settings a future version will read. But an unknown attribute inside a stanza that is read is a hard parse error, because there it is almost always a typo, and a silently ignored allowed_host_path would be a security control that quietly did nothing.

bind: the API and dashboard listener

Putting the listener here instead of in the unit means moving it later is an edit plus a restart, never a re-init, and a kanea init re-run will not overwrite it. When api_addr is declared and --listen was not passed, init skips the listen question and renders no listen flags into the unit - a unit that repeated the file's answer would turn the file off.

# /etc/kanea/kanea.hcl
bind {
  api_addr   = "192.168.1.10:8600"
  api_tls    = "self-signed"       # acme | self-signed | provided | plaintext
  # api_domain = "kanea.home.example"  # required by acme; names a self-signed cert
  # api_cert   = "/etc/kanea/api.pem"  # provided only; always with api_key
  # api_key    = "/etc/kanea/api.key"
}
FieldNotes
api_addrThe host:port to bind. Required by every other field: TLS with nothing to serve it on is a parse error.
api_tlsacme, self-signed, provided or plaintext - the same vocabulary services use. What each one gives you is in Exposing the API and dashboard. Unset resolves at the daemon: a declared pair means provided, loopback means plaintext, anything beyond loopback refuses.
api_domainRequired by acme (an IP cannot hold an ACME certificate). For self-signed it names the certificate, and it becomes required when api_addr binds every interface.
api_cert, api_keyYour own PEM pair, for provided only. Always together.
edge_http, edge_httpsAccepted, warned about, and not read. They are part of the fuller design sketch; setting one today does nothing but add its name to the startup warning.

Every contradiction is a parse error rather than a silent preference:

If you writeIt says
Only one of api_cert/api_keythey go together
A pair, or any TLS field, with no api_addrTLS with no api_addr to serve it on
provided with no pairprovided needs api_cert and api_key
plaintext beside a pairplaintext beside a TLS pair drops the pair; remove one
acme or self-signed beside a pairit issues its own certificate; remove the pair
acme with no api_domainan IP cannot hold an ACME certificate
self-signed on 0.0.0.0 with no api_domainbinds every interface; set api_domain so the certificate has a name
Any other api_tls valueuse one of acme, self-signed, provided, plaintext

A worked example: the dashboard on a public name

The most common thing anyone wants from this stanza - a real certificate on a real name, replacing the self-signed pair kanea init provisions:

# /etc/kanea/kanea.hcl
bind {
  api_addr   = "0.0.0.0:8600"
  api_tls    = "acme"
  api_domain = "kanea.apps.example.com"
}

The dashboard is then at:

https://kanea.apps.example.com:8600
The port is part of the URL, and leaving it off fails confusingly

https://kanea.apps.example.com (no port) goes to :443, which is kanea-edge, a different process that has never heard of the dashboard. The dashboard is served by kanead itself and is not a service, so no route to it can exist and the edge has nothing to forward to.

What you get is not a 404. On a node with no exposed services the edge holds no certificates at all, so the TLS handshake fails outright - SSL_ERROR_INTERNAL_ERROR_ALERT in Firefox, tlsv1 alert internal error from curl - which reads like a broken certificate rather than a wrong port. Check the port before you debug the certificate.

Three things have to be true for that certificate to issue, and only the first is in this file:

  1. An A record for kanea.apps.example.com pointing at the node. The name is validated by the CA connecting to it, so it must resolve before the first attempt.
  2. --acme-email on the daemon. It is a flag, and kanea init does not write it into the unit, so it goes in a drop-in. Without it nothing is ever requested from a CA, and the log says so at every start. Note there is no acme stanza in this file: writing one is accepted, ignored, and warned about by name, which is a confusing way to discover the flag exists.
  3. Port 80 reachable, with kanea-edge running. The edge answers /.well-known/acme-challenge/ on :80 even though the certificate is for kanead's listener, so a stopped edge means no dashboard certificate.
Two ways to get this wrong that look like something else

Binding loopback does not put the dashboard behind the edge. api_addr = "127.0.0.1:8600" makes it reachable only from the node itself; nothing bridges :443 to it, because a route's upstream is always a service VIP.

Until the first issuance the handshake fails by name, rather than serving something weaker. The listener binds a few seconds before the certificate arrives, so a restart has a brief window where the dashboard is hard-unreachable, and a node where ACME cannot issue stays that way with the real reason only in journalctl -u kanead. That is deliberate: a wrong certificate is worse than a refused connection.

Wanting the bare name with no port means giving the dashboard :443 and moving the edge with kanea edge --https :8443. Two processes cannot share a port, so that is the whole decision: every service you deploy then answers on :8443 instead. It is a reasonable trade on a node whose apps live elsewhere, and a bad one on a node that hosts anything.

storage: host-volume allowlist

A host volume mounts a directory that already exists on the node, and whether it may be mounted is not the spec's decision. The allowlist is empty by default, so host volumes do nothing until this stanza names a parent.

# /etc/kanea/kanea.hcl
storage {
  allowed_host_paths = ["/srv/kanea", "/mnt/media"]
}

The check runs after symlink resolution, on the object rather than the spelling. Three things are refused at startup: a prefix that is not absolute, a prefix that does not exist, and "/" itself - which "allows the entire filesystem; list the directories you actually intend to share". An explicit empty list is legal and means the same as no stanza: no host volumes.

A spec then mounts with storage "x" { type = "host" path = "/srv/kanea/x" }. A path outside every prefix fails the alloc. A path that does not exist is refused too, so a typo cannot become a silently empty volume - unless the spec sets create = true, which makes Kanea create it, still only inside a prefix you allowed. Creating is not owning: Kanea never chowns a host volume.

device and socket: passthrough grants

A job spec names a grant, never a path. There is no field in the spec for a device node or a socket path: not a validated one, none at all, because the node holds the mapping and the spec only asks for it by name.

# /etc/kanea/kanea.hcl
device "gpu" {
  nodes = ["/dev/dri/card0", "/dev/dri/renderD128"]  # a list; required
  allow = ["media"]                                 # projects that may claim it; required
  # mode = "rw"                                        # r, w, m; default "rw"
}

socket "containerd" {
  path  = "/run/kanea/containerd.sock"              # a single string; required
  allow = ["ops"]
}
AttributeOnNotes
The block labelbothThe grant name a spec references. Must be a DNS-1123 label, or no spec could name it. Duplicates within a kind are refused.
nodesdeviceA list of absolute device paths. Required and non-empty.
pathsocketA single string. Required, absolute, no .., not /.
allowbothRequired and non-empty: the projects that may claim the grant. Each must be a valid project name. There is no projects attribute - this is it.
modedeviceOptional cgroup permissions, some combination of r, w and m. Defaults to "rw"; m (mknod) only if you write it.

A spec claims one with device "dri" { grant = "gpu" } or socket "rt" { grant = "containerd" mount_path = "/run/containerd.sock" }. Paths are resolved and type-checked per alloc, not at load, so a grant naming a device that is absent or is not a device fails that alloc with a message naming the grant, rather than stopping the daemon. A transcoder silently running without its GPU would look healthy and do the wrong thing, so it fails instead. A grant the node does not hold is refused with a list of the ones it does.

A socket grant is root on this node

A container holding the runtime socket can start other containers without the hardening defaults. There is no containment story and none is claimed: nosuid,noexec,nodev restrict the filesystem entry, not the protocol spoken over it. The control is that granting it happens on the machine, in a file no spec author can write, and names one project.

dns: upstream resolvers

The internal resolver answers <service>.<project> names itself and forwards everything else. By default it forwards where the host forwards - every nameserver in /etc/resolv.conf, in the file's own order, read once at startup. This stanza pins that list instead, which is what you want on a node whose resolv.conf is DHCP's to rewrite:

# /etc/kanea/kanea.hcl
dns {
  upstreams = ["1.1.1.1", "10.0.0.53:5353"]  # a bare address gets :53
}

Each entry must be an address or a host:port pair, checked when the file parses. An empty list is refused by name: a stanza meaning "no upstreams" would silently turn every external name into SERVFAIL, so if that is what you want, remove the stanza. Allocs never see this list - their resolv.conf names Kanea's resolver alone, and it forwards on their behalf from the host's own namespace.

images: the default pull policy

Where images may come from for a service that does not say (R33). The one case this exists for is a node that is not supposed to reach a registry at all - air-gapped, or with images preloaded by an offline bundle - where the default's slow pull-and-fail blames a registry the node was never going to contact:

# /etc/kanea/kanea.hcl
images {
  pull_policy = "never"        # "if-not-present" (default) | "never"
}

Precedence is the file's usual one: --image-pull-policy on kanead wins and says so in the log, this stanza is next, and if-not-present when neither - which is what every node did before the field existed. A service's own task.pull_policy wins over all of it. always is refused here: it means per-service auto-update, and turning that on for every service on a node is not a default anybody asked for.

images.gc: cleaning up old images

Nothing else ever deletes an image, so every deploy, auto-update pin and rollback leaves its predecessor behind. The image GC collects them - on by default, and this block tunes or disables it:

# /etc/kanea/kanea.hcl
images {
  gc {
    enabled  = true      # --image-gc off on kanead also disables it
    interval = "12h"    # sweep cadence; floor 10m
    min_age  = "48h"    # younger images are never touched; floor 1h
    keep     = 2         # newest N per in-use repository, as rollback material
  }
}

An image is deleted only when all three rules agree: nothing references it (no service's declared image, pinned digest, rollback target or init step, and no running alloc), it is older than min_age, and it is not among the newest keep of a repository something still uses. A node whose default pull_policy is "never" refuses to collect at all: preloaded images cannot be re-pulled, so deleting one is unrecoverable. kanea images lists what the node holds with each image's in-use verdict, and kanea images --clean runs one sweep now (admin, audited).

variables: node-wide spec defaults

Values every spec on this node can reference as ${name}, so one node's domain or storage root does not have to be repeated in every file:

# /etc/kanea/kanea.hcl
variables {
  domain      = "home.lan"
  media_root  = "/srv/kanea/media"
}

Precedence is node under spec under caller: a spec's own variables block wins over the node's, and --var on the command line wins over both. Values are strings, numbers or bools - not lists or objects. A definition may reference the built-ins but never a sibling, so there is deliberately no ordering or cycle story. The R2 built-ins and service are reserved names.

Never put a secret here

GET /v1/vars serves this stanza to any authenticated caller, which is exactly why nothing secret may live in it. Secrets are secret: references, resolved per alloc into a tmpfs file; see Secrets.

Applying a change

The whole file is read once at startup, so there is one apply step and it is the same for every stanza:

sudo systemctl restart kanead

systemctl daemon-reload is not part of this. That is for changes to the systemd units, which this file is not. Running it after editing kanea.hcl is harmless but does nothing.

What a restart costs

Less than people expect, and it is worth knowing before you hesitate over a production node:

  • Running workloads keep running. The unit carries KillMode=process precisely so a control-plane restart does not take the containers with it; their shims outlive it.
  • Ingress keeps serving. kanea-edge is a separate process with no After= relationship in either direction, and it serves the last route snapshot it read from disk whether or not kanead is up.
  • What pauses is change, for the second or two it takes to come back: deploys, scaling decisions, certificate renewal, the API and the dashboard. See the process list for the same split across every daemon.
A bad edit stops the daemon starting

Malformed is fatal by design - no keep-last-good, no partial load - so a file with a typo means kanead refuses to come back, with the parse error naming the file and line in journalctl -u kanead. Fix the file and restart again. If you need the node up first and the config sorted out afterwards, add --config off to the unit's ExecStart (systemctl edit kanead), which starts the daemon as though the file did not exist, so you can fix the file with the dashboard and the API back up.

There is no offline validator for this file today: the restart is the check. On a node you cannot afford to have refuse a start, keep a known-good copy beside it before editing.

Confirming it took effect

A few lines in journalctl -u kanead answer nearly every "the file seems to do nothing" question:

Log lineMeans
server config loadedThe file was found, trusted and parsed. It carries the path.
server config carries stanzas this version does not readSomething in the file is being ignored, and the message names it. A misspelled stanza shows up here rather than as an error.
dns upstreams from server config
api/dashboard listener from server config
That stanza won and is in effect. These are the positive confirmations, and their absence is as informative as their presence.
… is not consulted (--flag wins)A flag on the unit is overriding that half of the file - there is one of these for --dns-upstream, --image-pull-policy, --listen, --allowed-host-paths and --passthrough-config. This is the single most common cause of a stanza appearing to do nothing; see precedence.

Flags, the file, and precedence

Every half of the file has a flag that overrides it, and the rule is the same throughout: an explicit flag wins, and the daemon says so in its log. Nothing is ever merged - whichever source supplies a setting supplies all of it, so you never end up with an address from one place and its certificate from another.

StanzaFlag that winsTurn it off withIf neither is set
bind--listen (plus --listen-cert/--listen-key)--listen noneUnix socket only
storage--allowed-host-paths--allowed-host-paths offNo host volumes
device, socket--passthrough-config <file>--passthrough-config offNo grants
dns--dns-upstream--dns-listen off (disables the resolver)The host's resolv.conf
images--image-pull-policyn/a - there is always a policyif-not-present
variablesnone - file-only--config offNo node variables
the whole file--config <file>--config off/etc/kanea/kanea.hcl if it exists
Two traps in that table

The disable words are not the same. It is --listen none but off everywhere else. none was chosen for the listener because it matches what kanea init asks you to type at the prompt.

--config off disables all six stanzas, not just the security ones. The file is never even stat-ed, so a present and perfectly good dns or variables stanza goes with it. And variables is the one half with no flag of its own, which is why the daemon probes the file whenever --config does not say off.

A flagged --passthrough-config file gets the same trust check as kanea.hcl itself: arriving by argv does not make a policy file more trustworthy. Note also that --allowed-host-paths does not merge with the stanza; it replaces it, and the daemon logs server config storage stanza is not consulted when both are present.

A complete file

Every stanza this version reads, in one file:

/etc/kanea/kanea.hcl
# /etc/kanea/kanea.hcl: the node's, never the repository's.
# Read once at startup. Edit, then: sudo systemctl restart kanead

bind {                                       # where the API and dashboard listen
  api_addr = "192.168.1.10:8600"
  api_tls  = "self-signed"                  # acme | self-signed | provided | plaintext
}

storage {                                    # parents that `host` volumes may mount from
  allowed_host_paths = ["/srv/kanea", "/mnt/media"]
}

dns {                                        # pin resolvers; omit to follow the host's resolv.conf
  upstreams = ["1.1.1.1", "9.9.9.9"]
}

variables {                                  # node-wide spec defaults; never secrets
  domain     = "home.lan"
  media_root = "/mnt/media"
}

images {                                     # the node default; a task's own pull_policy wins
  pull_policy = "if-not-present"
  gc {                                       # unused-image cleanup; on by default
    interval = "12h"
    min_age  = "48h"
    keep     = 2
  }
}

device "gpu" {                               # claimed as: device "dri" { grant = "gpu" }
  nodes = ["/dev/dri/renderD128"]
  allow = ["media"]
}

socket "containerd" {                        # root-equivalent for whoever holds it
  path  = "/run/kanea/containerd.sock"
  allow = ["ops"]
}

Names and DNS

Kanea deals in two separate name spaces, and almost every DNS question comes from them being mistaken for one. One is internal, resolved by Kanea, and needs nothing from you. The other is public, resolved by the world's DNS, and is the only part where you create records.

InternalPublic
Looks likeweb.shop.kaneaweb.shop.apps.example.com
Resolved byKanea's own resolverYour DNS provider
Answers withThe service's virtual IPThe node's public address
ReachesThe allocs directly, east-westThe edge on :443, then the allocs
Records you createNone. It is automaticOne wildcard, usually
Who can use itAllocs on this nodeAnything that can reach the node

Internal names, between services

Every alloc gets a generated /etc/resolv.conf pointing at Kanea's resolver, with a search path that makes the short forms work:

nameserver 10.244.0.1
search shop.kanea kanea
options ndots:1 timeout:2 attempts:2

So from a service inside project shop:

You writeIt resolvesWhen to use it
dbdb.shop.kaneaA service in the same project.
cache.datacache.data.kaneaA service in another project.
db.shop.kaneaitselfWhen you would rather be explicit.

The address that comes back is the service's virtual IP, and the connection is load-balanced across its healthy allocs at connect() by the datapath - there is no proxy in the path and no client-side round-robin to get wrong. The resolver lives at the node CIDR's .1 (10.244.0.1 by default, moving with --node-cidr), and it is an anchor address, never a wildcard.

Resolving is not reaching

A cross-project name resolves for anyone, because DNS is not the security boundary. The datapath is, and it is deny-by-default between projects: the connection is dropped unless the callee lists the caller in allow_from. This presents as a name that resolves fine and a connection that times out, which reads like a network fault rather than a policy decision. If a cross-project call hangs, check allow_from on the service being called before you look at DNS.

The base domain

An exposed service's public name is <service>.<project>.<base-domain>, so --base-domain apps.example.com gives the web service in project shop the name web.shop.apps.example.com. A spec can name its own domains instead, with expose { domains = [...] }; the base domain is what it falls back to when it does not.

kanea init does not set this, and the unit is rewritten

--base-domain, --acme-email and --tls-default are daemon flags that init never renders into kanead.service - it writes only seven, and these are not among them. Without a base domain a service has no generated FQDN at all, so this is the step between "deployed" and "reachable by name".

Set them in a drop-in, not by editing the unit: kanea init rewrites kanead.service on every run, so a direct edit is lost the next time you regenerate units, while a drop-in survives. sudo systemctl edit kanead and add:

[Service]
ExecStart=
ExecStart=/usr/local/bin/kanea agent --data-dir /var/lib/kanea \
  --log-dir /var/log/kanea/allocs --network ebpf \
  --node-cidr 10.244.0.0/24 --cluster-cidr 10.244.0.0/16 \
  --edge-group kanea-edge \
  --base-domain apps.example.com --acme-email you@example.com

The empty ExecStart= is required: systemd appends to that directive otherwise, and a unit with two ExecStart lines fails to start. Copy the existing line from systemctl cat kanead and add to it, rather than retyping it from here - your node's flags may differ. Then sudo systemctl daemon-reload && sudo systemctl restart kanead.

The records you create

One wildcard covers every service, present and future, which is why it is the usual answer:

RecordTypePoints atCovers
*.apps.example.comAThe node's public IPv4Every generated service name
*.apps.example.comAAAAThe node's public IPv6The same, for v6 clients
kanea.example.comAThe node's addressThe dashboard, if you gave it a name
shop.example.comAThe node's public IPv4One custom domains entry

The AAAA record is about the node's own public address and has nothing to do with the opt-in container IPv6 feature: the edge binds :443 on whatever the host has, so a node with public v6 serves v6 clients whether or not allocs have v6 addresses.

The dashboard is not a service and takes no expose block, so it is not covered by the wildcard unless you happen to name it under the base domain. Give it a name with bind.api_domain and point a record at the node; without one it is reached by IP, and its certificate is the self-signed pair kanea init provisioned. Remember that you reach it at the listener's port - https://kanea.example.com:8600 - because the edge, which owns :443, has no route to it.

Nothing needs a record per alloc, and nothing needs a record that changes when you scale or redeploy: the edge routes by Host header and the datapath handles the rest, so the DNS side is static once it is set up.

Certificates and the challenge

Which ACME challenge Kanea can use decides what you have to arrange, and the two are not interchangeable:

HTTP-01 (default)DNS-01
NeedsPort 80 reachable from the internet, and the name already resolving to the nodeAn RFC 2136 nameserver accepting TSIG-signed updates
ConfigureNothing beyond --acme-email--acme-dns-server and the TSIG flags
Wildcard certificatesImpossible - there is no single host to reachYes
Works on a private nodeNoYes

HTTP-01 is the default and needs no configuration: the edge serves the challenge path on :80, and kanead reaches its own edge first to confirm the challenge is actually being served before it asks the CA to try (--acme-verify-url). The record has to exist and point at the node before issuance, so create DNS first, deploy second.

DNS-01 is RFC 2136 dynamic update, with TSIG:

--acme-dns-server ns1.example.com:53
--acme-dns-zone example.com                      # default: the challenge name's parent
--acme-dns-tsig-key kanea-updates
--acme-dns-tsig-secret secret:shared/tsig        # a reference, never the literal
--acme-dns-tsig-algorithm hmac-sha256.

There is no provider integration - no Cloudflare, Route 53 or Azure DNS plugin. RFC 2136 is the whole DNS-01 story, which suits BIND, Knot and PowerDNS, and means a provider that only offers a proprietary API needs a zone you run yourself, or a CNAME delegation of _acme-challenge into one. The TSIG secret is required whenever --acme-dns-server is set, and must be a secret: reference: a literal is refused by name, and so is an absent one, because unsigned dynamic updates would let anyone on the network answer a challenge.

When Kanea switches to a wildcard

Certificates are per service until a node has 20 generated names, at which point it collapses them into one wildcard for the base domain - but only if a DNS-01 solver is configured, since HTTP-01 cannot validate a wildcard. The number comes from Let's Encrypt's limit of 50 certificates per registered domain per week, and every redeploy that changes a name spends one.

Custom domains are never collapsed. A name you declared in expose { domains } is somebody else's zone, and Kanea has no standing to ask a CA for *. of it. Use --acme-directory staging while you are getting this working: the staging CA has far looser limits, and its certificates are not trusted, which is the point.

Private networks and homelabs

No public name, no port 80 from the internet, and still real HTTPS. Two routes, and the difference is whether a browser has to be taught to trust anything:

  • The node's own CA (--tls-default self-signed). Point a wildcard record on your local resolver - a Pi-hole, a router, an internal BIND - at the node, then install the CA once per device with kanea ca show > kanea-ca.crt. No CA to reach, no rate limit to spend, works entirely offline. The cost is one trust-store install per device, and devices you do not administer cannot be taught.
  • Real certificates over DNS-01. If you own a public domain and can run RFC 2136 updates against it, the node gets publicly-trusted certificates for names that resolve only on your LAN. Nothing has to be reachable from the internet, because the challenge is answered in DNS rather than over HTTP. Every device trusts it with no setup.

Either way the record points at the node's LAN address, which is ordinary split-horizon DNS: the name resolves internally and either does not resolve externally or resolves elsewhere. Kanea does not care which, because it never resolves its own public names - the edge routes on the Host header it is given.

Prefer not to deal with names at all? Publish a port and reach the service at http://<node>:8096. That path needs no DNS and no certificate, and ip_restriction bounds who may use it.

Job spec reference

Specs are HCL v2. If you have written Nomad job files the shape will be familiar, though the vocabulary is Kanea's. Every rule below is enforced when the file is parsed (errors carry file, line and column) and kanea plan is where you see them.

The minimum

A service needs an image and nothing else. This is a complete, valid spec:

spec_version = 1

project "demo" {}

service "web" {
  project = "demo"
  task "app" {
    image = "nginx:1.27-alpine"
  }
}

Everything on this page is additive to that.

Top level

A spec file contains spec_version and any number of project, service and storage blocks, plus an optional variables block of shared values. Multiple files applied together are parsed as one set, and file order is irrelevant: a service may reference a project declared in another file.

FieldTypeNotes
spec_versionnumberCurrently 1. Future spec revisions are gated on this field, which is how an upgrade can tell an old file from a new one.
project "<name>"blockZero or more. Name is a DNS-1123 label.
service "<name>"blockZero or more.
storage "<name>"blockZero or more. May also be declared in the server config.
variablesblockOptional. Shared values the rest of the spec references as ${name} (R30).

variables

Declare a value once, reference it anywhere as ${name}, or as a bare identifier where HCL takes an expression, like count. The node may supply defaults from a variables stanza in its own /etc/kanea/kanea.hcl; the spec's block wins on a collision, and pipeline-supplied values (${GIT_SHA_SHORT} and friends) sit above both.

variables {
  domain   = "shop.example.com"
  replicas = 3
}

service "web" {
  project = "shop"
  count   = replicas
  expose {
    domains = ["${domain}", "www.${domain}"]
  }
}

Values are strings, numbers or bools: a list or object is refused by name, and so is redeclaring a name or shadowing a built-in. A variable's value may reference node variables and built-ins, never another spec variable. Variables are not secrets: the node's stanza is served to any signed-in caller over GET /v1/vars, so a secret stays a secret: reference; a variable whose value contains one is fine.

project

A named group of services, and the isolation boundary for network policy, secrets, containerd namespaces and DNS.

FieldTypeNotes
descriptionstringFree text, ≤ 512 chars. Shown in the dashboard.
gitblockOptional GitOps source for this project.
notificationsblockOptional channels and filters.

git

FieldTypeNotes
urlstringRequired. Clone URL: https://, ssh://, scp-form or a local path - http:// and git:// are refused (the deploy credential would travel in cleartext / nothing authenticates the content), and credentials in the URL's userinfo are refused in favour of auth_ref.
branchstringDefaults to the repository's default branch.
pathstringDirectory within the repository holding specs, e.g. .kanea/.
auth_refsecret refDeploy key or token. Project-scoped like every other reference.
webhook_secret_refsecret refHMAC key for the push webhook.
poll_intervaldurationHow often to poll when no webhook arrives.
require_approvalboolSync marks the project; a human promotes.
A repository speaks for its own project and no other

A synced spec that declares a different project is refused. It is the same boundary that scopes secrets, and it is the only thing between "can push to one repo" and "owns every service on the node".

The webhook never deploys from the request either. It marks the project as dirty and the sync loop re-reads the source over Kanea's own credential, so a forged body can at most cause a legitimate sync to happen sooner.

notifications

Up to five channel blocks; telegram, slack (Discord accepts the same shape on its /slack endpoint), ntfy, smtp, webhook; plus filters:

FieldTypeNotes
onlist(string)Event globs, e.g. ["deploy.failed", "scale.*"]. Validated against the known event vocabulary at parse time; a filter that could never match is a spec error, not a silence.
severitystringFloor: info, warning, error. Composes with on as an AND.
telegramblockchat_id, token_ref
slackblockurl_ref: an incoming-webhook URL is a credential in path form, so it is referenced, never inlined
ntfyblockurl, token_ref
smtpblockhost, port, from, to, username, password_ref
webhookblockurl, secret_ref (HMAC signature over the body)

Egress is checked at dial time against every resolved address, and redirects are refused. A hostname is not a destination.

storage

A named volume backend. Declared here or in the server config; services mount it by name with a volume block.

typeFieldsNotes
local-A directory on the node, managed by Kanea.
hostpath, createA directory the operator already owns, or, with create = true, one Kanea creates inside an allowed prefix. See R15: it does nothing until an operator allowlists its parent in /etc/kanea/kanea.hcl's storage { allowed_host_paths = […] } stanza (or --allowed-host-paths, which wins).
nfsserver, export, options
smbserver, share, auth_ref, options
s3bucket, endpoint, auth_ref, modemode = "ro" selects mountpoint-s3, "rw" selects s3fs.

auth_ref is R5-scoped like every other credential reference: every service mounting the storage must be allowed to read it (its own project or shared/), because mounting is reading. And an endpoint must be https:// - the mount helper sends the resolved credential to that address on every request.

S3 volumes are not a filesystem

Every file operation costs an object-store round trip; around 30 ms. Listing or creating a 200-file directory takes tens of seconds, no driver implements truncate (s3fs silently no-ops it), and a FUSE call against a dead backend blocks uninterruptibly for tens of seconds. Use them for bulk, read-mostly data. Never for a hot path or many small files.

service

FieldTypeDefaultNotes
projectstring-The owning project.
descriptionstring-Free text, ≤ 512 chars.
countnumber1Replicas. Must sit inside scaling's min/max if that block is present.
hardeningstringcompatibleThe service's posture (R13). restricted is the strong profile in one word: the task must declare a non-zero user, runs with no capabilities at all, and a capability grant beside it is refused. Init steps are deliberately unchanged - chown as root in a step, run the task restricted.
depends_onlist(string)-Start ordering. See R10.
taskblock-Exactly one in v1.
buildblock-Build from source instead of pulling.
networkblock-Ports and ingress policy.
exposeblock-Public exposure and middleware.
health_checkblock-Zero or more, each labelled.
volumeblock-Zero or more mounts.
scalingblock-Autoscaling policy.
updateblockrollingRollout strategy.
restartblockalwaysRestart policy.

task

FieldTypeNotes
imagestringOptional only if a build block is present (R8). Digests are supported and preferred.
commandlist(string)Overrides the image entrypoint. An argument array, never a shell string (R12).
argslist(string)Keeps the image's entrypoint and replaces its arguments (R12, v1.97). Beside command, the argv is command then args - the Kubernetes pair. args = [] is refused; omit the field instead.
capabilitieslist(string)Grants added to the baseline set; "none" starts from nothing instead (R13).
envmapValues may be secret: references or ${service.…} interpolations.
resourcesblockcpu in MHz (1000 = one core), memory in MiB. An omitted limit is unbounded: all cores / all allocatable memory. pids overrides the per-alloc process cap (256 when omitted; the cap is always on).
registry_auth_refstringsecret: reference to a docker config.json used to pull the image. Project-scoped like every other reference (R5).
pull_policystringWhere the image may come from (R33): if-not-present (default), never, or always. Omit it to take the node's default. See below.
read_only_rootfsboolMounts the container's root filesystem read-only. Volumes and files still mount as declared, so a writable /tmp is a volume away. kanea plan warns when it is off.
deviceblockRequests a host device the operator granted, by name (R17). See below.
socketblockRequests a host unix socket the operator granted, by name (R18). See below.

init

Setup that has to finish before the workload starts: a schema migration, a chown of a freshly-created volume, a config render. A service may declare any number of init blocks; they run in declaration order, one at a time, to completion, and the task is created only once the last has exited zero (R32).

Each step shares the alloc: the same network namespace (so ${service.…} and internal DNS resolve exactly as they do for the task, which is what makes a wait-for-database step possible), the same volumes at the same mount paths, and the same secrets. Everything else it declares for itself, and nothing is inherited from task.

service "api" {
  # The strong posture (R13): the chown happens in the root init step,
  # so the task itself runs as 999 with no capabilities at all.
  hardening = "restricted"

  # Runs as root to fix a directory the task will own as 999.
  init "fix-perms" {
    image        = "busybox:1.36"
    command      = ["chown", "-R", "999:999", "/data"]
    capabilities = ["CAP_CHOWN"]
    timeout      = "1m"
  }

  init "migrate" {
    image   = "registry.example.com/shop/api-migrate:1.4"
    command = ["/bin/migrate", "up"]

    env = {
      DATABASE_URL = "secret:shop/database-url"
      DATABASE_HOST = "${service.postgres.host}"   # a real dependency edge
    }

    timeout = "5m"
  }

  task "app" {
    image = "registry.example.com/shop/api:1.4"
    user { uid = 999 }
  }
}
FieldTypeNotes
imagestringRequired. Unlike a task an init container has no build block, so there is no pipeline to produce one.
commandlist(string)Overrides the image entrypoint. An argument array, never a shell string (R12's rule).
argslist(string)Keeps the step image's entrypoint and replaces its arguments (R12, v1.97), exactly as on a task.
envmapsecret: references and ${service.…} interpolations work exactly as they do on a task.
resourcesblockIts own. An omitted limit is unbounded (R11); nothing is inherited from the task.
userblockIts own numeric identity (R23). Independent of task.user on purpose: running as root to chown a directory the task owns as 999 is the canonical use.
capabilitieslist(string)R13's baseline and permitted set, same rule as a task's.
registry_auth_refstringThis step's pull credential, project-scoped like every other reference (R5).
pull_policystringR33's, minus always: there is one pinned-image field and it belongs to the task.
timeoutstringBounds this step. Omit it for no timeout. A floor, not a deadline: the kill lands within one reconcile pass.

A sequence re-runs from the first step whenever the alloc running it is created - a deploy, a crash restart, a spec change - so init steps must be idempotent. That is the same standing requirement the netns, the secrets tree and the staged host binds carry: an alloc is built, never resumed.

A sequence runs once per service, not once per alloc

It runs on the first alloc; every other alloc of the service waits for it and then starts with no sequence of its own. That is what makes a schema migration expressible: three replicas, one migration. Without it a first deploy of count = 3 creates three allocs in one pass and runs three copies of the migration at once, against the same database.

The gate is the first alloc's record leaving init at the same spec hash, so a deploy re-runs the sequence once and the others wait again. A waiting alloc says what it is waiting for on kanea ps. If the first alloc fails its sequence the service does not start, and the wait names it: the alternative is failing every replica for one migration's mistake.

The cost is per-alloc volumes. Local storage gives each alloc its own directory, so a step that prepares this alloc's volume - the classic chown - now prepares the first alloc's and no other. kanea plan warns when a service with count > 1 declares both init steps and a local volume. Steps that work on shared state are unaffected.

env_group

An environment declared once and taken by the services that want it. variables substitutes ${name} into text you write, so sharing an environment with them still means writing every key in every service; a group is injected.

env_group "common" {
  LOG_LEVEL = "info"
  REGION    = "eu-central-1"
}

env_group "db" {
  DATABASE_HOST = "${service.postgres.host}"
  DATABASE_URL  = "secret:shop/database-url"
}

service "api" {
  project  = "shop"
  env_from = ["common", "db"]

  task "app" {
    image = "registry.example.com/shop/api:1.4"
    env   = { LOG_LEVEL = "debug" }   # the service's own env wins
  }
}

Groups apply in the order env_from lists them, then the task's own env on top. It is opt-in per service rather than project-wide deliberately: environment is baked into a container, so a shared value changing rolls every service that takes it, and a blast radius that size should be something a spec states rather than something a service inherits by living in a project.

A group is evaluated once per service that takes it

Not once per spec, and the difference is visible. The ${service.…} namespace is scoped to a project, so the db group above resolves to postgres.shop.kanea for a service in shop and to a different address for one in another project. It also means the dependency edge lands on the taking service, which is what makes it start behind postgres. And a secret: in a group is scoped against every consumer: a group carrying shop's credential is refused for a service in another project.

file

Content Kanea places in the container, instead of baking it into an image or mounting a host volume and putting the file there yourself.

service "web" {
  file "nginx" {
    path   = "/etc/nginx/conf.d/app.conf"
    source = "./nginx.conf"        # read at parse, embedded in the record
  }

  file "pgpass" {
    path    = "/etc/app/pgpass"
    mode    = "0400"
    content = "db:5432:app:${secret.shop[\"database-password\"]}"
  }
}

content is an ordinary HCL string, so a whole config file goes in a heredoc; this is the shape most files actually take.

file "app-config" {
  path    = "/etc/app/config.yaml"
  content = <<-EOT
    server:
      addr: ":8080"
      upstream: "${service.api.host}:${service.api.port.http}"
    database:
      dsn: "postgres://app:${secret.shop[\"database-password\"]}@db:5432/app"
    log:
      # a literal dollar-brace the app expands itself
      format: "$${level} $${msg}"
  EOT
}

Three things there are worth naming. ${service.…} and ${secret.…} interpolate in the same file, and it is the second that decides the rest: because one line names a secret, the whole file is written 0400 on a tmpfs of its own and the record keeps the reference, never the value. Content of its own that looks like interpolation is escaped $${…}, which is HCL's rule rather than Kanea's. And file blocks are the one place a <<-EOT body reads better than a source path: a short config that belongs to the spec does not need a second file beside it.

FieldTypeNotes
pathstringRequired. Absolute, clean, ..-free. Refused under /dev, /proc, /sys, at /etc/resolv.conf, and where a volume or socket already mounts.
contentstringInline. ${service.…} and ${secret.…} interpolate; a literal dollar-brace is $${…}.
sourcestringA path beside the spec, read at parse and embedded. Relative only, no symlink in any component, mutually exclusive with content. Refused where the parser has no directory: the dashboard's spec editor and MCP, because that parse happens inside kanead as root and reading a path there would be an arbitrary file read.
modestringOctal. Default 0644, or 0400 for a file interpolating a secret. An execute bit is refused: a file block delivers configuration, not a program.
--Size. 64 KiB per file and 128 KiB per service, post-substitution (PRD §21). A Desired record is CDC-replicated and the change log holds a second copy of every value it ships, so bulk bytes here are paid for on every apply, forever.

Init containers must be idempotent. That is a requirement, not advice: a half-run sequence is abandoned rather than resumed, and the sequence re-runs from the first step whenever the alloc running it is created - a deploy, a crash restart, a spec change - for the same reason the network namespace and the secrets are rebuilt.

There is deliberately no device, socket, health_check, expose, scaling, count, depends_on, network or build field on an init block. An init container is a step, not a service, and the absence is the refusal.

A step's output goes to its own log: kanea logs shop/api -c migrate. Attempts append to one transcript per alloc, and each attempt opens with a separator line naming the step and when it started, so a tail says where the previous attempt ends. A step that runs and exits non-zero, or outlives its timeout, fails the alloc and spends the restart budget (R29), so a broken migration stops after attempts rather than hammering a database forever; a step that could not be pulled or created is retried every pass without spending it, because nothing ran.

pull_policy

Where an image may come from (R33), on a task or on an init block. Three values:

ValueMeaning
if-not-presentUse the node's content store, pull only what is absent. The default, and what every spec did before this field existed.
neverUse the content store or fail the alloc, naming the policy. For air-gapped nodes and ones whose images are preloaded: a pull attempt there is a slow failure that blames a registry the node was never going to contact.
alwaysRe-resolve the tag and, when the digest has moved, pin it and roll every replica through the update policy. Task only.

always does not mean re-pull at every container creation. That would let two replicas of one spec run different bytes, which is exactly what “a deploy is a spec-hash mismatch” exists to prevent. It lowers to update { auto = true }, so the tag is polled, a moved digest is pinned, and every replica rolls together through max_parallel, min_healthy and the health check. That also means it inherits auto-update's refusals: it cannot sit beside a build block, needs a tag rather than a digest, and cannot be combined with update { auto = false }. The unattended cadence is update { interval } (default 6h, floor 5m), but re-applying the spec forces an immediate re-check: push new bytes under the same tag, kanea apply, and the digest is re-resolved within about a minute instead of waiting out the interval.

Omitting pull_policy takes the node's default, which an operator sets in /etc/kanea/kanea.hcl:

images {
  pull_policy = "never"       # or --image-pull-policy, which wins
}
Files are mounted read-only with nosuid,noexec,nodev, and after volumes - so a file at a path inside a volume wins, which is what "declare a file at a path" has to mean.

A secret in a file never enters Kanea's state

${secret.<scope>.<name>} resolves when the spec is parsed to an opaque placeholder. The stored record keeps the reference, exactly as it does for an environment variable, and the value is substituted on the node when the container is created - so a rotation lands at the next replacement, and the credential is not in the state database, in a backup archive, or in GET /v1/services. A file carrying one is written 0400 onto a tmpfs of its own, owned by the workload's user.

Use the bracket form for any name that is not a bare identifier, which is most of them: ${secret.shop["database-password"]}.

Where source can be read

From the directory beside the spec for kanea run, and out of the commit for a synced repository. The dashboard's spec editor and the MCP spec tools parse text rather than files - inside the daemon - so they refuse source by name and want content. That refusal is the feature: a parser that opened files there would be reading the node's filesystem as root on behalf of whoever is signed in.

Content is capped at 64 KiB per file and 128 KiB per service: a service record is replicated in full on every deploy.

device and socket

A GPU for transcoding, a USB dongle, or the container runtime's socket for a watchtower-style updater. Both blocks name a grant and never a path: the operator defines grants on the node, and the spec asks for one by name. There is no field to write a device path into, which is the property the whole model rests on: a spec cannot request /dev/mem because there is nowhere to say it.

service "jellyfin" {
  task "app" {
    image = "jellyfin:10.9"

    device "dri" {
      grant = "gpu"          # defined by the operator, not here
    }
  }
}

service "watchtower" {
  task "app" {
    image = "watchtower:1.7"

    socket "runtime" {
      grant      = "containerd"
      mount_path = "/var/run/docker.sock"
    }
  }
}
FieldTypeNotes
grantstringRequired, on both blocks. Names a device/socket grant in the node's server config, /etc/kanea/kanea.hcl.
mount_pathstringsocket only, required. Where the socket appears in the container. A device appears at its host path.
read_onlyboolsocket only. Restricts the filesystem entry, not the protocol spoken over it.

The operator side lives in the node's server config, /etc/kanea/kanea.hcl (full reference: Node configuration → device and socket), probed automatically at daemon start: no flag, no unit editing, and a kanea init re-run never touches it. The default is that the file does not exist: a spec asking for a grant on a node with none fails its alloc rather than starting without it. Grants name the projects that may claim them:

# /etc/kanea/kanea.hcl; the node's, never the repository's
bind {                                  # the API/dashboard listener (v1.61)
  api_addr = "192.168.1.10:8600"
  api_tls  = "self-signed"              # acme | self-signed | provided | plaintext
}

variables {                             # node-wide spec-variable defaults (v1.63), never secrets
  domain = "home.lan"
}

device "gpu" {
  nodes = ["/dev/dri/renderD128"]
  allow = ["media"]
}

socket "containerd" {
  path  = "/run/kanea/containerd.sock"
  allow = ["ops"]
}

Setting it up, end to end:

  1. Create the file (init already made the directory): sudo install -o root -g root -m 0644 /dev/null /etc/kanea/kanea.hcl, then write the grant blocks above, and, if you use host volumes, the storage allowlist stanza in the same file. It must stay root-owned and writable only by its owner; kanead refuses a group- or world-writable policy file.
  2. sudo systemctl restart kanead: the file is read once, at startup, never polled. Running workloads and ingress both survive the restart; see applying a change.
  3. Verify: the startup log carries server config loaded with the path, and a warn line naming the passthrough consequence when grants are present. kanea ps shows the alloc converging.

A socket grant is root on this node for whoever holds it. A container with the runtime socket can start other containers without the hardening defaults. There is no containment story and none is claimed: the control is that granting it happens on the machine, in a file no spec author can write, and names one project.

network

network {
  port "http" { container = 8096 }

  publish "http" {
    host = 8096                                  # http://<node>:8096
    mode = "http"                                # "http" (default) | "tcp"
    ip_restriction { allow = ["192.168.0.0/16"] }
  }

  policy { allow_from = ["analytics/collector"] }
}
  • port "<name>": container is the port inside the container. The name is what expose and health_check refer to, and what ${service.x.port.<name>} resolves.
  • publish "<port name>": a node port the edge binds for this service, with or without a domain. The label names the port above; there is deliberately no field for a container port number, so a published port can never forward somewhere the service did not declare. See R21.
  • mode = "tcp" relays bytes for Postgres, a game server, or anything that is not HTTP. It keeps ip_restriction and nothing else: a rate_limit or headers block on one is a plan error rather than a control that is silently dropped. On a tcp listener the upstream sees the edge's address, not the client's, so ip_restriction is the whole mitigation and it is enforced at accept time.
  • mode = "udp" relays datagrams (game servers, DNS, syslog): one session per client source address, pinned to a backend by rendezvous hash so a conversation survives an edge restart. Because one spoofed datagram can open a session, two bounds assume forgery: at most 8 live sessions per source IP, and at most 4 KiB back to the client until it sends a second datagram - the proof its address is real. Both refusals are counted in kanea_edge_udp_refused_total.
  • Which node ports a spec may claim belongs to the node, not the spec (--publish-ports, unprivileged by default). A repository anyone can push to must not be able to take :22. See R22.
  • policy.allow_from: fully-qualified "<project>/<service>" peers permitted to reach this service. See R14.

expose

North-south exposure: the edge proxy, TLS, and the middleware chain. The block may repeat: each expose is one complete route with its own domains, port, TLS and middleware, so one service can serve its UI and its API on different names and ports. Only the first block may omit domains, and blocks that declare auth must declare the same auth.

FieldTypeNotes
domainslist(string)Omitted → <service>.<project>.<base_domain>. No two services may claim the same domain, counting generated ones.
portstringWhich declared network { port } the route proxies to. Omitted → the port named http, or the sole declared port. Explicit beats every convention; naming an undeclared or udp port is a plan error.
tlsblockmode: acme, self-signed, provided or plaintext; name selects one of the node's provided grants. Omit the block and the node's --tls-default decides. A mode names a source, never a path.
ip_restrictionblockallow, deny: lists of CIDRs. Empty allow means the world; deny wins.
rate_limitblockrequests, window, per, burst. per is ip, service, or header:<name>.
headersblockrequest_set, request_remove, response_set, response_remove.

The chain evaluates in a fixed order regardless of declaration order:

Host match → IP allow/deny → rate limit → header transforms → upstream
You cannot touch X-Forwarded-*

The headers block rejects any attempt to set or remove the X-Forwarded-* set, or the hop-by-hop headers. Those carry the client identity that IP restriction, rate limiting and the audit log are all keyed on: a spec able to rewrite X-Forwarded-For would be forging the thing every other control trusts.

health_check

health_check "http" {
  type     = "http"     # http | tcp | exec
  path     = "/healthz" # http only
  port     = "http"     # port name, required for http and tcp
  interval = "10s"
  timeout  = "2s"
  failures = 3
}

An exec check runs inside the task's container and takes command as an argument array, never a shell string: the same rule as task.command, for the same reason.

No check means no alloc is ever "healthy"

Alloc health is only ever written by a probe. A service without a health_check block has every alloc unhealthy forever, and that is correct: Kanea asks whether a check is configured before it asks whether one passed. It matters because min_healthy in the update block is meaningless without one.

volume

volume "data" {
  storage    = "local-ssd"    # a declared storage block
  mount_path = "/var/lib/data"
  read_only  = false
  size       = "10GiB"        # a budget, not a quota: see below
}

Mount failures fail the alloc loudly. A volume that silently is not there is worse than one that stops the deploy.

size declares a budget (R31). Kanea measures what the volume actually uses on a slow background schedule, shows it in kanea volume list, and emits volume.over_budget when it is crossed, and volume.under_budget when it comes back. It is deliberately not a quota and nothing enforces it: no filesystem quota mechanism exists on the node, and nfs, smb and s3 could not carry one if it did. Declaring it on an s3 volume is a plan error, because that driver is never walked: one request per directory, on a schedule, is not a measurement worth taking. Changing a size never redeploys the service: a budget is not baked into a container.

scaling

scaling {
  min = 2
  max = 10
  metric "cpu"            { target = 70 }   # percent of resources.cpu
  metric "rps"            { target = 500 }  # requests per second, per alloc
  metric "p95_latency_ms" { target = 800 }
  cooldown = "2m"
}
  • cpu comes from containerd cgroup metrics and is a percentage of the alloc's declared limit: the number the workload can actually use.
  • rps and p95_latency_ms come from the edge for exposed services, and from the datapath's own map counters for east-west traffic.
  • count must fall inside [min, max], otherwise the autoscaler would immediately contradict the spec.

update and restart

update {
  strategy     = "rolling"   # rolling | replace
  max_parallel = 1
  min_healthy  = "30s"

  auto     = true            # follow task.image's tag (R19); off by default
  interval = "6h"            # how often to re-resolve it; minimum 5m
  deadline = "10m"           # how long a new digest has to prove itself
}

restart {
  attempts = 5
  backoff  = "10s,30s,1m,5m"
}

max_parallel bounds allocs that are down, not replacements in flight: anything already unavailable spends the budget first, so a failing deploy halts rather than marching through every replica. min_healthy applies only to allocs this deploy has already replaced.

restart bounds crash restarts: each one waits the next backoff entry (the last repeats), and an alloc that exhausts attempts is failed and left alone until a deploy or kanea restart - either resets the count, because the budget belongs to the spec hash that spent it. The budget counts only crashes the daemon watched (v0.31.0): an alloc found already stopped at startup - a node reboot, a power loss - is recovered immediately without spending an attempt, however many the counter shows, and without waiting out a backoff armed before the outage. An alloc already marked failed stays failed across a reboot: that verdict came from crashes the platform saw, and the fix is the same deploy or restart as ever.

auto is the watchtower feature, as policy rather than as a container holding your runtime socket. Kanea re-resolves the tag the service already declares, and when the digest behind it moves it pins the new one: the tag stays in the spec, the digest is what runs. That is an ordinary deploy nobody typed: it rolls with max_parallel, waits min_healthy and is gated by the service's health check, because it goes through the same machinery every other deploy does. If the new digest does not converge within deadline, the previous one is re-pinned and the service goes back to what worked.

It is refused on a service whose image is already a digest (nothing to follow) and on one with a build block (the pipeline owns that image). Private registries need task.registry_auth_ref. Both outcomes emit events: image.updated and image.update_failed.

build

build {
  context           = "./web"
  dockerfile        = "Containerfile"   # optional; auto-detected when omitted
  target            = "registry.example.com/shop/web"   # optional; omitted, the node's
                                                        # internal registry fills in
                                                        # <registry>/shop/web
  tag               = "${GIT_SHA_SHORT}"
  cache_repo        = "registry.example.com/shop/web-cache"
  registry_auth_ref = "secret:shop/registry"
}
  • Builds run on rootless BuildKit, unprivileged end to end.
  • An omitted target means the node's internal registry. Every node runs a loopback-only build registry (--registry, default 127.0.0.1:5100; off disables it, and --buildkit off implies that), and the target is composed on the node as <registry>/<project>/<service>, so build { context = "." } is a complete build configuration: no external registry, no push credential. Pushes to it are credential-gated per boot; pulls are loopback plain HTTP, which containerd already speaks. An explicit target behaves exactly as before, and cache_repo may not point at the internal registry.
  • Containerfile and Dockerfile both work, and Containerfile wins when both exist.
  • context and dockerfile are paths inside the checkout: absolute and .. forms are refused, and the context must be a real directory, never a symlink.
  • target and cache_repo must parse as image references, and tag may not contain a comma, an = or whitespace - buildctl reads those flags as comma-separated option lists, so a comma is option injection, not a name.
  • The push credential is materialised as a config.json for the build and never enters the build context.
  • A service with a build block and no task.image is legitimate: the reconciler skips it until the first build pins a digest.
  • Builds are serialised, and refused when full rather than queued forever. Isolation is collective, so a second concurrent build would share the first's budget. queued is a real state and shutdown cancels what is still waiting.

Interpolation

Three kinds of ${…} appear in a spec, and they resolve at different times.

FormResolvedNotes
${VAR}Parse timeFrom the spec's variables block, the node's variables stanza in /etc/kanea/kanea.hcl, and built-ins such as KANEA_PROJECT. Node under spec under caller (R30).
${GIT_SHA_SHORT}Checkout timeSurvives parsing as a literal reference: its value does not exist until a commit is checked out.
${service.<n>.host}Alloc startBecomes <n>.<project>.kanea. Validated at plan time; resolved as a DNS name, never an IP, so load-balancer reprogramming cannot break it.
${service.<n>.port.<p>}Alloc startThe named port's frontend port.
secret:<path>Alloc startNot interpolation: a reference the reconciler resolves. Project-scoped (R5).

Removing a service from a spec

Deleting a service block and re-applying does not stop the service. An apply replaces the desired state of the services it names and leaves everything else alone, so that kanea run web.hcl can never delete what db.hcl declares. Removing a service is therefore a deliberate act, and there are two ways to perform it.

One serviceEverything a spec dropped
Commandkanea remove shop/webkanea run --remove-orphans app.hcl
Scopethe service you namethe spec's project blocks
Dry run-kanea plan --remove-orphans app.hcl

--remove-orphans makes the spec authoritative for the projects it declares a project block for: anything stored in one of them that the file no longer declares is deleted, in the same atomic batch as the applies. Projects the spec does not mention are never touched, so a spec for shop cannot reach into data.

$ kanea plan --remove-orphans app.hcl
- destroy shop/legacy-worker (count 1, image legacy:v3)
    volumes           - scratch  local /scratch rw
                        (the mount goes; the volume's data is NOT deleted)

~ update shop/web
    image             web:v1 -> web:v2  (rolls allocs)

Plan: 2 change(s) - 1 update, 1 destroy; 1 replace running allocs.
Run `kanea run` to apply.

$ kanea run --remove-orphans app.hcl
[the same block]

Apply? [Y/n]
applied shop/web
removed shop/legacy-worker

1 service(s) removed. Volume data was not deleted.
Refused with a selector, and with --image

kanea run app.hcl shop/web --remove-orphans is refused by name. A selector sends part of the spec, so the apply cannot claim to be the whole of any project - and if it did, it would read as "delete everything in shop except web". --image is refused for the same reason from the other end: it declares one service and no project at all.

Note that flags must come before file arguments: kanea run --remove-orphans app.hcl, not the other way round.

What a prune destroys, and what survives

GoesStays
The containers, and the alloc records behind kanea psVolume data, local and network
The service's virtual IP (released, and another service may take it)Certificates, including issued ACME ones
Its routes, and its network volume mountsSecrets under the project
The restart generation and any pinned image digestIts internal datapath id, which is keyed by name

Volume data is never deleted, which cuts both ways and is why the command says so every time. A service pruned by mistake comes back with its data by re-applying it, because the directory is derived from project/service/index and nothing removes it. And a prune frees no disk: reclaiming it is a manual rm under the volume directory, deliberately, since Kanea has never deleted a volume it did not create and R15 says as much for host paths.

Each removal emits service.removed and is named in the audit log, so a prune is not a quiet operation even when nobody is watching the terminal.

A GitOps sync is unaffected: a synced repository still applies additively, and removing a service from a synced spec leaves it running until someone prunes or stops it.

Validation rules

The rules a spec is checked against, all enforced at parse or plan time with file/line/column diagnostics. They are numbered because the error messages cite them, and the numbering is the PRD's; the cards below cover the ones most people meet, not every rule the parser knows.

R1

Names are DNS-1123 labels. Project and service names: lowercase alphanumeric and -, alphanumeric at both ends, ≤ 63 characters. Parse errors abort the run.

R2

Variables. Flat ${VAR} interpolation in any attribute, from the spec's variables block, the node's variables stanza and built-ins: node under spec under caller. Reserved names (the built-ins, service) cannot be declared, values are primitives, and a definition never references a sibling.

R3

Secrets are referenced, never inlined. Resolved at alloc start. Primary injection is a tmpfs file at /run/kanea/secrets/<alloc>/<name>; env vars are supported and documented as weaker.

R4

kanea plan is a real dry run and shows the create/change/destroy diff before anything is applied.

R5

Secret references are project-scoped. A service may name secret:<own-project>/… or secret:shared/… and nothing else. Cross-project references are rejected. Git, registry, storage and notification credentials follow the same scoping.

R6

spec_version = 1 is declared by every file; future revisions are gated on it.

R7

Health check types are http, tcp and exec. exec runs inside the container and takes an argument array, never a shell string.

R8

The minimal service is image-only. At least one of task.image or build must be present. When both are, the pipeline-built digest wins and task.image is the pre-first-build value.

R9

Service references are same-project in v1, validated at plan against the whole applied set: the referenced service and port must exist, and file order is irrelevant. Resolved as DNS names, never IPs. Cycles are rejected, with the cycle shown in the diagnostic.

R10

Dependencies gate starts, not stops. depends_on and every implicit reference edge mean a dependent will not start until its dependencies are healthy. If a dependency degrades afterwards, dependents keep running and events are emitted: no cascading stops.

R11

A declared limit is enforced; an omitted one is the node's capacity. Omit resources and the alloc gets all cores and all allocatable memory, bounded by the workload parent cgroup, not by a per-alloc number nobody typed. A default pids.max applies regardless; resources.pids declares a service's own cap (256 when omitted), and changes it like any other limit. A declared memory.max breach OOM-kills the alloc, emits an event, and the restart policy applies. Scaling on cpu or memory (percent-of-limit) requires the corresponding limit and is refused at plan without it. Declared values are the admission units counted against the node's workload budget. Functions keep small defaults (cpu = 100, memory = 64): the wasm sandbox's caps are promises.

R12

task.command is an argument array. The first element must be non-empty; later ones may be empty, because some programs use that meaningfully; redis-server --save "" is how you disable snapshots. task.args keeps the image's entrypoint and replaces its arguments (v1.97): alone, argv is the image's ENTRYPOINT then args, resolved on the node from the image itself; beside command, argv is command then args - the Kubernetes pair. Its elements may all be empty (they are arguments, not a program), but args = [] is refused: the record cannot carry "declared empty" apart from "absent", and command is the spelling for "run the bare entrypoint". Both apply to init blocks too; a function takes neither.

R13

task.capabilities adds to a baseline, and "none" takes the baseline away. Every runc alloc starts from the baseline set; CHOWN, DAC_OVERRIDE, FOWNER, FSETID, KILL, SETGID, SETUID, so PUID-style images start without a capability line. Binding :80 needs no capability at all: every alloc netns sets ip_unprivileged_port_start=0, which is why NET_BIND_SERVICE is no longer baseline. The effective grant is the union of baseline and declared list; ["none"] is the full drop-ALL posture, ["none", "CAP_NET_RAW"] exactly one grant. Declarable beyond the baseline: SETPCAP, SETFCAP, NET_BIND_SERVICE, NET_RAW (never baseline; the datapath's identity is the IP), SYS_CHROOT, MKNOD, AUDIT_WRITE. One honest caveat on SETFCAP: file capabilities are xattrs, so a binary given them inside an rw host volume keeps them on the host for whoever executes it later; no-new-privileges contains the grant inside the container, not the file. Privilege-equivalent ones; SYS_ADMIN, SYS_MODULE, SYS_PTRACE, BPF, PERFMON and friends; are rejected at parse time: granting them would be the privileged escape hatch v1 deliberately does not have. Effective capabilities go into the bounding, effective and permitted sets, never inheritable or ambient. hardening = "restricted" on the service names the strong posture: a non-zero user required, drop-ALL enforced, any declared grant refused - and init steps unchanged, because the canonical shape is a root init step that chowns and exits before a restricted task.

R14

allow_from only ever adds reachability. Each entry is a fully-qualified "<project>/<service>"; the datapath's ingress rules only ever union, so an entry can never weaken the project default-deny. There is no wildcard: "analytics/*" is a parse error, because naming the peer is the point. Same-project entries are accepted and redundant.

R15

host volumes are operator-gated. The path is validated as absolute, clean and ..-free at parse time, but whether it may be mounted is not the spec's decision. kanead refuses any path outside storage.allowed_host_paths in the server config /etc/kanea/kanea.hcl (or --allowed-host-paths, which wins when set), whose default is empty. The check is applied after symlink resolution, and the directory must already exist - unless the storage block sets create = true, which makes Kanea create it. The default has not moved: without that flag a missing path is still refused, because creating one on demand turns a typo into a volume that is silently empty. Creation is allowlist-gated first - the nearest existing parent is resolved and checked before anything is written - so a create outside a permitted prefix leaves no directory behind. Creating a directory is not owning it: Kanea still never chowns a host volume.

R16

expose fails closed. A service may only be exposed if it declares a port, and the upstream port must be unambiguous; declared with port = "<name>", or named http, or the only one declared. The block may repeat, each one a complete route validated independently; only the first may omit domains, and blocks that declare auth must agree. Every domain is validated as a hostname, and no two services may claim the same one, counting generated FQDNs. Middleware is checked here too: CIDRs must parse, rate_limit needs a positive requests and a valid window, and headers may not touch the hop-by-hop or X-Forwarded-* sets.

R17

task.device names a grant, not a device. There is no field for a device path, so a spec cannot ask for one. Parse time checks only that the grant name is a DNS-1123 label. The node refuses a grant it does not have, a grant whose allow list does not name the requesting project, and a path that is no longer a character or block device: checked after symlink resolution, at every alloc start. The device appears at its host path, and the grant carries the cgroup permissions (rw by default, never m unless written). A failed grant fails the alloc; it never starts without what it asked for.

R18

task.socket is R17 for unix sockets, and is privilege delegation. mount_path is validated as absolute, clean and ..-free, may not sit under /dev, /proc or /sys, and may not collide with another socket or a volume. The bind carries nosuid, noexec and nodev. None of that makes it safe and none of it is meant to: a container holding the container runtime's socket can create containers without the hardening defaults, so the grant is equivalent to root on the node. The server config is the only control over it, which is why it is project-scoped and empty by default.

R32

init blocks run to completion before the task. In declaration order, one at a time; the task is created only once the last has exited zero. Each shares the alloc's network namespace, volumes and secrets, and declares its own image, command, args, env, resources, user, capabilities and timeout - nothing is inherited from task. There is no device, socket, health_check, expose, scaling, count, depends_on, network or build field: an init container is a step, not a service. Names are DNS-1123 labels, unique within a service, and compose into a container id and a log file. A sequence runs once per service, not once per alloc (v1.92): on alloc 0, with every other alloc waiting on that alloc's record leaving init at the same spec hash, which is what makes "three replicas, one migration" expressible. A local volume is one directory per alloc, so a step preparing this alloc's own volume prepares the first's alone; kanea plan warns on that shape. They must be idempotent: a half-run sequence is abandoned rather than resumed, and it re-runs whenever the alloc running it is created. A step that ran and failed, or outlived its timeout, spends the restart budget (R29); one that could not be pulled or created is retried every pass without spending it, because nothing ran.

R33

pull_policy says where an image may come from. if-not-present (the default), never (content store only - the air-gapped case) or always. always is not a per-create re-pull, which would let two replicas of one spec run different bytes: it lowers to update { auto = true }, so a moved digest is pinned and every replica rolls together, and it inherits that rule's refusals. It is refused on an init block, because there is one pinned-image field and it belongs to the task. The node supplies the default through /etc/kanea/kanea.hcl's images stanza; an omitted policy means "the node decides".

R34

An env_group is declared once and taken by a service. A top-level block names environment variables; a service opts in with env_from = ["common", "db"], and the groups apply in that order with the task's own env on top. Opt-in per service rather than project-wide, because environment is baked into a container and a shared value changing rolls every service that takes it. A group is evaluated once per consuming service: ${service.…} is project-scoped, so one group taken from two projects resolves differently, and the dependency edge lands on the service that took it. A secret: inside a group is scoped against every consumer, so a group carrying one project's credential is refused for a service in another.

R35

A file block puts content in a container, and a secret in it stays a reference. path is absolute, clean and ..-free, refused where a volume or socket already mounts and under the system paths a volume destination is; the bind is read-only with nosuid,noexec,nodev, and an execute bit in mode is refused. Files mount after volumes, so one inside a volume's path wins. Exactly one of content and source; source is read where the spec is parsed and is refused where there is no directory (the dashboard editor, MCP), because a parser reading files there would be reading the node's filesystem as root. ${secret.<scope>.<name>} resolves to an opaque placeholder at parse and is substituted on the node at container create, so the value never enters the state database, a backup, or the API; such a file is written 0400 on a tmpfs of its own. Content is SpecHash material - editing a config file rolls the service, which is the feature - and is capped at 64 KiB per file and 128 KiB per service.

R19

update.auto follows the tag the service declares. Off by default. Kanea re-resolves task.image's tag every interval (default 6h, minimum 5m) and pins the digest behind it when it moves; the declared tag is never overwritten, because it is what the next poll re-reads. The pinned digest is server-owned and survives kanea apply: except when you edit image or turn auto off, which both hand authority back to the spec. An apply does re-arm the poll, though: the tag is re-resolved within about a minute of one, so a re-pushed tag lands with a push and an apply rather than an interval later. A failed update reverts to the digest that was running if the new one has not converged within deadline: converged means healthy where a check block exists, and running without crash-looping where it does not. Refused on a digest-pinned image and on a service with a build block.

Full example

Everything above, in one file. This parses and validates against Kanea's own parser.

shop.hcl
# shop.hcl: everything for one project
spec_version = 1

project "shop" {
  description = "E-commerce storefront stack"

  git {
    url      = "https://github.com/example/shop-deploy.git"
    branch   = "main"
    path     = ".kanea/"
    auth_ref = "secret:shop/github-deploy-key"
  }

  notifications {
    slack { url_ref = "secret:shop/slack-webhook" }
    on       = ["deploy.failed", "service.unhealthy", "scale.*"]
    severity = "warning"
  }
}

storage "local-ssd" {
  type = "local"
}

service "postgres" {
  project = "shop"

  task "db" {
    image = "postgres:16-alpine"
    env = {
      POSTGRES_PASSWORD = "secret:shop/postgres-password"
    }
    resources {
      cpu    = 1000
      memory = 1024
    }
  }

  network {
    port "pg" { container = 5432 }
  }

  volume "data" {
    storage    = "local-ssd"
    mount_path = "/var/lib/postgresql/data"
  }

  health_check "tcp" {
    type     = "tcp"
    port     = "pg"
    interval = "10s"
  }
}

service "web" {
  project     = "shop"
  description = "Storefront frontend"
  count       = 3
  depends_on  = ["postgres"]

  build {
    # No target: the node's internal registry builds, pushes and deploys
    # this service without any external registry.
    context = "./web"
  }

  task "app" {
    image = "registry.example.com/shop/web:latest"

    env = {
      NODE_ENV      = "production"
      DATABASE_URL  = "secret:shop/database-url"
      DATABASE_HOST = "${service.postgres.host}"
      DATABASE_PORT = "${service.postgres.port.pg}"
    }

    resources {
      cpu    = 500
      memory = 512
    }
  }

  network {
    port "http" { container = 3000 }
  }

  expose {
    domains = ["shop.example.com", "www.shop.example.com"]
    tls { mode = "acme" }

    ip_restriction {
      deny = ["198.51.100.7/32"]
    }

    rate_limit {
      requests = 100
      window   = "1m"
      per      = "ip"
      burst    = 20
    }

    headers {
      response_set    = { Strict-Transport-Security = "max-age=63072000; includeSubDomains" }
      response_remove = ["Server", "X-Powered-By"]
    }
  }

  health_check "http" {
    type     = "http"
    path     = "/healthz"
    port     = "http"
    interval = "10s"
    timeout  = "2s"
    failures = 3
  }

  scaling {
    min = 2
    max = 10
    metric "cpu" { target = 70 }
    metric "rps" { target = 500 }
    cooldown = "2m"
  }

  update {
    strategy     = "rolling"
    max_parallel = 1
    min_healthy  = "30s"
  }

  restart {
    attempts = 5
    backoff  = "10s,30s,1m,5m"
  }
}

Sample stacks

Five complete stacks, from a static site to a Kafka cluster. Every one parses and validates against Kanea's own parser: copy the file, change the names and domains, and start with kanea plan, which will tell you about anything the copy broke before anything runs.

A static site

The smallest production-shaped thing: two replicas behind Let's Encrypt, with a health check. The check is not decoration: it is what a rolling deploy waits on (min_healthy has nothing to measure without one), what anything declaring depends_on this service waits for, and the difference between a replica that stops answering showing up in kanea status and staying a mystery.

site.hcl
spec_version = 1

project "web" {}

service "site" {
  project = "web"
  count   = 2

  task "nginx" {
    image = "nginx:1.27-alpine"

    resources {
      cpu    = 200
      memory = 128
    }
  }

  network {
    port "http" { container = 80 }
  }

  expose {
    domains = ["example.com", "www.example.com"]
    tls { mode = "acme" }
  }

  health_check "http" {
    type     = "http"
    path     = "/"
    port     = "http"
    interval = "10s"
    timeout  = "2s"
    failures = 3
  }
}

This works the moment example.com resolves to the node and ports 80/443 reach it: the ACME HTTP-01 flow needs nothing else. From here the natural next step is replacing the stock image with your own: add a build block and the pipeline builds and pins a digest on every push.

kanea plan site.hcl
kanea run site.hcl --wait=60s
kanea status web/site

Two apps and a database

A frontend and an API sharing PostgreSQL and Redis. This is the shape most multi-service deployments take, and it exercises the machinery that matters: depends_on gates the apps until their backends are healthy (which is why postgres and redis have checks) and the ${service.…} references resolve to internal DNS names at alloc start, so nothing here hard-codes an address. One secret, referenced from two services, never written in the file.

paste.hcl
spec_version = 1

project "paste" {
  description = "Pastebin: frontend, API, database, cache"
}

storage "db-data" {
  type = "local"
}

service "postgres" {
  project = "paste"

  task "db" {
    image = "postgres:16-alpine"
    env = {
      POSTGRES_DB       = "paste"
      POSTGRES_USER     = "paste"
      POSTGRES_PASSWORD = "secret:paste/db-password"
    }
    resources {
      cpu    = 1000
      memory = 1024
    }
  }

  network {
    port "pg" { container = 5432 }
  }

  volume "data" {
    storage    = "db-data"
    mount_path = "/var/lib/postgresql/data"
  }

  health_check "up" {
    type     = "tcp"
    port     = "pg"
    interval = "10s"
  }
}

service "redis" {
  project = "paste"

  task "cache" {
    image   = "redis:7-alpine"
    command = ["redis-server", "--save", ""]
    resources {
      cpu    = 250
      memory = 256
    }
  }

  network {
    port "redis" { container = 6379 }
  }

  health_check "up" {
    type     = "tcp"
    port     = "redis"
    interval = "10s"
  }
}

service "api" {
  project    = "paste"
  count      = 2
  depends_on = ["postgres", "redis"]

  task "app" {
    image = "registry.example.com/paste/api:1.4.2"
    env = {
      DB_HOST     = "${service.postgres.host}"
      DB_PORT     = "${service.postgres.port.pg}"
      DB_USER     = "paste"
      DB_PASSWORD = "secret:paste/db-password"
      REDIS_ADDR  = "${service.redis.host}:${service.redis.port.redis}"
    }
    resources {
      cpu    = 500
      memory = 512
    }
  }

  network {
    port "http" { container = 8080 }
  }

  expose {
    domains = ["api.paste.example.com"]
    tls { mode = "acme" }

    rate_limit {
      requests = 300
      window   = "1m"
      per      = "ip"
      burst    = 50
    }
  }

  health_check "http" {
    type     = "http"
    path     = "/healthz"
    port     = "http"
    interval = "10s"
    timeout  = "2s"
    failures = 3
  }

  update {
    strategy     = "rolling"
    max_parallel = 1
    min_healthy  = "30s"
  }
}

service "web" {
  project    = "paste"
  count      = 2
  depends_on = ["api"]

  task "app" {
    image = "registry.example.com/paste/web:1.4.2"
    env = {
      API_URL = "http://${service.api.host}:${service.api.port.http}"
    }
    resources {
      cpu    = 250
      memory = 256
    }
  }

  network {
    port "http" { container = 3000 }
  }

  expose {
    domains = ["paste.example.com"]
    tls { mode = "acme" }
  }

  health_check "http" {
    type     = "http"
    path     = "/"
    port     = "http"
    interval = "10s"
    timeout  = "2s"
    failures = 3
  }
}

Create the secret first, then deploy; order within the file never matters:

kanea secret put paste/db-password     # value on stdin
kanea run paste.hcl --wait=90s

Details worth stealing: redis-server --save "" is a command as an argument array with a meaningfully empty argument (R12; a shell string could not say that); the rate limit lives only on the API route, because the frontend serving assets at API rates would be rate-limiting your own pages; and postgres and redis declare no expose and no publish, so they are reachable from this project's services and from nothing else on the network.

Jellyfin with local media

A media server is the stack where the node itself gets a say: the library is a directory the operator already owns, and hardware transcoding needs a GPU device. Both cross the line a job spec cannot cross alone: a host volume does nothing until its path is allowlisted on the node (R15), and the device block names a grant, never a path (R17).

media.hcl
spec_version = 1

project "media" {}

storage "config" {
  type = "local"
}

storage "library" {
  type = "host"
  path = "/srv/media"
}

service "jellyfin" {
  project = "media"

  task "app" {
    image = "jellyfin/jellyfin:10.9.11"

    device "dri" {
      grant = "gpu" # hardware transcoding; the grant is defined on the node
    }

    resources {
      cpu    = 4000
      memory = 4096
    }
  }

  network {
    port "http" { container = 8096 }

    publish "http" {
      host = 8096 # http://<node>:8096, LAN only
      ip_restriction { allow = ["192.168.0.0/16"] }
    }
  }

  volume "config" {
    storage    = "config"
    mount_path = "/config"
  }

  volume "media" {
    storage    = "library"
    mount_path = "/media"
    read_only  = true
  }

  health_check "http" {
    type     = "http"
    path     = "/health"
    port     = "http"
    interval = "15s"
    timeout  = "5s"
    failures = 3
  }
}

The node's half, in the server config: without it the spec is valid and the alloc fails, loudly:

# /etc/kanea/kanea.hcl; the node's, never the repository's
storage {
  allowed_host_paths = ["/srv/media"]
}

device "gpu" {
  nodes = ["/dev/dri/renderD128"]
  allow = ["media"]
}

A failed grant fails the alloc rather than starting without the GPU, because a transcoder silently falling back to software looks healthy and does the wrong thing. The library is mounted read_only (a media server has no business writing to it) while /config is an ordinary local volume Kanea manages. The publish block binds :8096 on the node for the LAN, with the edge enforcing the CIDR allowlist; for access from outside, add an expose block with a domain instead of widening the CIDR.

SMB and S3 volumes

The same volume block, backed by things that are not on the node: a file browser over a NAS share and a read-only bucket. The credential for either driver is one secret whose value is <user>:<secret> (username and password for SMB, access key and secret key for S3) resolved at mount time into a 0600 file, never onto a command line. Omit auth_ref entirely for a public bucket or an open share.

files.hcl
spec_version = 1

project "files" {}

# A share on the NAS. The secret's value is "<username>:<password>".
storage "nas" {
  type     = "smb"
  server   = "192.168.1.20"
  share    = "documents"
  auth_ref = "secret:files/nas"
}

# A bucket, read-only. The secret's value is "<access-key>:<secret-key>".
storage "archive" {
  type     = "s3"
  bucket   = "household-archive"
  endpoint = "https://minio.internal:9000" # omit for AWS S3
  auth_ref = "secret:files/archive"
  mode     = "ro"                          # mountpoint-s3; "rw" selects s3fs
}

service "filebrowser" {
  project = "files"

  task "app" {
    image = "filebrowser/filebrowser:v2.32.0"

    user {
      uid = 1000
      gid = 1000
    }

    resources {
      cpu    = 500
      memory = 256
    }
  }

  network {
    port "http" { container = 80 }
  }

  expose {
    domains = ["files.example.com"]
    tls { mode = "acme" }
  }

  volume "documents" {
    storage    = "nas"
    mount_path = "/srv/documents"
  }

  volume "archive" {
    storage    = "archive"
    mount_path = "/srv/archive"
    read_only  = true
  }

  health_check "http" {
    type     = "http"
    path     = "/health"
    port     = "http"
    interval = "15s"
    timeout  = "5s"
    failures = 3
  }
}

Create both secrets first:

kanea secret put files/nas          # username:password on stdin
kanea secret put files/archive      # access-key:secret-key
kanea run files.hcl

The user block does double duty here (R23, R24): the process runs as 1000:1000, and the volumes inherit that ownership; no uid/gid on the volume needed. On a mounted filesystem there is nothing to chown, so the ownership travels in the mount options instead, which is exactly why it works on a share the NAS controls. The inheritance is resolved at parse time, not on the node, so a spec means the same thing everywhere. (host and nfs volumes are the exception: those drivers cannot carry ownership, so declaring it on one is a plan error, and inheritance skips them.)

mode on the S3 storage selects the driver ("ro" is mountpoint-s3 and the default, "rw" is s3fs) and the storage reference's warning applies in full: an object store is not a filesystem, so keep it for bulk, read-mostly data. A custom endpoint points at MinIO or any S3-compatible store; omit it for AWS. An nfs export is the same shape with server and export. For all of them, a mount that cannot be established fails the alloc loudly rather than starting the service beside an empty directory, and a mount that dies later is supervised and remounted.

A Kafka cluster (KRaft)

Three brokers, no ZooKeeper. Each broker is its own count = 1 service rather than one service with count = 3, and that is the load-bearing decision: a Kafka broker has an identity (a node id and an advertised name that must be stable across restarts) and replicas of one service are deliberately interchangeable. Three services give each broker its own DNS name, its own data volume, and its own line in the quorum.

kafka.hcl
spec_version = 1

project "kafka" {
  description = "Three-broker KRaft cluster, no ZooKeeper"
}

storage "kafka-data" {
  type = "local"
}

service "kafka-1" {
  project = "kafka"

  task "broker" {
    image = "apache/kafka:4.0.0"
    env = {
      KAFKA_NODE_ID                          = "1"
      KAFKA_PROCESS_ROLES                    = "broker,controller"
      KAFKA_LISTENERS                        = "PLAINTEXT://0.0.0.0:9092,CONTROLLER://0.0.0.0:9093"
      KAFKA_ADVERTISED_LISTENERS             = "PLAINTEXT://kafka-1.kafka.kanea:9092"
      KAFKA_CONTROLLER_LISTENER_NAMES        = "CONTROLLER"
      KAFKA_LISTENER_SECURITY_PROTOCOL_MAP   = "PLAINTEXT:PLAINTEXT,CONTROLLER:PLAINTEXT"
      KAFKA_CONTROLLER_QUORUM_VOTERS         = "1@kafka-1.kafka.kanea:9093,2@kafka-2.kafka.kanea:9093,3@kafka-3.kafka.kanea:9093"
      KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR = "3"
      KAFKA_LOG_DIRS                         = "/var/lib/kafka/data"
      CLUSTER_ID                             = "MkU3OEVBNTcwNTJENDM2Qg"
    }
    resources {
      cpu    = 1000
      memory = 2048
    }
  }

  network {
    port "client"     { container = 9092 }
    port "controller" { container = 9093 }
  }

  volume "data" {
    storage    = "kafka-data"
    mount_path = "/var/lib/kafka/data"
  }

  health_check "up" {
    type     = "tcp"
    port     = "client"
    interval = "15s"
    timeout  = "5s"
    failures = 3
  }
}

service "kafka-2" {
  project = "kafka"

  task "broker" {
    image = "apache/kafka:4.0.0"
    env = {
      KAFKA_NODE_ID                          = "2"
      KAFKA_PROCESS_ROLES                    = "broker,controller"
      KAFKA_LISTENERS                        = "PLAINTEXT://0.0.0.0:9092,CONTROLLER://0.0.0.0:9093"
      KAFKA_ADVERTISED_LISTENERS             = "PLAINTEXT://kafka-2.kafka.kanea:9092"
      KAFKA_CONTROLLER_LISTENER_NAMES        = "CONTROLLER"
      KAFKA_LISTENER_SECURITY_PROTOCOL_MAP   = "PLAINTEXT:PLAINTEXT,CONTROLLER:PLAINTEXT"
      KAFKA_CONTROLLER_QUORUM_VOTERS         = "1@kafka-1.kafka.kanea:9093,2@kafka-2.kafka.kanea:9093,3@kafka-3.kafka.kanea:9093"
      KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR = "3"
      KAFKA_LOG_DIRS                         = "/var/lib/kafka/data"
      CLUSTER_ID                             = "MkU3OEVBNTcwNTJENDM2Qg"
    }
    resources {
      cpu    = 1000
      memory = 2048
    }
  }

  network {
    port "client"     { container = 9092 }
    port "controller" { container = 9093 }
  }

  volume "data" {
    storage    = "kafka-data"
    mount_path = "/var/lib/kafka/data"
  }

  health_check "up" {
    type     = "tcp"
    port     = "client"
    interval = "15s"
    timeout  = "5s"
    failures = 3
  }
}

service "kafka-3" {
  project = "kafka"

  task "broker" {
    image = "apache/kafka:4.0.0"
    env = {
      KAFKA_NODE_ID                          = "3"
      KAFKA_PROCESS_ROLES                    = "broker,controller"
      KAFKA_LISTENERS                        = "PLAINTEXT://0.0.0.0:9092,CONTROLLER://0.0.0.0:9093"
      KAFKA_ADVERTISED_LISTENERS             = "PLAINTEXT://kafka-3.kafka.kanea:9092"
      KAFKA_CONTROLLER_LISTENER_NAMES        = "CONTROLLER"
      KAFKA_LISTENER_SECURITY_PROTOCOL_MAP   = "PLAINTEXT:PLAINTEXT,CONTROLLER:PLAINTEXT"
      KAFKA_CONTROLLER_QUORUM_VOTERS         = "1@kafka-1.kafka.kanea:9093,2@kafka-2.kafka.kanea:9093,3@kafka-3.kafka.kanea:9093"
      KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR = "3"
      KAFKA_LOG_DIRS                         = "/var/lib/kafka/data"
      CLUSTER_ID                             = "MkU3OEVBNTcwNTJENDM2Qg"
    }
    resources {
      cpu    = 1000
      memory = 2048
    }
  }

  network {
    port "client"     { container = 9092 }
    port "controller" { container = 9093 }
  }

  volume "data" {
    storage    = "kafka-data"
    mount_path = "/var/lib/kafka/data"
  }

  health_check "up" {
    type     = "tcp"
    port     = "client"
    interval = "15s"
    timeout  = "5s"
    failures = 3
  }
}

The quorum voters and advertised listeners are written as literal internal DNS names (<service>.<project>.kanea), not ${service.…} references: deliberately. A reference is also a start dependency (R10), and three brokers referencing each other is a cycle kanea plan rejects (R9). Kafka's quorum is designed to form as the peers come up, so the ordering edge is not wanted; the literal names are exactly the strings the interpolation would have produced, minus the edge. Note the shared storage "kafka-data" block is safe: each service's volume gets its own directory beneath it, so the brokers never share a log dir.

A client in the same project bootstraps with all three names: kafka-1.kafka.kanea:9092,kafka-2.kafka.kanea:9092,kafka-3.kafka.kanea:9092; a service in another project also needs allow_from naming it on each broker (R14). The cluster is deliberately not exposed north-south: Kafka's protocol hands clients the advertised names to dial, and those resolve only inside the node; publishing :9092 would let an external client connect once and then fail on redirect. Generate your own CLUSTER_ID (kafka-storage.sh random-uuid) rather than shipping the example's.

kanea run kafka.hcl --wait=120s
kanea ps --project=kafka
kanea logs kafka/kafka-1 --tail=50    # watch the quorum form

CLI reference

One binary, one command tree. The CLI talks to kanead over a unix socket (--socket overrides it on every client command) and the daemon commands are the ones systemd runs for you.

The socket is root-owned, so client commands run under sudo, or without it, after joining the kanea group init creates and logging in again: sudo usermod -aG kanea <user>. Membership is root-equivalent, exactly like docker's group, and is never granted by Kanea itself.

Services are addressed as project/service throughout, or as a bare service name with --project.

Every client command also takes a remote endpoint instead of the socket: --url / $KANEA_URL, --token / $KANEA_TOKEN and --ca-cert / $KANEA_CA_CERT. That is how a laptop or a CI runner drives a node it is not sitting on; see Working with a remote node for the setup and the pipeline examples.

The CLI also installs through Homebrew, on Linux and on macOS. kanea plan parses and validates a job spec with file-and-line diagnostics before anything dials a socket, so authoring needs no daemon at all - and since v1.82 the rest of the CLI is not tied to the socket either: with KANEA_URL and a token it drives a remote node completely. A Mac is a first-class client; it is just never the node, which is Linux. See Working with a remote node.

Setup

kanea init

Interactive first install: preflight checks, configuration, the master-key ceremony, the systemd units, and, since v0.5, the rest of the way to a working platform: it asks for the dashboard's listen address (loopback by default; none keeps the API socket-only), starts kanead, creates the first admin account over the local socket, and ends with a summary of what it built; the dashboard URL, your account, the internal DNS address and the subnet layout. Run it once, as root.

sudo kanea init [--data-dir=/var/lib/kanea] [--log-dir=/var/log/kanea/allocs]
                [--unit-dir=/etc/systemd/system] [--prefix=/usr/local/lib/kanea]
                [--network=ebpf|netns] [--containerd=…|external] [--buildkit=…|off]
                [--node-cidr=10.244.0.0/24] [--cluster-cidr=10.244.0.0/16]
                [--node-cidr6=…] [--cluster-cidr6=…] [--service-cidr6=…]
                [--listen=127.0.0.1:8600|none] [--listen-cert=…] [--listen-key=…]
                [--reserve=256M] [--admin-user=…] [--timeout=2m]
                [--bundle=…] [--no-install] [--no-start] [--skip-checks] [--skip-units]

A full walk-through, with an annotated run and what a re-run keeps, is in Installing → Setting up the node. Note there is no --service-cidr here: the v4 service pool is fixed at 10.201.0.0/16 and is never rendered into the unit; change it by editing kanead.service.

A non-loopback --listen is served over TLS, and since v0.23 init provisions it: with no explicit --listen-cert/--listen-key pair it mints a ten-year self-signed pair at /etc/kanea/api.crt and api.key (once; a pair already there is left alone) and points the unit at it. Bring your own with the flags, or take the listener over with the bind stanza below. The one refusal left is an unspecified host (0.0.0.0, :8600): a certificate's SAN needs a host to name. --admin-user plus a piped password makes it scriptable; --no-start writes the files and stops, which is the pre-v0.5 behaviour. Re-running is safe: an existing master key and an existing account are left alone.

Since v1.61 the listener can live in the server config instead: a bind { api_addr = … api_tls = … } stanza in /etc/kanea/kanea.hcl, where api_tls is acme, self-signed, provided or plaintext; the same modes services use (the full field set is in Exposing the API and dashboard). When it is declared and --listen was not passed, init skips the listen question and renders no listen flags into the unit: the file owns the listener, and moving the API and dashboard later is an edit to the file plus systemctl restart kanead, never a re-init.

The key is shown once

The ceremony prints the master key and requires you to type it back; if that fails it is discarded and nothing is written. Without that key every backup this node ever makes is unreadable. Have somewhere to record it before you start.

kanea doctor

kanea doctor [--data-dir=…] [--prefix=…] [--network=ebpf|netns] [--containerd=…]
             [--buildkit=…|off] [--node-cidr=…] [--cluster-cidr=…] [--service-cidr=…]
             [--offline]

Verifies the node: dependencies and their versions against the pinned matrix, the containerd socket, bpffs and the cgroup2 mount, kernel version, cgroups v2 and slice placement, the effective memory floor, subnet overlap, the state database's permissions, the build socket, disk headroom and clock synchronisation. It also names known interference, like a host firewall dropping on the hooks alloc traffic crosses (docker, ufw, firewalld). Safe to run any time.

A failing check exits non-zero; warnings do not. A check this user may not perform - the containerd socket, the bpf pin root and the state database are all root-owned - is reported SKIP rather than FAIL, and the run says once at the end how many need root. "I could not look" is not "it is broken", and reporting it as one sends an operator to reinstall a healthy node. --offline skips the one network probe, which is a reachability check for the pinned artefacts rather than a requirement - on an air-gapped node it is the expected answer, not a problem.

kanea firewall

kanea firewall [--manager=ufw|firewalld|nft|iptables] [--all]
               [--prefix=…] [--node-cidr=…] [--cluster-cidr=…] [--network=ebpf|netns]

Prints the host-firewall rules this node's workloads need, derived from this node's own cluster CIDR and resolver address rather than from an example. Two allowances, because alloc traffic crosses two hooks a firewall owns: forward, on its way off the node, and input, to reach the internal resolver - a query to the resolver is a new inbound connection to the host on a veth. With no argument it prints for the manager that appears to own the ruleset; --all prints every one.

It prints and never applies. Kanea owns exactly the kanea nftables table and writes nothing outside it: a rule placed in a manager's ruleset is flushed away by that manager on its next reload, so applying one would be a fix that silently stops being applied. A published port needs its own inbound allow, which this command cannot know without asking the daemon.

kanea install

kanea install [--list] [--dry-run] [--only=containerd,runc,…] [--force]
              [--bundle=…] [--containerd=external] [--arch=amd64|arm64]
              [--prefix=…] [--conf-dir=…] [--data-dir=…] [--run-dir=…]
              [--unit-dir=…] [--skip-units]
              [--node-cidr=…] [--cluster-cidr=…] [--node-cidr6=…] [--cluster-cidr6=…]

Places the pinned host components: containerd, runc, rootless buildkitd and the wasmtime shim. kanea init runs this for you, so you normally only reach for it directly to inspect (--list, --dry-run) or to repair one component (--only, --force). Versions and SHA-256 hashes are compiled into the binary and never fetched. Full notes are in Installing → Host components.

kanea bundle

kanea bundle create [--arch=amd64|arm64] [-o=…] [--dir] [--no-images] [--containerd=…]

Builds an offline bundle of the host components on a machine that has a network, for a node that does not. --dir writes a directory instead of a tarball and --no-images omits the image components (leaving the binaries alone). The bundle carries no hashes of its own: it is verified against the installing node's binary. See Air-gapped nodes.

kanea ca

kanea ca show [--out=…] [--quiet]   # the CA certificate in PEM
kanea ca info                       # its subject, fingerprint and validity

show prints the node's own CA certificate, which is what devices need in order to trust self-signed services and the dashboard: kanea ca show > kanea-ca.crt. There is deliberately no kanea ca rotate and no route that returns the CA key: rotation means re-trusting every device, and a command for it would imply it is cheap.

kanea upgrade

kanea upgrade [--check] [--version vX.Y.Z] [--require-signature] [--no-fetch] [--skip-backup] [--dry-run] [--timeout=2m]

One command, both halves: it fetches the latest release (or --version), verifies it; sha256 against the release's checksums.txt always, the cosign keyless signature over that file when cosign is installed, a loud note when it is not (--require-signature turns both soft endings into refusals); installs it atomically over its own path, then takes a pre-upgrade backup, restarts kanea-edge and then kanead in that order, runs any state migrations, and waits for health. Already at the target version means nothing to download, so running it twice is safe by construction.

Owned by a package manager?

--no-fetch restarts onto whatever binary is already installed (the orchestration half alone) and --check only reports the running, installed and latest versions. Air-gapped nodes get their binary from the offline bundle flow; the fetch refuses with that pointer rather than hanging.

One thing it deliberately never does: rewrite systemd units. When release notes say the units changed, re-run sudo kanea init after upgrading (idempotent: the master key, accounts and settings are kept) then sudo systemctl daemon-reload.

Deploying

kanea plan

kanea plan app.hcl [more.hcl …] [selector …]
kanea plan --image=nginx:1.27-alpine --name=web --project=demo [--count=1]

A real dry run: one block per service, one row per resource that would be added, changed or removed, the resulting workload budget, and every validation rule. --remove-orphans adds the - destroy lines kanea run --remove-orphans would act on, computed from the same scope, so the plan and the run cannot disagree. Multiple files are parsed as one set, so file order does not matter. This is where you find out about a cycle, a cross-project secret, a duplicate domain or a rate limit that would fail open. It also warns, never fails, on the weak defaults: an effective uid of 0, a writable root filesystem, and an image following a moving tag without opting into auto-update - the shapes hardening = "restricted" and read_only_rootfs exist to replace.

$ kanea plan app.hcl
~ update shop/web
    image             web:v1 -> web:v2  (rolls allocs)
    env               + DB_URL, ~ LOG_LEVEL  (rolls allocs)
    volumes           + cache  s3:backups /var/cache rw 10 GiB budget  (rolls allocs)
                      ~ data  local /data rw -> local /srv/data rw
    expose            + api.example.com  (tls acme, port 8080)
    publish           - 8443/tcp  port http
    scaling           min 1, max 5, rps 100 -> min 1, max 10, rps 100

+ create shop/worker (count 2, image worker:v3)
    volumes           + queue  local /queue
    files             + /etc/app.conf  (412 B, content 3f2a1b9c, mode 0644)
    check             http /healthz on port 8080 every 10s, timeout 2s, 3 failure(s)

Plan: 2 change(s) - 1 create, 1 update; 1 replace running allocs.
Run `kanea run` to apply.

Rows marked (rolls allocs) replace running containers, the ones without are applied to the service as it is or republished to the edge. That is the difference between an edit that costs a rolling restart and one that costs nothing, and it is the number the summary line counts. It is derived from the same rule the reconciler uses to decide a deploy, so the plan and the daemon cannot disagree about which is which.

What a plan never prints

An environment value, because it may be a secret-env: reference, and a config file's content, because it carries secret placeholders. You get the keys, the paths, and a short digest of the content: enough to see that something changed, never enough to leak it into a terminal scrollback or a pasted issue.

kanea run

kanea run app.hcl [more.hcl …] [selector …] [--wait=60s] [--yes|-y]
kanea run --image=nginx:1.27-alpine --name=web --project=demo [--count=1]

Applies the spec. kanea apply is an alias: same flags, same behaviour. --wait is how long to wait for allocs to reach running before returning; 0 returns immediately. Running it twice with an unchanged spec does nothing: a deploy is a spec-hash mismatch, not an invocation.

It shows the plan and asks first. The block above is printed by the same code kanea plan uses, so what you confirm is exactly what you were shown, and then:

Apply? [Y/n]

Enter applies. Anything that is not y or yes aborts and nothing is sent. --yes (or -y) skips the question.

A script is never asked a question

The prompt appears only when stdin is a terminal. A piped or redirected stdin (every CI job, every pipeline, every kanea run … < /dev/null) applies without asking, exactly as it did before the prompt existed. No existing automation needs --yes added to it.

An apply is additive. A service the spec no longer declares keeps running, because kanea run on one file must not delete what another file declares. --remove-orphans opts out of that, for the projects this spec declares a project block for:

kanea run --remove-orphans app.hcl

Anything stored in one of those projects that the spec does not declare is deleted, in the same atomic batch as the applies, so a rename never leaves both or neither in place. It is refused with a selector (kanea run app.hcl shop/web --remove-orphans) and with --image: a selector sends part of the spec and --image declares no project, so neither can honestly claim to be the whole of anything. What a prune destroys is worth reading once.

--image creates a bare service and cannot update a real one. It builds a record from project, name, count and image alone, so applying it over a service that declares ports, expose, env, volumes, a health check or scaling would delete them. It refuses when that would happen, naming what would be lost; kanea deploy is the verb for changing the image of a service that already exists.

A selector scopes both commands to part of the file: kanea run app.hcl shop/web applies one service, shop alone a whole project, and several selectors union. An argument that exists on disk is a spec file; only a non-existent one is read as a selector, and one that is neither is refused by name. The whole file is still parsed and validated (a selector never changes what a spec means, only how much of it is sent) and every selector must match at least one service in the file. An apply is additive either way: services not in the request are never touched, so a scoped run cannot delete anything.

kanea deploy

kanea deploy [--project P] <[project/]service> <image> [--wait 60s] [--no-wait]

Point an existing service at a new image and leave the rest of its spec alone. It reads the record, changes the image and writes the whole record back, because there is no route that sets an image on its own and round-tripping is what stops a deploy dropping a field. An init step declaring the task's previous image moves with it (v0.31.1) - a migration on the app's own image must not run yesterday's bytes against today's application - and the output names the steps that followed; a step on any other image is untouched. Waits for the new image to be running by default, which is what makes it usable as a pipeline's last step. This is the CI verb; see Deploying a new image.

kanea stop

kanea stop [--project=p] <[project/]service> [--rm]

Scales to zero. --rm also deletes the service declaration, behind the same confirmation kanea remove asks; a piped stdin is never prompted, so scripted removals run unchanged.

kanea remove

kanea remove [--project=p] [--yes] <[project/]service>

Deletes one service declaration (alias: rm). Asks [y/N] on a terminal; --yes/-y skips. The containers, alloc records, VIP, routes and mounts go; volume data is kept, so re-applying the spec brings the service back with its data.

kanea start

kanea start [--project=p] <[project/]service> [count]

stop's counterpart: scales a stopped service back up. The daemon does not remember the pre-stop count (a stopped record says zero) so it starts one replica unless a count is given, or an autoscaled service's own floor (the scale route refuses a count outside the declared bounds). A service already running is left exactly as it is: start is idempotent, never a second spelling of scale.

kanea restart

kanea restart [--project=p] <[project/]service>

Rolls the service's allocs through its update policy: a generation bump, the same route the dashboard uses, not a second path into the runtime. It is also the way out of an exhausted crash-restart budget: the bump is a new spec hash, and the restart count belongs to the hash that spent it.

Inspecting

kanea ps

kanea ps [--project=p] [--service=s] [-a]

The alloc table: id, service, state, health, restarts, age, address. A removed alloc leaves no record (only failed-and-still-declared ones persist to explain themselves), so a stopped service is invisible here; -a adds what is declared but not running: services scaled to zero (stopped) and slots the reconciler has not created yet (pending).

kanea describe

kanea describe [--project=p] <[project/]service>

One service in full: the declared spec beside what is actually true; image and its pinned digest under auto-update, routes from every expose block and published port, volumes and grants, the alloc table with health verdicts, a stats snapshot, and the service's recent events. Stats and health render absent as absent (-), never as zero: a missing metric and an idle service are different facts.

The alloc table's REASON column says why each alloc last stopped; OOMKilled, Signalled, Error, Completed, or why one that never ran did not: ImageFailed, VolumeFailed, GrantFailed, NetworkFailed, CreateFailed, StartFailed. An OOM kill is read from the alloc's cgroup, never guessed from the exit code (a forced stop exits 137 exactly like a memory kill does) so a stop reports as Signalled, and a kill against a service that declared no memory limit names the node's collective ceiling rather than a number nobody typed. The reason shows whatever the alloc's current state is: a running row carrying OOMKilled means it is up now and was killed for memory last time.

kanea status

kanea status [--project=p] [[project/]service]

Health, recent events, current and desired counts, and the scaling picture.

kanea logs

kanea logs [--project=p] <[project/]service> [-f] [--tail=N] [--alloc=ID] [-c NAME] [--previous]

Merged across allocs by default; --alloc narrows to one. --tail shows the last N lines before following. -c reads an init container's log instead of the task's, by its block name. --previous reads the log files a stopped or removed service left on disk - a torn-down alloc keeps no record, but its log file survives, so the last output before a stop, a crash-loop's end, or a removal is still readable. Log drains are non-blocking with drop counters: a slow reader can never stall a workload's write().

kanea exec

kanea exec [--project=p] [--alloc=ID] [--user=UID] [-it] <[project/]service> -- <command…>

A debug shell inside an alloc. Admin-only, and audited whether or not the session establishes: "someone tried to open a shell on production" is worth keeping either way, so the attempt and the requested command are both recorded.

  • The -- is required. The command crosses the wire as separate arguments rather than one joined string, because every rule for splitting a string back into arguments is wrong for something somebody will eventually pass.
  • -it allocates a terminal and forwards stdin. A shell needs it.
  • --user takes a numeric uid only. Resolving a name would mean reading the container's own /etc/passwd, and a container-controlled file deciding which uid the control plane runs a process as is not a thing to build.

kanea ui

kanea ui [--addr=…] [--open]

Prints the dashboard URL; --open launches a browser.

Scaling and builds

kanea scale

kanea scale [--project=p] <[project/]service> <count>

Writes the desired count and returns; the reconciler converges. This is the same route the autoscaler uses, which is why manual and automatic scaling can never disagree about mechanism.

kanea build

kanea build [--project=p] <[project/]service> [--deploy=true] [--follow=true]

Triggers the service's build pipeline. --deploy rolls the built digest out on success; --follow streams the build log. Builds are serialised: a second one is queued, and refused rather than blocked when the queue is full.

kanea project

kanea project sync <project>             # re-read the git source now
kanea project builds <project> [--service=s] [--limit=N]

kanea images

kanea images [--clean] [--json]

The node's containerd images, per project, with each image's size, age and whether anything still references it. --clean runs one image GC sweep now (admin-only, audited) and reports what it removed and reclaimed; the rules it deletes under, and the gc block that tunes them, are in the node config reference. On a node whose default pull_policy is "never" the sweep answers with a refusal naming why: preloaded images cannot be re-pulled.

kanea volume

kanea volume list [--json]

Every storage resource with the mounts using it nested underneath: driver, where the bytes live, measured usage, declared budget and mount state. It is what to reach for when a disk is filling up - a flat list would repeat an NFS export's address once per service and still not say they were the same export.

Usage is sampled in the background, so a volume reads - until it has been measured, and s3 volumes are never walked. That dash is an absence, not a zero: a volume nobody has looked at and an empty one are different facts. A local volume appears once per alloc, because that is how many of it there are - each with its own contents, and each judged against the budget separately.

There are no other subcommands, deliberately. A volume exists because a spec declares it, so creating or deleting one here would be a second way to change desired state that the reconciler would immediately undo.

Secrets and accounts

kanea secret

kanea secret put [--from-file=path] <project>/<name>   # value on stdin
kanea secret ls [<project>]
kanea secret rm <project>/<name>
There is no get

Not for an operator, not over the API, and not for an AI agent at any tool tier. ls lists names. The value goes in and is only ever resolved into a running alloc.

kanea user and kanea token

kanea user add [--role=admin|viewer] <name>
kanea user ls
kanea user rm <name>                  # also revokes the account's sessions
kanea user revoke-sessions <name>     # sessions only; the account stays

kanea token create [--role=viewer] [--expires-in=720h] <name>
kanea token ls
kanea token rm <id>

Accounts live in the Store, not in a config file. Tokens default to never expiring; --expires-in takes a Go duration. The first admin is created by kanea init itself; OIDC and LDAP identities never appear in user ls: they are ephemeral, a session and nothing else.

Backup and restore

kanea backup create [--reason="on-demand"]
kanea backup list
kanea backup verify <archive-id>

kanea restore --from s3://bucket/prefix [--snapshot=ID] [--target=path]
              [--s3-endpoint=…] [--s3-region=…] [--s3-access-key=…] [--s3-path-style]
              [--master-key=path] [--data-dir=…]

verify reads the archive and checks its hashes and authentication tags without restoring anything: an archive that cannot be verified is one you find out about now rather than during an outage.

A restore is staged, never performed in place

The command stages the restore; it is performed at the next daemon start, before anything opens the Store. That is the interface rather than a safety check: the API has no method that restores at all, and there is no restore button in the dashboard, because a restore replaces everything on the node and belongs at a terminal.

Daemons

Normally systemd runs these. The flags are here because systemctl cat will show them to you.

kanea agent

kanea agent [--config=/etc/kanea/kanea.hcl|off] [--data-dir=…] [--log-dir=…]
            [--volume-dir=…] [--socket=/run/kanea/kanead.sock]
            [--containerd=…] [--network=ebpf|netns] [--node-cidr=…] [--cluster-cidr=…]
            [--service-cidr=10.201.0.0/16] [--bpf-dir=/sys/fs/bpf/kanea]
            [--dns-listen=…|off] [--dns-upstream=…] [--registry=127.0.0.1:5100|off]
            [--allowed-host-paths=…|off] [--passthrough-config=…|off]
            [--listen=…|none] [--listen-cert=…] [--listen-key=…]
            [--base-domain=…] [--tls-default=acme|self-signed|provided|plaintext]
            [--tls-certs-config=…] [--acme-email=…] [--edge-group=kanea-edge]
            [--publish-ports=1024-65535] [--secrets-providers-config=…]
            [--dashboard=true] [--autoscale=true] [--log-level=info]
            [--backup-s3-endpoint=…] [--backup-s3-region=…] [--backup-s3-access-key=…]
            [--oidc-client-id=…] [--ldap-url=…] [--acme-dns-tsig-key=…]

--listen beyond loopback requires --listen-cert and --listen-key. Unset, the listener comes from the server config's bind stanza when one is declared (below); --listen none forces socket-only regardless of the file. Credential-shaped options such as --oidc-client-secret, --ldap-bind-password and --acme-dns-tsig-secret take secret: references, never literals.

--ldap-url enables directory logins beside local accounts and OIDC: ldaps:// (or ldap:// with StartTLS forced; there is no insecure option), a user search under --ldap-user-base-dn with --ldap-user-filter, and group-to-role mapping through --ldap-admin-groups/--ldap-viewer-groups, deny-by-default. A local account with the same name always wins, and the login rate limit runs before any bind reaches the directory.

The operator-owned settings live in the server config, /etc/kanea/kanea.hcl: which directories host volumes may come from, which devices and sockets are granted to which projects, where the API and dashboard listen, which resolvers the internal DNS forwards to, and node-wide spec variables. It is read once at startup, refused if anyone but its owner could have written it, and absent by default. The complete reference - every stanza, every field, the precedence rules and how to apply an edit - is Node configuration.

Every one of those has a flag that overrides it, and the disable words differ (--listen none, but off elsewhere). The full precedence table is in Node configuration → Flags and precedence.

What init writes into the unit, and what it does not

kanea init renders exactly seven of these flags into ExecStart: --data-dir, --log-dir, --network, --node-cidr, --cluster-cidr, --edge-group, and --listen (with its cert pair, and only when the bind stanza does not own the listener). The three *6 flags join them only when IPv6 is enabled, so a v4-only unit stays byte-identical.

Everything else on this page is set by editing the unit (systemctl edit kanead) or by the server config. That includes two that look like they should be init's, because init has flags by the same name and uses them only for the install: --containerd and --buildkit. A non-default containerd socket or buildkit address needs a unit edit.

kanea edge

kanea edge [--routes=/run/kanea-edge/routes.json] [--certs=…] [--http=:80]
           [--poll=…] [--drain=…] [--memory-limit=128MiB] [--log-level=info]

Its own process and its own systemd unit, with no After=kanead.service: north-south traffic surviving a control-plane restart is the entire reason it is separate. --drain is how long in-flight requests get on shutdown.

MCP

kanea mcp [--socket=/run/kanea/kanead.sock] [--verbose]

A stdio MCP server for an AI agent running on the node: 24 tools in read, mutate and destructive tiers. The same server is available over streamable HTTP at /mcp on the API listener for agents running anywhere else. Setting it up for Claude Code, opencode or Codex is AI agents (MCP), which has the client configuration for both transports.

Tools reach the platform only by making requests against the API's own handler, so an agent is never more privileged than the credential it was given. Tiers are advertised as well as enforced, and the advertisement fails closed. A refusal comes back as a tool result rather than a protocol error, because the model is what has to react to it.

kanea version

Prints the version stamped in at build time. kanea upgrade compares it against what the running daemon reports.

AI agents (MCP)

Kanea ships a Model Context Protocol server, so an agent can look at the node and operate it through the same API you do. It speaks both transports: stdio, for an agent running on the node itself, and streamable HTTP on the API listener, for one running anywhere else.

There is nothing to install. The server is the same kanea binary, and the tools reach the platform only by making requests against the API's own handler - which is what makes the next section a statement about capability rather than about good intentions.

An agent is never more privileged than its credential

A tool's only verb is "send this request", so nothing in the MCP server can hold a Store, a secrets store or an auth store, and no tool can do something the token behind it could not do at the REST API. Which tools an agent can even see follows from that token's role, and there is no secrets tool at any tier - not a redacted one, none.

Which transport

stdioStreamable HTTP
Command / endpointkanea mcphttps://<node>:8600/mcp
Agent runsOn the nodeAnywhere
Authenticates byUnix socket accessA bearer token
Needsroot, or kanea group membershipAn API listener (--listen or a bind stanza)
RoleAlways admin: the socket is root-equivalentWhatever the token has

The practical rule: if your editor runs on the node, use stdio. If you are on a laptop and the node is a server - which includes every Mac, since kanead is Linux-only - use HTTP. Prefer HTTP when you want the agent to be read-only, because that is the transport where you can choose a role.

stdio, on the node

kanea mcp [--socket=/run/kanea/kanead.sock] [--verbose]

It talks to kanead over the unix socket and speaks MCP on stdin and stdout. --verbose logs protocol activity to stderr; stdout is the protocol channel, and a single stray line on it corrupts the stream in a way that presents as a client which connects and then does nothing.

The socket is root-owned, and this is the setup step people miss

Your editor runs as you, not as root, so kanea mcp fails to reach the daemon unless you have joined the group kanea init created:

sudo usermod -aG kanea $USER   # then log out and back in

Membership is root-equivalent, exactly like docker's group. An agent on this transport is therefore always an admin, with every tool available to it. If that is more than you want to hand over, use HTTP with a viewer token instead.

Claude Code

claude mcp add kanea -- kanea mcp

Add --scope project to write it into the repository's .mcp.json instead of your own config, so everyone working on that project gets it. The file form is the usual one:

{
  "mcpServers": {
    "kanea": {
      "command": "kanea",
      "args": ["mcp"]
    }
  }
}

opencode

In opencode.json at the project root, or your global ~/.config/opencode/opencode.json:

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "kanea": {
      "type": "local",
      "command": ["kanea", "mcp"],
      "enabled": true
    }
  }
}

Codex

codex mcp add kanea -- kanea mcp

Which writes to ~/.codex/config.toml; the equivalent by hand is:

[mcp_servers.kanea]
command = "kanea"
args    = ["mcp"]

Streamable HTTP, from anywhere

The same server is mounted at /mcp on the API listener - deliberately not under /v1, because that path is this API's versioned surface and MCP carries its own version in the initialize handshake. It exists only when the node has an API listener at all: a socket-only node (--listen none) has no HTTP transport to offer.

Mint a token for the agent, and choose its role deliberately:

sudo kanea token create --role viewer agent-readonly    # can look, cannot touch
sudo kanea token create --role admin  agent-deploy      # can deploy, scale and stop
sudo kanea token create --role viewer --expires-in 720h agent-30d

The secret is printed to stdout once and is not recoverable, so … > token.txt captures exactly it while the commentary goes to stderr. kanea token ls shows what exists and kanea token rm <id> revokes one.

Claude Code

claude mcp add --transport http kanea https://192.168.1.10:8600/mcp \
  --header "Authorization: Bearer $KANEA_TOKEN"

opencode

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "kanea": {
      "type": "remote",
      "url": "https://192.168.1.10:8600/mcp",
      "enabled": true,
      "headers": { "Authorization": "Bearer <token>" }
    }
  }
}

Codex

[mcp_servers.kanea]
url = "https://192.168.1.10:8600/mcp"
bearer_token_env_var = "KANEA_TOKEN"
Two things that will bite you here

A JSON entry with a url and no type is a configuration error, not a remote server. Claude Code reads a typeless entry as stdio and skips it with a message saying so; opencode needs "type": "remote" for the same reason. Copying an mcpServers block from another client is where this usually happens.

If the node serves a self-signed certificate - which it does by default, whether from kanea init's provisioned pair or the node CA - a client that does not trust it will fail the TLS handshake before MCP is ever spoken. Install the CA with kanea ca show on the machine running the agent, or point the client at a name covered by an acme certificate.

The tools

Twenty-five tools in three tiers. The tier an agent gets is its credential's role, and the listing is filtered, not merely refused: a viewer token does not see the mutating tools in tools/list at all, so the model cannot propose calling one.

TierNeedsTools
Read
13 tools
viewerlist_projects, get_project, list_services, get_service, list_allocs, get_logs, get_events, get_node_stats, get_service_stats, list_pipelines, list_storage, list_backups, get_audit*
Mutate
9 tools
adminplan_spec, apply_spec, scale_service, restart_service, stop_service, deploy_service, run_pipeline, create_backup, test_notification
Destructive
3 tools
admin, plus an explicit confirmrestore_backup, delete_service, delete_project

* get_audit is the one read tool a viewer cannot actually use: the audit log is admin-only at the API, so the tool is listed but the call comes back 403. That is the tier system being advisory and the API being authoritative, in the one place where they disagree.

A destructive tool called without confirm: true is refused with a result telling the model to call it again only if the operator has explicitly asked for this, and that it cannot be undone. The gate is enforced in the MCP server rather than the API, because it is not an authorization rule - an admin is allowed to delete a project and the API will let them. It is a rule about agents: a destructive action has to be arrived at deliberately rather than by pattern-matching a tool name, and a human reading the transcript afterwards can see that it was. A refusal comes back as a tool result rather than a JSON-RPC protocol error, because a protocol error is handled by the client library and never reaches the model, which is the one party that has to react to it. An unknown tool is the exception, and deliberately so: that is a bug in the client, not a refusal for a model to reason about.

Every call lands in the audit log with the token's id, exactly like a REST call, because it is one. The dashboard's audit tab, GET /v1/audit and the get_audit tool all show it - which means an agent can read its own trail, and so can you.

Deciding what to hand over

  • Start with a viewer token. It covers the whole diagnostic loop - what is deployed, what is failing, the logs, the events, the stats, the audit trail - which is most of what an agent is actually useful for, and it cannot change anything.
  • Give admin only when you want the agent deploying. apply_spec, scale_service and deploy_service are real deploys, with the same rolling-update rules and health gates a human's would have; nothing about a change arriving from an agent makes it a different kind of change.
  • Prefer HTTP with a scoped token over stdio when the distinction matters. The stdio transport authenticates by socket access, and socket access is root-equivalent, so there is no read-only version of it.
  • Set an expiry. kanea token create warns when a token never expires, for good reason: --expires-in 720h costs nothing and bounds the blast radius of a leaked editor config.
  • Secrets are unreachable by construction, so do not plan around them. There is no tool that reads one at any tier; a spec references secret:<project>/<name> and the agent never sees a value.

Working with a remote node

The CLI reaches kanead two ways. Locally it uses the unix socket, and your credential is membership of the kanea group. From anywhere else it uses the node's HTTPS listener, and your credential is a bearer token. Everything below is the second one - a laptop, or a CI runner that has just built an image.

LocalRemote
Reaches the daemon bythe unix sockethttps://node:8600
Credentialkanea group membershipa bearer token
Set up withusermod -aG kanea $USERKANEA_URL, KANEA_TOKEN, KANEA_CA_CERT
Get the CLI fromthe install scriptHomebrew, or a release archive
Authorized asadmin, alwaysthe token's role

A remote CLI is not a new privilege. It presents the same token the REST API and the MCP server have always accepted, and the role decides exactly what it decided there: a viewer token can look and cannot deploy. What was missing until v1.82 was a client that spoke it, not a permission.

Setting it up

1. Get the CLI

On a laptop, or in a CI image, Homebrew is the easy channel - it needs no root and no node:

brew tap m18h/kanea
brew trust m18h/kanea   # brew ≥ 6 refuses formulae from untrusted third-party taps
brew install kanea

For a container base without Homebrew, take the release archive; it is one static binary and needs nothing installed beside it. The verify-by-hand steps are the same.

2. Mint a token, on the node

sudo kanea token create --role admin ci --expires-in 720h

The secret goes to stdout once and is not recoverable, so … > token.txt captures exactly it while the commentary goes to stderr. Deploying needs --role admin; a viewer token can read everything and change nothing, which is the right choice for a job that only inspects. kanea token ls shows what exists and kanea token rm <id> revokes one, immediately.

Set an expiry

kanea token create warns when a token never expires, for good reason: --expires-in 720h costs nothing and bounds how long a leaked CI variable is worth anything. Rotating is one command on each side.

3. Trust the node's certificate

A Kanea node serves its own CA's certificate or a self-signed one unless you gave it an acme name, so a fresh client does not trust it and the TLS handshake fails before anything else happens. Export the CA once:

sudo kanea ca show > kanea-ca.crt        # on the node
export KANEA_CA_CERT=/path/to/kanea-ca.crt

KANEA_CA_CERT takes a file path or the PEM itself, because a CI system hands secrets to a job as values and writing them to a file first is a step that gets skipped. Both of these work:

KANEA_CA_CERT=/etc/kanea/ca.crt
KANEA_CA_CERT="$(cat kanea-ca.crt)"     # or a CI variable holding the PEM

If the node's certificate comes from Let's Encrypt (an acme bind stanza), skip this entirely: the public roots already trust it. There is deliberately no --insecure-skip-verify; the whole reason the inline form above exists is that a skip flag is what people reach for when the honest path is awkward.

4. Point the CLI at the node

export KANEA_URL=https://kanea.apps.example.com:8600
export KANEA_TOKEN=$(cat token.txt)

kanea ps
kanea status
kanea logs -f shop/web

Every client command takes these, as environment variables or as --url, --token and --ca-cert. A flag you actually pass beats the environment; an explicit --socket keeps a command local even with KANEA_URL exported, so a node's own shell is never redirected by a variable someone set for another tool.

Deploying a new image

kanea deploy [--project P] <[project/]service> <image> [--wait 60s] [--no-wait]

kanea deploy points an existing service at a new image and leaves everything else exactly as declared. It reads the service's record, changes the image, and writes the whole record back - which matters, because there is no API route that sets an image on its own, and round-tripping the record is what makes it impossible for a deploy to drop a field it does not know about. One deliberate exception rides along (v0.31.1): an init step whose image is exactly the task's previous one is pointed at the new image too, so a migration that runs the app's own image deploys in lock-step with the app. Byte-for-byte equality on the declared reference, nothing cleverer; a step on its own image never moves. The same rule applies when a GitOps build deploys its result, so a step that starts on the app's image stays on it build after build.

kanea deploy shop/web ghcr.io/acme/web@sha256:9f2a…

It waits for the new image to be running and fails if it does not, so a pipeline goes red at the deploy rather than the next morning. --no-wait returns as soon as the change is accepted. Deploying the image a service already declares is reported and does nothing, so re-running a pipeline on an unchanged commit is not an error.

Prefer a digest to a tag. A tag can move under you, so two allocs replaced a minute apart can run different code; a digest is the thing that was built.

Not kanea run --image

That flag builds a service from nothing - project, name, count, image - so applying it over an existing service used to delete its ports, expose block, env, volumes, health check and scaling, silently, and you found out when traffic stopped arriving. It now refuses and names what would be lost. Use it to create a bare service; use kanea deploy to change one that exists.

In a pipeline

The CLI ships as a container image, ghcr.io/m18h/kanea, which is what a runner wants: linux/amd64 and linux/arm64 in one manifest list, tagged vX.Y.Z and latest. Pin the version - that is what tags are for. See the container image for what is in it and how to verify it.

GitLab:

.gitlab-ci.yml
deploy:
  stage: deploy
  image:
    name: ghcr.io/m18h/kanea:vX.Y.Z           # pin the version
    entrypoint: [""]                          # GitLab runs the job script as the container command
  variables:
    KANEA_URL: https://kanea.apps.example.com:8600
  script:
    - kanea deploy shop/web "$CI_REGISTRY_IMAGE@$IMAGE_DIGEST"
  # KANEA_TOKEN: masked project variable
  # KANEA_CA_CERT: a file variable, or the PEM as a masked variable

GitHub Actions:

.github/workflows/deploy.yml
- name: Deploy
  # or run the image directly: docker run --rm -e KANEA_URL -e KANEA_TOKEN \
  #   ghcr.io/m18h/kanea:vX.Y.Z deploy shop/web "$IMAGE"
  env:
    KANEA_URL: https://kanea.apps.example.com:8600
    KANEA_TOKEN: ${{ secrets.KANEA_TOKEN }}
    KANEA_CA_CERT: ${{ secrets.KANEA_CA_PEM }}
  run: kanea deploy shop/web "ghcr.io/acme/web@${{ steps.build.outputs.digest }}"
If the job fails before it deploys

unknown command "sh" - the image's entrypoint is kanea and GitLab passes the job script to the container as its command; use the map form above with entrypoint: [""]. KANEA_TOKEN is set but no endpoint - you set the token and not the URL. … has no credential without a token - the secret was not exported and arrived empty, which is why an empty variable counts as unset rather than as a token. … certificate from an unknown authority - set KANEA_CA_CERT. refusing to send a token … over plain HTTP - the endpoint is http:// and not loopback, which would put the token on the wire in clear text.

What stays on the node

Two commands refuse a remote endpoint, and both are about the machine rather than the platform:

  • kanea upgrade restarts this host's kanead and kanea-edge units and installs a binary over this host's. Reading a remote daemon's version and then restarting local services is the worst outcome available here, so it is refused by name. Upgrade over ssh, or on the node.
  • kanea mcp serves MCP over stdio, and its credential is the local socket. A remote agent already has the node's own /mcp endpoint with a bearer token, so a second spelling would add nothing.

kanea init, install, bundle and doctor never take an endpoint at all: they act on the host they run on. Everything else works remotely, including kanea exec (over wss://, with the same token) and kanea logs -f.

The alternative: let the node pull

A token in CI is not the only way. Kanea's own design answer is the other direction: the node watches a git repository, and CI commits the digest rather than calling out.

  1. Your pipeline builds and pushes the image, then commits the new digest into the spec repository the project's git block names.
  2. The push fires the webhook (POST /v1/webhooks/git/<project>, authenticated with GitLab's X-Gitlab-Token or GitHub's X-Hub-Signature-256 against the project's webhook_secret_ref).
  3. Kanea re-reads the repository over its own credential and applies what it finds. Nothing is deployed from the request, so a forged delivery cannot choose an image.

It needs no token in CI and leaves a git history of every deploy; it costs a spec repository and a commit per release. kanea project sync <project> forces the read immediately rather than waiting for the poll. Which one to prefer is a question of whether you would rather your deploys be a push or a record.

Troubleshooting

Where to look when something is wrong, in the order that usually finds it: the workload first, then the daemons underneath it, then the node. Everything in this section is read-only and safe to run on a live node.

Checking services

Start with what Kanea believes is true:

kanea ps -a                  # every alloc; including stopped services and pending slots
kanea status shop/web        # health, recent events, current vs desired counts
kanea describe shop/web      # the full picture: spec, routes, volumes, allocs, stats, events

Three columns in ps carry most of the signal. State: pending means the reconciler has not created the slot yet; usually an image still pulling, or a dependency that is not healthy. Health: a - means the service declares no health check, which is a different fact than failing one; a check-free service is never reported unhealthy, only running or not. Restarts: a climbing count is a crash loop; when it stops climbing the restart budget is exhausted and the alloc has been failed and left alone (see below for the way out). A node reboot moves no counter: allocs found dead at boot are recovered without spending the budget.

Then the processes underneath. A standard install runs four units:

systemctl status kanead kanea-edge kanea-containerd kanea-buildkit
  • kanead: the control plane. Workloads and north-south traffic both survive it being down; what stops is change: deploys, scaling, certificate renewal, the API and dashboard.
  • kanea-edge: the ingress proxy. Deliberately independent of kanead (no After= in either direction): it serves the last route snapshot it read from disk whether or not the control plane is up.
  • kanea-containerd: Kanea's own containerd, socket at /run/kanea/containerd.sock. Restarting it does not stop running containers (KillMode=process: shims outlive it), but nothing can be created or probed while it is down. Absent when the node adopted an existing daemon with --containerd external.
  • kanea-buildkit: the rootless build daemon. Only builds need it.

Finally the node itself: kanea doctor verifies dependencies and their pinned versions, the containerd socket, bpffs and the cgroup2 mount, slice placement and the effective memory floor, the build socket, disk headroom and clock sync, and it names known interference, like a firewall FORWARD-drop policy (docker, ufw) eating east-west traffic. Safe to run any time.

Viewing logs

Workload logs stream through the CLI:

kanea logs shop/web -f               # merged across allocs, follow
kanea logs shop/web --tail=200       # the last 200 lines first
kanea logs shop/web --alloc=<id>     # one alloc only (ids from kanea ps)

On disk they are one file per alloc under /var/log/kanea/allocs/ (<alloc-id>.log); a replaced alloc starts a fresh file under its new id. Drains are non-blocking with drop counters, so a burst of logging can be dropped but can never stall the workload's write(): if lines are missing under load, that is the drop counter doing its job, not a lost file.

Daemon logs go to stderr, which under systemd means the journal:

journalctl -u kanead -e              # control plane: reconciler, deploys, certificates, backups
journalctl -u kanea-edge -e          # ingress: TLS, routing, published ports
journalctl -u kanea-containerd -e    # runtime: image pulls, task create failures
journalctl -u kanea-buildkit -e      # the build daemon

journalctl -u kanead -f --since "15 min ago"

A deploy that goes wrong is usually legible in kanead's journal; a task that will not create at all (a missing shim, a device the node did not grant) often explains itself one level down in kanea-containerd's.

Build logs are their own stream: kanea build shop/web --follow live, kanea project builds shop for history, and one file per run under /var/lib/kanea/builds/.

Slices and resources

Everything Kanea runs sits in one of two cgroup slices, and that split is the resource-isolation story (architecture): kanea.slice holds the control plane (all four units above) with a kernel-guaranteed memory floor (MemoryMin, default 256 MiB; raise it with kanea init --reserve on a node that runs builds), and kanea-workloads.slice holds every alloc under a collective ceiling of total RAM minus that reserve.

systemctl status kanea.slice             # the control plane, with its live memory number
systemctl status kanea-workloads.slice   # every alloc as a child cgroup
systemd-cgls kanea-workloads.slice       # the tree, one cgroup per alloc
systemd-cgtop                            # live CPU/memory per slice

The ceiling is computed and applied by kanead at startup (it depends on how much memory the node has, which a unit file cannot know) so read it from the cgroup filesystem, which is the ground truth either way:

cat /sys/fs/cgroup/kanea.slice/memory.min             # the floor
cat /sys/fs/cgroup/kanea-workloads.slice/memory.max   # the ceiling
cat /sys/fs/cgroup/kanea-workloads.slice/memory.current

Per-alloc cgroups live one level down (/sys/fs/cgroup/kanea-workloads.slice/<alloc>/) with their own memory.max, memory.current and cpu.max. A declared resources limit is enforced exactly there; an omitted one reads as max: unbounded within the collective ceiling, by design, never a filled-in default.

If kanead was OOM-killed, the slices are the first suspect

The floor and the OOM score adjustments live in the unit files, not in the Go code: a Kanea started outside its units runs without the guarantee, and the first time the node is under memory pressure the kernel picks whatever is largest, which is usually kanead. kanea doctor checks slice placement and the effective floor for exactly this reason; journalctl -k | grep -i oom shows what the kernel actually chose.

Common situations

  • A service is crash-looping. kanea logs for why. Once the restart budget is exhausted the alloc is failed and left alone on purpose: kanea restart shop/web clears it, because the restart count belongs to the spec hash that spent it and a restart is a new one. So does deploying a fix. Only crashes the daemon watched count: after a power loss or reboot, every alloc found stopped comes back on its own with the budget untouched (v0.31.0; earlier versions charged one attempt per outage and could leave services failed at boot).
  • A deploy did nothing. A deploy is a spec-hash mismatch, not an invocation: re-running an unchanged spec is a no-op by design. kanea plan shows the diff the daemon would see: an empty one means the spec really is what is running.
  • The site is unreachable but the service is healthy. Work outward: kanea describe shows the routes the service declares, then systemctl status kanea-edge. The edge reads its world from two files (/run/kanea-edge/routes.json and its certificate bundle) so a route missing from that snapshot is a publishing problem in kanead's journal, and a route present in it is a proxying problem in the edge's.
  • Browsers show a certificate error. Deliberate fail-closed behaviour, not breakage: a domain whose certificate is not issued yet refuses the handshake, and a provided certificate that stops resolving serves plaintext; neither ever silently falls back to a weaker certificate. kanead's journal has the ACME or resolution failure.
  • Containers cannot reach each other. kanea doctor first: a foreign FORWARD-drop firewall policy (docker, ufw) is a finding it names. Then remember policy is deny-by-default: a cross-project call needs allow_from on the callee.
  • The bind stanza is ignored: the dashboard stays on localhost. An explicit --listen always beats the file, and an init run from before the stanza existed rendered --listen 127.0.0.1:8600 into the kanead unit. systemctl cat kanead | grep -- --listen confirms it; remove the flag (and --listen-cert/--listen-key if present) from ExecStart, then systemctl daemon-reload and restart. A re-run of kanea init will not put it back, with bind.api_addr declared it renders no listen flags at all.
  • kanead will not start on a fresh cloud node. If the journal says /etc/resolv.conf lists only loopback resolvers, the node is older than v0.23.2 and is hitting a fixed bug: systemd-resolved's 127.0.0.53 stub is the only nameserver a stock Debian or Ubuntu server has, and it was being discarded, leaving no upstreams and a refusal. sudo kanea upgrade is the fix - the stub is a perfectly good upstream, since kanead forwards from the host's own namespace. A dns stanza pins resolvers if you want to, but it is an override, never a requirement.
  • The server config seems to do nothing. Three things, in order. Did it load? journalctl -u kanead | grep 'server config' shows server config loaded with the path, or nothing at all if the file is absent. Is a stanza being ignored? The same grep shows carries stanzas this version does not read with the offending names - that is where a misspelled stanza lands, since only unknown attributes inside a read stanza are errors. Is a flag beating it? Look for is not consulted: an explicit flag on the unit always wins, and --config off disables the file entirely. Remember the file is read once, so an edit needs systemctl restart kanead.
  • Builds fail with exec: "buildctl": executable file not found in $PATH. The binary is installed - kanea doctor will confirm it is at its pinned version - but kanead cannot find it, because buildctl lives in Kanea's own bin dir and systemd's default PATH does not include it. Nodes whose units were written before this was fixed need the units regenerated: sudo kanea init (idempotent) then sudo systemctl daemon-reload && sudo systemctl restart kanead. kanea doctor names it under the buildkit check. Upgrading alone will not fix it: kanea upgrade deliberately never rewrites units.
  • The dashboard fails the TLS handshake, or the browser says the certificate is broken. Check the port first. The dashboard is kanead's own listener, so it answers on bind.api_addr's port (https://<name>:8600); the bare name goes to :443, which is kanea-edge, and on a node with no exposed services the edge holds no certificates at all, so the handshake dies with SSL_ERROR_INTERNAL_ERROR_ALERT rather than a 404. If the port is right, the other cause is a certificate that has not issued yet: the listener refuses the handshake rather than serving something weaker, and journalctl -u kanead | grep 'api listener certificate installed' says whether it arrived. See the worked example.
  • A service has no public name, or the name does not resolve. Three separate things have to be true, and they fail differently. kanea describe shows the domains the service actually claims: if it shows none, the node has no --base-domain - which kanea init never sets, so it is a drop-in on the unit, not a spec change. If it shows a name that does not resolve, the DNS record is yours to create; a wildcard covers every service at once. If it resolves but the certificate never issues, the record has to point at the node before issuance for HTTP-01, and --acme-email has to be set or nothing is requested at all.
  • A cross-project call resolves but times out. That is policy, not DNS. Names resolve for everyone because DNS is not the security boundary; the datapath is, and it denies between projects by default. The callee needs allow_from naming the caller. A timeout rather than a refusal is what a dropped packet looks like, which is why this reads as a network fault.
  • The autoscaler stopped scaling. Its circuit breaker trips on purpose after repeated failed actions, and says so in the journal, on the dashboard, and as kanea_circuit_breaker_open in /v1/metrics. The trip survives a daemon restart by design: restarting kanead is not a way around it; fixing the cause is. kanea scale still works meanwhile: the breaker pauses the automatic decisions, not the route they travel.