Documentation
Overview
Kanea runs containers on one machine and gives that machine the things a platform normally needs a cluster for: service discovery, load balancing, network policy, TLS, autoscaling, GitOps, backups. It is one static binary, and the concepts it asks you to learn number about six.
If you have not installed it yet, Installing is the next section: one command for the binary, one for the node. The rest of this page assumes you have a node running and want to know what you are looking at.
Concepts
Kanea borrows from Nomad and Kubernetes and keeps the smallest set of concepts that covers the work. If you know either system, the middle columns tell you what maps to what; if you are coming from a managed cloud service, the ECS and Container Apps table is below.
| Kanea | Nomad | Kubernetes | What it is |
|---|---|---|---|
| Node | Client + server agent | Node + control plane | One machine running kanea agent. In v1 there is exactly one. |
| Project | Namespace | Namespace | A named group of services, and the isolation boundary: network policy, secrets, containerd namespace and DNS all key off it. |
| Service | Job + group | Deployment + Service | A declarative long-running workload with count replicas. |
| Task | Task | Container | The container inside a service. Exactly one per service in v1: sidecars are a v1.1 concept. |
| Alloc | Allocation | Pod | One running instance of a service. count = 3 means three allocs. |
| Storage | CSI volume | PV / PVC | A named volume backend (local, host, nfs, smb or s3) that services mount by name. |
| Pipeline | - | Tekton / CI job | A build run producing an image, from a build block or a git push. |
The one that surprises people is project. It is not a label: it is a boundary
that four subsystems enforce independently. A service in project shop
cannot reach a service in analytics unless somebody wrote that down, it
cannot read analytics' secrets even by naming them, its containers live in
a separate containerd namespace, and its DNS names sit under a separate suffix.
Coming from ECS or Container Apps
The two managed services closest to Kanea in shape are Amazon ECS and Azure Container Apps. (On Azure, Container Apps is the right comparison: AKS is the Kubernetes column above, and Container Instances is a single container with no orchestration around it.) Both map onto Kanea closely enough to be worth a table - but read the ECS column carefully, because one of its terms means the opposite of what it means here:
| Kanea | Amazon ECS | Azure Container Apps |
|---|---|---|
| Node | Cluster, plus its container instances | Environment |
| Project | - (a cluster is the nearest thing) | Environment, again |
| Service | Service | Container app |
| Task | Container definition | Container |
| Alloc | Task | Replica |
| Spec file | Task definition | App YAML / ARM template |
| Deploy (spec-hash change) | New task definition revision | New revision |
| Storage | Volume: bind mount, EFS, FSx | Volume mount: Azure Files, ephemeral |
| Internal DNS | Service Connect / Cloud Map | Built-in environment DNS |
expose + the edge | ALB target group + listener rule | Ingress |
scaling | Application Auto Scaling | KEDA scale rules |
| Pipeline | CodePipeline + CodeBuild | ACR Tasks / GitHub Actions |
| Secrets | Secrets Manager / SSM, by ARN | Container app secrets / Key Vault |
In ECS a task is a running instance - the unit a service keeps
desiredCount of, and the direct equivalent of a Kubernetes pod. In
Kanea that is an alloc. A Kanea task is the container inside
a service, which ECS calls a container definition and puts inside a task
definition.
So "three tasks" means three replicas on ECS and is not a sentence you can say in Kanea, where a service has exactly one task and three allocs. It is the single most reliable way to misread this documentation coming from ECS.
Where the mapping breaks down
- Project has no ECS equivalent. ECS groups by cluster and otherwise relies on IAM and tags, so there is no boundary that carries network policy, secret scope and DNS at once the way a Kanea project does. On Container Apps the environment is the closest thing, which is why it appears twice in the table above: it is both the hosting boundary and the network boundary, where Kanea splits those into the node and the project.
- One node, not a region. Both services schedule across availability zones and survive a machine dying; Kanea v1 is a single machine, and if it stops, everything on it stops. That is the trade the whole design makes, and state replication and restore is the answer it offers instead of failover: an encrypted snapshot and change log in S3, restored onto a new box.
-
Nothing is billed per control plane, because there is no control plane to
rent.
kaneadis a process on your machine inside a 256 MiB reservation. The flip side is that patching, capacity and the node's own uptime are yours. -
Load balancing is in the datapath, not a resource you create. A service
gets a virtual IP the moment it exists, and east-west traffic is rewritten at
connect()by eBPF rather than crossing a proxy at all. Container Apps gives you ingress built in and is closest here; on ECS the equivalent is a load balancer, target group and listener rules to declare, wire up and pay for. North-south traffic goes through the edge, which is a process on the same box, not a product. -
The spec is one file for the whole project. An ECS deployment is typically
a task definition plus a service definition plus target groups plus listener
rules, assembled by Terraform or CDK; a Kanea project is one HCL file that
kanea planvalidates offline, covering the containers, routing, TLS, volumes, scaling and builds together.
Naming
Project and service names must be DNS-1123 labels: lowercase alphanumeric and
-, starting and ending alphanumeric, at most 63 characters. This is
checked when the spec is parsed, not when something fails later.
The reason is that names compose into DNS without an escaping step. A service
web in project shop is web.shop.kanea internally
and web.shop.<base_domain> publicly. A name that needed encoding to
become a hostname would be a name that means two different things in two places.
Every project and service takes a description: free text, up to 512
characters, shown in the dashboard. That is where the human-readable detail belongs;
the name stays a label.
Lifecycle
Nothing in Kanea is applied directly to the runtime. A spec becomes desired state in the Store, and a reconciler converges the world toward it: continuously, not once.
Job spec (HCL) ──parse/validate──▶ Desired state (Store)
│
Reconciler loop
│
┌───────────────┼────────────────┐
▼ ▼ ▼
containerd eBPF datapath Edge proxy
(tasks/images) (policy/LB) (routes/TLS)
│ │ │
└───────────────┴────────────────┘
▼
Actual state / events / metrics
Consequences worth knowing up front:
- Drift is repaired. Delete a container by hand and it comes back. The reconciler is comparing, not remembering what it did.
- Restart policies are
always(default),on-failurewith backoff, andnever. - Updates are rolling by default, health-gated, bounded by
max_parallel. - Dependencies start first. A
depends_onedge (or any${service.…}reference, which creates one implicitly) means the dependent does not start until its dependency is healthy. If the dependency degrades later, dependents keep running: no cascading stops, just events. - Storms are capped. Per-service restart rate limits plus a node-wide circuit breaker that pauses rollouts and scale actions when failure rates spike. A trip emits an event and a notification.
What counts as a deploy
A deploy is a spec-hash mismatch. The reconciler hashes the parts of a service
that are baked into a container at creation time; if an alloc's recorded hash differs
from the current one, that alloc is replaced under the update policy. If it matches,
nothing happens, however many times you run kanea run.
This is why kanea restart is not a separate path into the runtime: it
bumps a generation counter that participates in the hash, and the ordinary rolling
update does the rest.
max_parallel bounds allocs that are down, not replacements in
flight. Anything already unavailable spends the budget first, so a deploy that
starts going wrong stops, instead of walking through every replica.
min_healthy applies only to allocs the current deploy has already
replaced, and health means a probe said so. A service with no
health_check block never reports healthy for any alloc, which is fine,
and is why the update logic asks whether a check is configured before it asks whether
it passed.
Secrets
Secrets are never written in a spec. They are referenced as
secret:<path> and resolved when an alloc starts:
env = {
DATABASE_URL = "secret:shop/database-url"
}
The variable then carries the path of a file in a per-alloc tmpfs mount
(/run/kanea/secrets/<alloc>/shop/database-url, read-only,
owned by the service's user) - the value never sits in an environment
block. Software that can only read environment variables opts into the
weaker form, secret-env:shop/database-url, which inlines the
value (visible via /proc/<pid>/environ); the record keeps
the reference either way, and a rotated secret takes effect at the next
replacement.
Two rules make that reference safe rather than merely tidy:
- References are project-scoped at validation time. A service in
shopmay namesecret:shop/…orsecret:shared/…and nothing else. A spec reaching for another project's secret fails to parse: it does not fail at runtime, where the failure would be a log line nobody reads. - The default injection is a tmpfs file, at
/run/kanea/secrets/<alloc>/<name>. Environment variables work and are documented as the weaker option, because they are visible in/proc/<pid>/environ, in runtime inspect APIs, and to every child process.
The API and the MCP server are write-only for secrets. There is no get,
not for an operator, not for an agent, at any permission tier.
Your first service
The smallest useful thing needs no file at all:
kanea run --image nginx:1.27-alpine --name web --project demo kanea ps -p demo kanea logs -f demo/web
When you want it written down, the same deployment as a spec is three lines plus a
wrapper, and from there every block in the
job spec reference is additive. Run
kanea plan first; it prints the create/change/destroy diff, and it is
where every validation rule fires.
Installing
Two commands: one puts the binary on the node, one turns the node into a platform.
The first only ever writes /usr/local/bin/kanea; everything else
(the runtime, the keys, the units, the accounts) belongs to
kanea init, which you can read before you run.
| Requirement | What it must be |
|---|---|
| Architectures | linux/amd64, linux/arm64. Anything else is refused before a byte is downloaded. |
| Kernel | ≥ 5.10, cgroups v2 unified hierarchy. |
| Init system | systemd. |
| Clock | NTP-synchronised (certificates and audit timestamps depend on it). |
That is the whole list. containerd, runc, rootless buildkitd
and the wasmtime shim are not prerequisites: kanea init installs
them at versions pinned by SHA-256 inside the binary. The network layer needs no
component at all, because the eBPF datapath is compiled into kanea.
The install script
curl -fsSL https://m18h.github.io/kanea/install.sh | sudo bash
It resolves the latest release, downloads the archive for this architecture, verifies it, and installs one binary. Then it stops. It generates no keys, starts no services, writes no units and installs no runtime, so running it on a node that is already serving traffic changes nothing until you act on what it prints.
There are no arguments. The three knobs are environment variables, and
sudo does not forward them by default - use
sudo VAR=value bash or sudo -E:
| Variable | Default | Meaning |
|---|---|---|
KANEA_VERSION | latest | The release tag to install, e.g. vX.Y.Z. latest is resolved from the redirect GitHub already serves, so no jq and no API rate limit. |
KANEA_PREFIX | /usr/local/bin | Directory the binary is installed into, mode 0755. |
KANEA_REPO | m18h/kanea | The owner/repo releases come from, so a fork installs from itself. |
KANEA_REQUIRE_SIGNATURE | 0 | Set to 1 to make the cosign verification mandatory: a missing cosign binary or an unsigned release becomes fatal instead of a note. The script never downloads cosign for itself. |
curl -fsSL https://m18h.github.io/kanea/install.sh | sudo KANEA_VERSION=vX.Y.Z bash
It needs curl, tar and install on the node,
and it must run under bash, not sh. Four things are
refused up front rather than discovered after a download: a non-Linux kernel, an
architecture outside amd64 and arm64, an unreachable release endpoint, and a
repository with no published release (which would otherwise resolve to a
plausible-looking archive name and a 404 nobody can read).
What is verified
The checksum is mandatory and there is no flag to skip it. The script
downloads checksums.txt, takes the line for the archive it fetched, and
runs it through sha256sum -c. A mismatch stops with
checksum mismatch: do not run this binary.
The Sigstore signature is checked when cosign is on
PATH: the keyless signature over checksums.txt, bound
to this repository's release workflow through the GitHub OIDC issuer. The
asymmetry is deliberate and worth being precise about:
- cosign present, signature valid - prints
Signature verified. - cosign present, signature invalid - fatal. The install stops.
- cosign present, no signature published - a note, then continues on the checksum alone.
- cosign absent - prints
cosign not found; checksum verified but signature not checkedand continues. Installing cosign first is the stronger path; verifying by hand is the same check written out.
KANEA_REQUIRE_SIGNATURE=1 turns the last two endings into refusals: a
node that cares gets to insist on the signature, and kanea upgrade
--require-signature is the same posture for every upgrade after this one.
The script creates the kanea-edge system account if it is missing
(useradd --system --no-create-home --shell /usr/sbin/nologin). The
ingress process runs as that user rather than root, and a failure here is a
warning, never fatal. Nothing else on the node is touched.
Setting up the node
sudo kanea init is the rest of the way: preflight checks, the host
components, the master-key ceremony, the systemd units, the first admin account,
and a summary of what it built. Run it once, as root. It is safe to re-run -
see what a re-run does below.
$ sudo kanea init
kanea init: <version>
API/dashboard listen address [127.0.0.1:8600] ("none" for socket-only): 192.168.1.10:8600
Listener certificate already present at /etc/kanea/api.crt; leaving it alone.
Checking this node:
ok platform linux/amd64
ok cgroups v2 unified hierarchy with cpu, memory and pids
ok kernel 6.12.43-amd64
ok clock synchronised (systemd-timesyncd)
ok systemd running
ok data directory /var/lib/kanea
ok edge user kanea-edge
Installing the host components:
kanea install: <version>
Source: upstream
Prefix: /usr/local/lib/kanea
up-to-date containerd 2.3.3 container runtime daemon
up-to-date runc 1.5.1 OCI runtime
up-to-date wasmtime-shim 0.6.1 wasmtime containerd shim; functions
Wrote /etc/kanea/containerd/config.toml
Wrote /etc/systemd/system/kanea-containerd.service
Starting containerd to pull the image components…
up-to-date buildkit v0.32.0-rootless rootless build daemon (the only build driver)
Wrote /etc/systemd/system/kanea-buildkit.service
Created /var/lib/kanea, /var/log/kanea/allocs and /etc/kanea
Master key already present at /var/lib/kanea/master.key; leaving it alone.
Wrote /etc/systemd/system/kanea.slice
Wrote /etc/systemd/system/kanea-workloads.slice
Wrote /etc/systemd/system/kanead.service
Wrote /etc/systemd/system/kanea-edge.service
Starting kanead:
waiting for the control plane…
First admin username: <name>
password for <name>:
again:
Created admin account "<name>".
Restarting kanead so the settled listener takes effect…
───────────────────────────────────────────────────────────────
Kanea is running.
Dashboard https://192.168.1.10:8600 (log in as "<name>")
Internal DNS 10.244.0.1:53 (allocs resolve <service>.<project> here)
Addressing
node CIDR 10.244.0.0/24 this node's allocs
cluster CIDR 10.244.0.0/16 routed, masqueraded as internal
service CIDR 10.201.0.0/16 service VIPs
Deploy something: kanea run <spec.hcl>
CLI without sudo: sudo usermod -aG kanea <user> # root-equivalent; log in again
Configure a backup destination before you need one; see docs/DR_RUNBOOK.md.
Versions and the kernel string are placeholders; the shape is what a run looks
like. This one is a re-run on a node that already had a key and a listener
certificate, which is why two lines say "leaving it alone" and the components read
up-to-date. A first run mints both and installs each component.
What it asks
-
The listen address, defaulting to
127.0.0.1:8600. Answernoneto keep the API on its unix socket only. This is asked only when--listenwas not passed and stdin is a terminal, so a scripted install never has a prompt eat one of its lines. Abindstanza in the server config replaces the question entirely. - The master key, typed back. It is printed once and you must retype it exactly; a mismatch writes nothing and stops. There is no flag to accept it unseen.
-
The first admin's username, unless
--admin-userwas given. - That account's password, twice.
Without it, every backup this node ever makes is unreadable and every stored secret is unrecoverable. Have somewhere to record it before you start. Every prompt in a run shares one reader, so a piped install works exactly as a typed one does.
TLS for the dashboard
Beyond loopback the listener is served over TLS, and init provisions it rather
than refusing: with no explicit --listen-cert/--listen-key
pair it mints a ten-year self-signed pair at /etc/kanea/api.crt and
api.key (key mode 0600, with a real IP SAN when the address is an IP)
and points the unit at it. It is minted once and never re-minted, because
re-minting would flip the fingerprint an operator has already accepted.
Bring your own with the two flags, or hand the whole listener to the
bind stanza, which also gets you
acme and the node CA. One case is still refused: an unspecified host
(0.0.0.0, :8600), because a certificate's SAN needs
something to name - bind a specific address, or set
bind.api_domain.
What a re-run does
| Kept | Rewritten |
|---|---|
| The master key, always. An existing one is never regenerated. | The four systemd units, every time, at mode 0644. |
The api.crt/api.key pair. Half a pair is a hard error, never a guess. | The containerd config and component units. |
| The first admin. If any account exists the step is skipped, never re-prompted. | |
/etc/kanea/kanea.hcl. Init never writes it, on a first run or a re-run. |
That is exactly why a release whose notes mention changed units is a re-run of
sudo kanea init followed by systemctl daemon-reload:
it is the supported way to regenerate them, and it costs you neither the key nor
the accounts. kanead is only restarted when the settled listener actually differs
from the running one, so an unchanged re-run restarts nothing.
Useful flags
The full list is in the CLI reference; these are the ones that come up:
| Flag | Default | Why |
|---|---|---|
--listen | 127.0.0.1:8600 | Skips the prompt. none keeps the API socket-only. |
--admin-user | prompt | With a piped password, makes the whole run scriptable. |
--containerd external | install our own | Adopt the containerd already on the node instead of installing one. |
--bundle | - | Install the host components from an offline bundle. |
--reserve | 256M | The control plane's memory floor. A node that runs builds wants 512M: buildkitd alone holds ~157 MiB. |
--node-cidr | 10.244.0.0/24 | This node's container subnet. Its .1 is also the internal DNS address. |
--no-start | false | Write everything and stop, without starting kanead or creating an account. |
--skip-checks | false | Run the ceremony without the preflight. The checks gate, so this is how you override one you disagree with. |
Host components
containerd, runc, rootless buildkitd and the wasmtime
shim are installed by Kanea, pinned by version and SHA-256 in a manifest compiled
into the binary. Hashes are never fetched, so moving a component is a code
change with a review behind it, and the same manifest is the version matrix
kanea doctor checks the node against.
kanea install --list # the pinned versions, straight from the binary sudo kanea install --dry-run # resolve and verify every artefact, write nothing sudo kanea install # normally you never run this; kanea init does
Nothing at a distribution's paths is touched. Kanea's containerd lives under
/usr/local/lib/kanea on its own socket at
/run/kanea/containerd.sock, which is why a node that ran Docker
yesterday runs it tomorrow.
| Flag | Meaning |
|---|---|
--list | Print the pinned matrix and exit. Needs no network and no root. |
--dry-run | Download and verify everything, write nothing. The honest pre-flight for an upgrade. |
--only | Comma-separated component names. They are reordered into manifest order, because install order is a dependency: containerd and runc first, containerd starts, and only then is the buildkit image pulled through it. |
--force | Reinstall components already at the pinned version. |
--bundle | Install from an offline bundle. This turns network fetching off entirely; it is not a preference. |
--containerd external | Adopt an existing containerd: drops containerd from the set and writes a slice drop-in instead. |
--arch | Target architecture. Meaningful only with --dry-run; installing for a foreign architecture is refused. |
--prefix, --conf-dir, --data-dir, --run-dir, --unit-dir | Where binaries, configuration, state, sockets and units go. |
--skip-units | Binaries only. Also skips the image components, since containerd is never started. |
Installing by hand
Every release publishes per-architecture archives, an SPDX SBOM beside each one,
a kanea_<version>_source.spdx.json for the build's own graph
(where the embedded dashboard's npm dependencies are listed), a
checksums.txt, and a keyless cosign signature over that file. The
SBOMs are listed inside the checksums, so the one signature covers them too:
VERSION=vX.Y.Z; ARCH=amd64 # the tag you want, e.g. from the releases page
BASE=https://github.com/m18h/kanea/releases/download/$VERSION
curl -fLO $BASE/kanea_${VERSION#v}_linux_$ARCH.tar.gz
curl -fL -O $BASE/checksums.txt -O $BASE/checksums.txt.sig -O $BASE/checksums.txt.pem
cosign verify-blob \
--certificate checksums.txt.pem \
--signature checksums.txt.sig \
--certificate-identity-regexp "https://github.com/m18h/kanea/" \
--certificate-oidc-issuer "https://token.actions.githubusercontent.com" \
checksums.txt
sha256sum --ignore-missing -c checksums.txt
tar xzf kanea_${VERSION#v}_linux_$ARCH.tar.gz
sudo install -m 0755 kanea /usr/local/bin/kanea
There is no long-lived signing key to guard. The signature is bound by Sigstore to
the release workflow in this repository, and the proof is in a public transparency
log. Then continue at sudo kanea init.
Homebrew
brew tap m18h/kanea brew trust m18h/kanea # brew ≥ 6 refuses formulae from untrusted third-party taps brew install kanea
Homebrew is a CLI channel, never a node channel. On macOS you get the
authoring half, where kanea plan validates job specs with
file-and-line diagnostics and needs no daemon. On Linux the formula installs the
same full binary, but a node belongs to the install script above:
root-owned at /usr/local/bin, where kanea upgrade owns
the swap. A brew-owned binary upgrades with brew upgrade kanea, then
sudo kanea upgrade --no-fetch for the restart-and-migrate half.
The container image
docker run --rm ghcr.io/m18h/kanea:latest version # the shape a pipeline uses: the spec on a mount, the node over the network docker run --rm -v "$PWD:/workspace:ro" \ -e KANEA_URL -e KANEA_TOKEN \ ghcr.io/m18h/kanea:vX.Y.Z run shop.hcl # pin it; that is what tags are for
linux/amd64 and linux/arm64 in one manifest list, tagged
vX.Y.Z and latest; latest moves only for a
plain vX.Y.Z, so a prerelease never becomes what everyone pulls by
default. Like Homebrew it is a CLI channel, never a node channel: it is for
the pipeline that deploys to a node, so every verb that acts on a host, its
systemd and its own binary - agent, edge,
init, install, doctor,
upgrade - is not what it carries.
The image contains the released binary, not a rebuild of it. The release
workflow builds the image after it signs, from the same
linux_amd64/linux_arm64 archives you can download, after
verifying checksums.txt under its signature and each archive against
it - the two checks the by-hand install
performs, in that order. So the manifest of hashes you can verify yourself
describes what is inside the image, with no second build to reconcile.
The digest is signed with the same keyless identity as the checksums, and carries an SPDX attestation beside the signature:
cosign verify ghcr.io/m18h/kanea:vX.Y.Z \ --certificate-identity-regexp 'https://github.com/m18h/kanea/' \ --certificate-oidc-issuer 'https://token.actions.githubusercontent.com'
It runs as an unprivileged user (uid 65532) with /workspace as
its working directory, so mount your spec there. And if your network terminates
TLS with its own CA, either drop the root into
/usr/local/share/ca-certificates and run
update-ca-certificates in a derived image, or pass the node's own CA
with KANEA_CA_CERT - which replaces the system pool
rather than adding to it, so it is the node's CA or the public ones, never both.
Upgrading
sudo kanea upgrade alone is the whole upgrade: it downloads and
verifies the release itself before restarting anything. Re-running the install
script works too, and is what you do when the binary is what you want replaced
first; it changes the binary only, and says so:
$ curl -fsSL https://m18h.github.io/kanea/install.sh | sudo bash
Upgrading kanea <installed> -> <latest> (linux/amd64)
Verifying the checksum
cosign not found; checksum verified but signature not checked
Installed /usr/local/bin/kanea
The running daemons are still on the old binary. Next:
sudo kanea upgrade
It backs up, drains and restarts kanea-edge, then restarts kanead, which
runs any state migrations. Running workloads are untouched throughout.
The sequence kanea upgrade performs, in order: resolve and verify the
release (sha256 always, cosign when present, fatal if present and failing),
install it atomically over its own path, take a pre-upgrade backup, restart
kanea-edge, restart kanead, run state migrations, and wait
for health. The edge goes first because it drains in seconds and must be current
before kanead publishes a new-format route snapshot; kanead goes last because it
runs the migration once everything else has settled.
Running allocs are not touched at any point. Already being at the target version means nothing to download, so running it twice is safe by construction.
An admin can run the same flow from the dashboard's Updates page, reached
from a pinned sidebar item whose badge marks a newer published release or a
pending OS reboot. The page checks for updates and installs one
(GET and
POST /v1/upgrade, admin-only, audited). The daemon performs the
sequence above on itself - pre-upgrade backup, fetch and verify, atomic
install - then answers the request, restarts kanea-edge, and exits
into systemd's Restart=always; the page reloads once the new
version answers health, which is also what swaps in the new embedded dashboard.
The check never runs unprompted: the node asks GitHub only while an admin has
the dashboard open, and caches the answer for an hour, so a node never phones
home on its own. There is no downgrade over the API
(--allow-downgrade stays CLI-on-the-node), and outside systemd
nothing restarts: the binary is installed and the response says
restart_required. KANEA_REQUIRE_SIGNATURE=1 on the
unit gives this path the --require-signature posture.
| Flag | Meaning |
|---|---|
--check | Report the running, installed and latest versions, and stop. |
--version vX.Y.Z | Go to a specific release instead of the latest. |
--no-fetch | Download nothing; restart onto whatever binary is already installed. This is the half a package manager leaves for you. |
--require-signature | Refuse the release unless its cosign signature verifies: a missing cosign binary, or a release published without one, becomes fatal instead of a note. Never downloads cosign itself - a verifier fetched over the same channel would be a second trust root vouching for the first. |
--allow-downgrade | Permit a target older than the running daemon. Refused otherwise, because the Store's schema is forward-only. |
--skip-backup | Skip the pre-upgrade backup. The schema migration still takes its own local copy. |
--dry-run | Print what would run and stop. |
That is deliberate: your units may carry local edits. When release notes say the
units changed, re-run sudo kanea init (idempotent - key,
accounts and settings are kept) and then
sudo systemctl daemon-reload. Host components are pinned by the new
binary, so sudo kanea install --dry-run shows whether any need to
move and sudo kanea doctor confirms the node agrees with the matrix.
Air-gapped nodes
A node with no egress is a supported installation, not a workaround. Build a bundle where there is a network, carry it across, install from it:
kanea bundle create --arch amd64 -o kanea-bundle.tar.gz # connected machine sudo kanea init --bundle kanea-bundle.tar.gz # air-gapped node
The bundle carries no hashes of its own. Its contents are verified against the ones compiled into the installing node's binary; a bundle that supplied its own would be a bundle that authenticates itself. For image components that means a digest comparison rather than a name lookup, because containerd names image records from the archive's own annotation verbatim.
Releases publish one bundle per architecture, covered by the same signed
checksums.txt. This covers Kanea's own components; your workload
images still come from a registry the node can reach. On an air-gapped node
kanea doctor --offline skips the single network probe, and
kanea upgrade refuses to fetch with a pointer to this flow rather
than hanging.
Architecture
One binary produces two long-running processes and a CLI. The datapath that networks
your services is not one of them: it is a set of eBPF programs compiled into the binary
and loaded into the kernel, not a daemon Kanea drives. Everything else on the node
(containerd, buildkitd) is software Kanea drives rather than
software it contains, but it is no longer software you have to find: kanea
init installs each of them at a version pinned by SHA-256 in the binary, under
Kanea's own prefix and on Kanea's own sockets, so nothing already on the node changes.
The shape
┌──────────────── kanead (control plane) ────────────────┐
│ │
Browser ──HTTPS──▶ │ ┌────────────┐ ┌───────────┐ ┌──────────┐ ┌─────────┐ │
CLI ──────HTTPS──▶ │ │ API server │ │ Dashboard │ │Reconciler│ │Autoscale│ │
Webhooks ────────▶ │ │ REST + WS │ │ (embedded)│ │ │ │ (eBPF) │ │
│ └─────┬──────┘ └───────────┘ └────┬─────┘ └────┬────┘ │
│ └─────────────┬─────────────┴────────────┘ │
│ ┌────────┴─────────┐ │
│ │ Store (bbolt) │ │
│ └────────┬─────────┘ │
│ ┌─────────┬───────┼────────┬──────────┐ │
│ │ Runtime │Network│ GitOps │ Notifier │ │
│ │containerd│ eBPF │BuildKit│ │ │
│ └────┬────┴───┬───┴────┬───┴──────────┘ │
└────────┼────────┼────────┼───────────────┬─────────────┘
│ │ │ ▼
│ │ │ State replicator ──▶ S3
▼ ▼ ▼
┌── kanea-edge (separate process; reads a projection, never the Store) ──┐
│ L7 routing · TLS termination · middleware · request metrics │ ◀── :80/:443
└────────────────────────────────────────────────────────────────────────┘
External: containerd · buildkitd · Linux kernel (eBPF, cgroups v2, netfilter, bpffs)
Two processes, one binary
kanead is the control plane: API, dashboard, reconciler, autoscaler,
GitOps, notifications, backups. kanea-edge is the ingress proxy, and it
is a separate systemd unit running as a separate unprivileged user.
That split is the single most consequential design decision in the system, and it buys two things:
- Restarting, upgrading or crashing the control plane does not interrupt public
traffic. The edge unit deliberately has no
After=kanead.service. - The process that terminates untrusted public traffic cannot mutate the platform. It has no Store access and no write path at all.
bbolt takes a lock on the whole database file, so a second process opening it
(even read-only) would block until kanead exits. So the edge does not
open it. kanead projects what the edge needs into
/run/kanea-edge/: routes.json (0644, host → service
frontend; nothing secret, the domains are in public DNS) and
certs.json (0640, private keys). Two files, two permission sets,
so neither has to compromise for the other.
Both are written temp-then-rename(2), so a half-written file is never
observable. The projection carries the Store index it was built from, which is how a
stale snapshot can be recognised rather than merely suspected.
A missing or stale snapshot is not an outage: the edge keeps serving the last table
it loaded for as long as kanead is away, and starts with an empty table
rather than refusing to start. "The control plane is down" must never become "the
site is down".
The Store
One embedded bbolt database holds every mutation, behind a Store interface
with monotonic indexes. Buckets: projects, services, allocs, events, certs, secrets,
pipelines, audit, kv.
- Single writer. All mutations serialise. Reads are bounded and paginated, because a long read transaction blocks the writer.
- Raft-shaped on purpose. The interface and its index semantics are what a Raft FSM would need, so a clustered implementation can replace it without touching call sites.
- Metrics and logs never touch it. Time series live in a bounded in-memory ring; logs go to file pipelines with non-blocking drains. This is a hard constraint, not a performance preference: a metrics write that contends with the reconciler's writer would make observability and convergence share a failure mode.
- Migrations are explicit. The Store does not migrate itself at open. A
migration rewrites state in place, and the copy that makes a bad migration survivable
needs the database open and the migration not yet started: exactly one window,
between
OpenandMigrate.
Runtime driver: containerd
- Driven over its socket with the official Go client. One containerd namespace per
project (
kanea-<project>), which makes image and container isolation free rather than enforced. - Responsibilities: image pull (credentials from the secrets store, digest pinning supported), task lifecycle, per-alloc netns setup, cgroup metrics sampling, and stdout/stderr capture.
- Hardening is not opt-in. Every alloc starts from a baseline capability set;
the uid-switching grants PUID-style images need to chown a volume and drop to their
configured user, and nothing more (
CAP_NET_RAWis deliberately excluded, and binding :80 needs no capability at all - the alloc netns has no privileged-port floor); plusno-new-privileges, the default seccomp profile, and its own PID and IPC namespaces.capabilities = ["none"]drops to nothing; a capability a workload genuinely needs beyond the baseline is named in the spec and drawn from a permitted set; the privilege-equivalent ones are rejected when the spec is parsed. No job spec can lift any of it on its own: there is noprivilegedfield. The one way past these defaults is a host device or socket the operator granted on the node (R17, R18), which a spec requests by name and can never define. - The strong posture has a name.
hardening = "restricted"on a service requires a non-zerouser, runs the task with no capabilities at all, and refuses any grant declared beside it; a service that owns a volume chowns it in a root init step first.task.read_only_rootfs = truemounts the image read-only on top of that. Neither is the default - the baseline exists so stock images start - butkanea planwarns on the weak shapes it replaces: an effective uid of 0, a writable root filesystem, and an image following a moving tag without opting into auto-update. - Disk hygiene is part of the driver. Image GC, build-cache caps across both content stores (containerd's and the rootless buildkitd user's) per-service log caps, and watermark alerts at 80% and 90%. One disk holds images, logs, state and volumes; pressure must never surprise the control plane.
Network driver: eBPF datapath
The datapath is Kanea's own: three small eBPF programs, a handful of pinned maps and
plain netlink plumbing, all loaded and written by kanead from one object
compiled into the binary. There is no network agent, no kvstore and no CNI. The loader
is the standalone github.com/cilium/ebpf library: importing
github.com/cilium/cilium would pull the Kubernetes client graph, which is
the one dependency the project does not have. The programs are compiled ahead of time
and committed, so go build needs no clang and the node needs no BTF: they
read only UAPI context types, so there is no CO-RE and no vmlinux.h.
IP is identity
kanead allocates every alloc address from the node CIDR and every service
VIP from the service CIDR, durably in the Store. Because the platform that hands out
addresses is the same one that writes the kernel's maps, the identity map (alloc
IP → {project, service}) is written by the allocator itself. There is no
identity-allocation protocol, no label race and no settle window. The numeric project
and service ids behind the maps are Store-allocated, monotonic and never reused,
which is what keeps a pinned map meaningful across a kanead restart.
Attach is deny-closed by construction
Per alloc, in order: netns → identity map write → veth created with the host side
down → policy programs attached at tc → addresses and static neighbours → link up → the
host /32 route last. The first moment a packet can reach or leave the
alloc, policy is already enforcing and identity is already resolved; a skipped step
fails closed, because an identity miss is a drop. The deny-by-default guarantee is
structural, not temporal: there is no unlabelled window to hold shut with
retries, the way the previous design held its reserved:init state closed.
Attach has no wait loop and completes in milliseconds.
Load balancing is connect-time
One cgroup connect4 program at the root cgroup, held by a pinned
bpf_link, rewrites VIP:port → backend at
connect(2); for host processes and containers alike, which is the one
property the edge depends on: it dials a VIP with a plain dialer. There is no per-packet
NAT and no conntrack entry per flow, and an established connection never consults a
map: kanead can restart, or recreate every map, without touching live
traffic. Backends update by generation flip: the new set is written under the
next generation and one atomic map update commits it, so a concurrent
connect() sees a complete old set or a complete new one, never a torn one.
A VIP with no backends refuses at connect() rather than black-holing into a
timeout. Service ports are TCP-only in v1, refused at plan otherwise.
Policy is SYN-gated map entries
The tc program on each alloc's host-side veth admits a source carrying the host identity
(the edge's upstream dials, kanead's DNS replies and probes), a source in
the same project, or a source named by an allow_from edge (R14), and drops
every cluster-internal source the identity map does not know. A source outside
the cluster CIDR carries no identity by construction: it is the internet answering a
connection the alloc opened, un-NATed by conntrack on the way back in, and passes
(v1.65); nothing unsolicited arrives that way, because the pod CIDR is unroutable from
off-node and published ports terminate at the edge. The mirror-image rule guards the
other direction: a packet leaving an alloc must carry a cluster source, so a
forged external address cannot ride that pass. There is no policy
file, no selector language and no translation step that could make a rule silently match
nothing; policy is map entries the datapath enforces directly, and rules only ever
union, so allow_from can never weaken the project default-deny.
Enforcement is per connection attempt: TCP that is not a SYN passes, which is what lets cross-project replies flow without a conntrack. That is deliberately weaker than stateful tracking: an in-node ACK probe traverses the filter and is stopped only by the receiving stack's RST, and it is stated in the threat model rather than hidden. The upgrade to an LRU conntrack map is additive.
A second, small egress program is load-bearing for A10: it drops the cloud-metadata
range 169.254.0.0/16 in the kernel with a per-alloc drop counter (not
asserted in a policy file) along with any service-CIDR destination that escaped
connect-time rewrite, and it counts per-endpoint traffic. Masquerade for routed pod
traffic is one nftables rule in an owned kanea table, and since v1.65 the
rule and net.ipv4.ip_forward are re-asserted every thirty seconds:
a firewall reload that flushes the ruleset costs seconds, not a daemon restart. A
drop policy another tool installs can still eat that traffic, on the forward
hook and on the input hook alike - the second is how an alloc reaches the
internal resolver, and a default-deny ufw eats it while the host's own
dig keeps working, since every manager accepts lo
unconditionally. kanea doctor detects and names both, passing once the
ruleset carries an accept for the cluster CIDR, and kanea firewall
prints the rules; a missing kanea table and forwarding turned off are
findings of their own.
Internal DNS
Allocs resolve through kanead's own resolver, bound to the datapath's host
anchor (the node CIDR's .1) and never a wildcard, on UDP and TCP
alike. It answers <service>.<project>.kanea names
authoritatively and forwards everything else upstream. The upstream list is, in order of precedence: the
--dns-upstream flag; the server config's dns stanza
(v1.66); the host's own /etc/resolv.conf, read once at startup, taken
exactly as listed. That includes systemd-resolved's 127.0.0.53 stub,
which is the only nameserver a stock Debian or Ubuntu server has:
kanead forwards from the host's own namespace, so the stub is as
reachable for it as for any other process, and a workload inherits the node's
cache and DNSSEC posture with it (an alloc is never handed these addresses; its
resolv.conf names Kanea's resolver). The stanza is how you pin resolvers on a node
whose resolv.conf is DHCP's to rewrite:
# /etc/kanea/kanea.hcl
dns {
upstreams = ["1.1.1.1", "10.0.0.53:5353"] # a bare address gets :53
}
Entries are validated when the file parses; an empty list is refused by name; a
stanza that meant "no upstreams" would silently turn external resolution into
SERVFAIL. An explicit --dns-upstream wins over the stanza and the daemon
says so in its log, the same precedence every half of the server config follows. The
stanza's own reference, and how to apply an edit, is in
Node configuration.
An alloc's query to the resolver is a new inbound connection to the host on a
veth, so it crosses the input hook like any other arrival and a default-deny
ufw or firewalld eats it. Nothing about it looks like a firewall problem: the
host's own dig @10.244.0.1 keeps answering, because every manager
accepts lo unconditionally, and the resolver's logs show a healthy
server nobody is reaching. kanea doctor names it and
kanea firewall prints the rules.
TCP is served beside UDP (v1.86) because three things set the truncation bit - an internal answer over 512 bytes, an upstream reply that filled the read buffer, and an upstream reply that already carried TC - and each one tells a client to retry over TCP. Serving only UDP made every one of those a dead end. The TCP half is bounded the way the UDP half is: a hard concurrent-connection cap whose refusals are counted, deadlines on every connection, and a TCP query forwarded over TCP, so the retry can carry what the datagram could not.
Datapath state is derived state. Programs, maps and the cgroup link are pinned
under /sys/fs/bpf/kanea with a schema stamp; a kanead restart
leaves the dataplane untouched, and a stamp mismatch recreates and repopulates the maps
inside the first reconcile pass: safe precisely because established flows bypass them.
Nothing under the pin root is ever backed up.
The edge
A Go reverse proxy: host-based L7 routing to service frontends, TLS termination with Let's Encrypt certificates, WebSocket and gRPC support, HTTP→HTTPS redirects and security headers.
The middleware chain runs in a fixed order, per service, from the spec's expose block:
Host match → IP allow/deny → rate limit → header transforms → upstream proxy
All of it is validated at kanea plan time and fails closed. An ingress
control that silently does nothing at runtime is worse than one that is absent.
Hardening is mandatory rather than tunable: read/header/idle timeouts and header size
caps against slowloris, per-route upstream timeouts, bounded connection pools,
client-supplied X-Forwarded-* stripped, and an unknown Host
answered with 404, which is also the DNS-rebinding defence for the co-hosted API.
Because it already sits in the request path, the edge is the primary source of L7 metrics for exposed services: requests per second and latency percentiles at no extra data-plane cost. The datapath's own map counters cover east-west.
ACME runs in kanead, never in the edge: obtaining a certificate means
writing one, and the edge does not write.
Resource isolation
The control plane must survive anything a workload does. Enforcement is cgroups v2, arranged as two sibling slices:
/sys/fs/cgroup
├── kanea.slice # kanead, kanea-edge (+ containerd and buildkitd, via their own units)
│ memory.min = system_reserve_memory # kernel-protected floor (default 256 MiB; build nodes raise it)
│ memory.swap.max = 0 # the floor is RAM, not swap
│ cpu.weight = 10000
│ OOMScoreAdjust = -900 # the global OOM killer picks workloads first
└── kanea-workloads.slice # every alloc lives under this one parent
memory.max = total RAM − system_reserve_memory
memory.swap.max = 0
cpu.weight = 100
└── per-alloc: memory.max · cpu.max · pids.max from the spec's resources block
Calling mlockall on a Go control plane is rejected outright: the GC grows
the heap unpredictably and RLIMIT_MEMLOCK turns pin overflow into hard
allocation failure; the lock itself could crash kanead. The guarantee
comes from memory.min (the kernel refuses to reclaim the floor under
pressure), the OOM score, and no swap in the slice.
Per-alloc limits are enforced where declared. Omit resources
(or a field of it) and the alloc is unbounded: all cores, all allocatable memory,
capped only by the workload parent's collective ceiling, which, with the
control-plane floor, is the isolation that actually protects the platform. A
default pids.max applies to every alloc; resources { pids = 2048 }
declares a different value the same way the other limits are declared (the cap is
always on, omission means 256). Declared resources are
also the admission units: kanea plan renders the workload budget and
apply refuses a total above it unless the operator has explicitly enabled
oversubscription.
Metrics and autoscaling
Three scrapers feed one bounded in-memory time series: containerd cgroup metrics (CPU, memory), the edge (requests per second, latency percentiles for exposed services), and the datapath's own per-CPU map counters (east-west flows and drops, on by default). Roughly 27 MiB at the 2 000-alloc target, and there is a test that says so.
A missing metric and an idle service lead to opposite decisions, so the distinction is preserved end to end: in the time series, in the evaluator, in the Prometheus exporter and in the dashboard. Each layer has a test asserting it.
The evaluator applies the scaling block with guardrails and a circuit
breaker, and the budget is 20 seconds from a sustained breach to a decision: a
15-second averaging window (three samples at the 5-second scrape resolution) plus one
evaluation tick. A large spike decides sooner.
The autoscaler is not a second scheduler. It writes one number (the desired
count) and the reconciler converges. kanea scale uses the same route,
which is why manual and automatic scaling cannot disagree about mechanism.
State replication and restore
- Change data capture. Store mutations produce change segments; snapshots and segments are replicated to any S3-compatible bucket. The client is hand-written against the REST API (SigV4 and four verbs) and CI exercises it against MinIO in both addressing styles, path-style and virtual-hosted.
- Archives are chunked AEAD, with keys HKDF-derived from the master key rather than being it. The last chunk is sealed under different additional data, which is the only thing standing between a truncated snapshot and a restore that decrypts cleanly into half a platform.
- Manifests are unencrypted and hashes are over the ciphertext, so someone holding the bucket and no key can still see what is there and whether it is intact, which is the question you ask before going to find the escrowed key.
- Restore is staged, never performed in place. It can be requested over the API, and the daemon performs it at the next start, before anything opens the Store. The API has no method that restores; that is the interface, not a check.
- The replication cursor is derived from the sink, never stored: writing it to the Store would emit a change that needs shipping, which would write it again.
- The destination can change at runtime; the dashboard's Settings page, or
PUT /v1/settings/backup. A new destination is probed with a real test write before anything commits, so a typo is a refusal with the old replication untouched; the old destination receives one final segment ship, the new one an immediate full snapshot. The unit's flags remain the seed: a settings record, once written, wins, and deleting it reverts.
The master key is generated by kanea init, shown once, and must be typed
back; it is discarded if that fails. Without it, every archive is unreadable. The
DR runbook is
worth reading before you need it.
API, dashboard and MCP
The API is REST plus one multiplexed WebSocket, and every route is deny-by-default:
including the WebSocket and every MCP route. The exceptions are enumerable:
/login, the ACME challenge path, and /healthz - and even
the last one keeps its answers short. Unauthenticated, it returns the status and
the OIDC issuer (a login screen needs that before it holds a credential) and
nothing else; version, PID, store index, listen address and uptime need an
identified caller. A bad token still gets the slim 200, so a load balancer
forwarding a stale Authorization header cannot flap a health check.
There is exactly one authentication mechanism that is not the §13 one: the git webhook route, which uses a per-project HMAC over the raw body, rejects replays, and is audited. It never deploys from the request: it marks the project, and the sync loop re-reads the source over Kanea's own credential.
The dashboard is a React SPA embedded in the binary with go:embed and served
by the daemon. A service page charts CPU, memory, request rate and p95 over a real
time axis, seeded on the first frame of the subscription that carries them
(since v1.79) so a chart draws its window on arrival rather than growing one point per
scrape, and an empty panel distinguishes still arriving from
no samples yet, which is also the honest answer for the ten seconds after
a restart, during which no rate can exist. It streams logs live (virtualized, filterable, copy, download, and a
full-screen view that keeps your place in the buffer across the toggle), shows a
restart or deploy as rollout progress (the same spec-hash rule the planner uses,
served on the wire since v1.64) and gives an admin a shell into any running
alloc, over the same exec websocket the CLI uses: the CSRF token rides the
handshake as a subprotocol, since a browser cannot set the header there. Its Settings
page shows the node's configuration and lets an admin change
what changes at runtime: the backup destination and the notification channels
(node-wide defaults and per-project, each with a test button); plus accounts, API
tokens and the audit log, one tab each; what belongs to the unit (listen address,
subnets, DNS, the
port policy) is shown read-only with a note saying so. A pinned Updates page
holds the rest of the node's software: check for or install a Kanea release (the
upgrade flow, run by the node on itself), and read
the OS's pending packages, security count, list freshness and reboot flag beside
the manifest-pinned host components - reported, never installed: the page names
the apt command to run on the node. The MCP server exposes 24 tools
in three tiers (read, mutate, destructive) over stdio and streamable HTTP.
Every tool reaches the platform by making an HTTP request against the API's own handler. A tool's only verb is "send this request", so it can never be more privileged than the credential its caller presented: nothing in the MCP package may hold a Store, a secrets store, or an auth store. That is what makes "no side channels" structural rather than a rule somebody has to remember.
There are also no secret tools at any tier. The safety requirement is that no tool returns a secret value; the implementation goes further and gives an agent no secrets verb whatsoever, and a test fails if one appears.
Exposing the API and dashboard
The API and the dashboard are one listener: the dashboard is served by
kanead on the same address the REST API, the WebSocket and the MCP
HTTP transport bind, in front of kanea-edge entirely; neither is a
service, and neither takes an expose block. The consequence worth
internalising is that the dashboard's URL carries the listener's port
- https://<name>:8600 by default, not :443,
which belongs to the edge and has no route to the dashboard. There is a
worked example in the bind stanza's
own section. Exposing them means
setting where that listener binds, and there are two ways: the
--listen/--listen-cert/--listen-key flags
(asked by kanea init and rendered into the
unit), or (since v1.61) a bind stanza in the server config, which
survives re-runs and makes moving the listener later an edit plus
systemctl restart kanead, never a re-init:
# /etc/kanea/kanea.hcl
bind {
api_addr = "192.168.1.10:8600"
api_tls = "self-signed" # acme | self-signed | provided | plaintext
# api_domain = "kanea.home.example" # required by acme; names a self-signed cert
# api_cert = "/etc/kanea/api.pem" # provided only; always with api_key
# api_key = "/etc/kanea/api.key"
}
| Field | Type | Notes |
|---|---|---|
api_addr | string | The host:port the API and dashboard bind. Required by every other field in the stanza: TLS with nothing to serve it on is a parse error. |
api_tls | string | acme, self-signed, provided or plaintext; the same mode vocabulary services use. Unset resolves at the daemon: a declared pair means provided, loopback means plaintext, and anything beyond loopback refuses there. |
api_domain | string | Required by acme (an IP cannot hold an ACME certificate). For self-signed it names the certificate, and it is required when api_addr binds every interface (":8600", 0.0.0.0): otherwise the certificate would have no name at all. |
api_cert, api_key | string | Your own PEM pair, for provided only. Always together, and refused beside a managed mode or plaintext; a control that cannot act is refused, never silently dropped. |
What each mode gives you:
self-signed: a certificate from the node's own CA, the onekanea ca showinstalls on your devices, with a real IP SAN when the address is bare. Renewed automatically. The right default on a home network.acme: a Let's Encrypt certificate forapi_domain, issued and renewed by the same account and pass that serve your services. Needs--acme-emailconfigured on the daemon, refused at startup without it.provided: your ownapi_cert/api_keypair, for a certificate something else manages.plaintext: explicit HTTP. Allowed beyond loopback because you typed it, logged loudly, and it implies the insecure-cookie posture; a Secure cookie over plain HTTP is a login that silently fails.
Precedence: an explicit --listen always beats the stanza, and
--listen none keeps the node socket-only regardless of the file. When
the stanza is declared and --listen was not passed,
kanea init skips the listen question and renders no listen flags into
the unit: the file owns the listener. On the flags path, a non-loopback
--listen with no explicit --listen-cert/--listen-key
pair no longer refuses: init provisions a default 10-year self-signed pair at
/etc/kanea/api.crt/api.key (minted once, never re-minted)
and points the daemon at it. The one refusal left is an unspecified host
(0.0.0.0): a SAN needs something to name, so bind a specific address
or set api_domain in the stanza. The stanza's full field list and every
refusal it can raise are in
Node configuration → bind.
One consequence worth knowing when adopting the stanza on an existing node: an init
run from before it existed rendered --listen 127.0.0.1:8600 into the
kanead unit, and that flag shadows the file; the stanza appears to do nothing.
Remove the flag from the unit's ExecStart, then
systemctl daemon-reload and restart; the
troubleshooting page has the steps.
Node configuration
/etc/kanea/kanea.hcl is the node's own file: the settings that belong
to whoever owns the machine rather than to whoever writes a job spec. A spec can
reference what this file permits and can never add to it, which is the whole
point of the split - GitOps deploys specs automatically, so anything a spec
could declare, anyone who can push to a synced repository could declare.
It does not exist by default, and that default is "off": a node with no file grants
nothing, allows no host paths, and takes its listener and resolvers from flags.
kanea init creates the /etc/kanea directory but
never writes this file, on a first run or a re-run.
The file
Four properties decide how it behaves, and each one surprises somebody:
-
It is read once, at daemon start. One stat, never a poll. There is no
reload, no
SIGHUPand no watcher, so an edit means a restart - see applying a change. A grant is a decision, not a thing that rotates behind your back. - Absent is off, but malformed is fatal. A missing file is every setting's zero value and no error. A file that does not parse, or breaks a stanza's rules, stops the daemon starting. There is deliberately no keep-last-good and no partial load: a policy file half-applied is worse than one refused.
- It is trust-checked before it is parsed. It must be a regular file (not a symlink, not a directory), owned by root or by the daemon's own uid, and writable only by its owner. Anything else refuses startup. World-readable is fine; this is policy, not a secret.
- An unknown stanza is a warning; an unknown attribute is an error. See the stanza list for why the asymmetry is intentional.
Creating it, if it is not there:
sudo install -o root -g root -m 0644 /dev/null /etc/kanea/kanea.hcl sudo $EDITOR /etc/kanea/kanea.hcl sudo systemctl restart kanead
What may go in it
Seven stanzas are read. Everything else is accepted and warned about by name.
| Stanza | Repeats? | What it decides |
|---|---|---|
bind | once | Where the API and dashboard listen, and how that listener gets its certificate. |
storage | once | Which host directories a host volume may mount. |
dns | once | Which resolvers the internal DNS forwards external names to. |
images | once | The node's default for where a service's images may come from, and the image GC's gc block. |
variables | once | Node-wide defaults for spec variables. |
device | many | A named device grant, and the projects that may claim it. |
socket | many | A named socket grant, and the projects that may claim it. |
A top-level block this version does not read is loaded, ignored, and
named at startup in a server config carries stanzas this version
does not read warning - never silently swallowed. That is what lets
the file carry settings a future version will read. But an unknown attribute
inside a stanza that is read is a hard parse error, because there
it is almost always a typo, and a silently ignored
allowed_host_path would be a security control that quietly did
nothing.
bind: the API and dashboard listener
Putting the listener here instead of in the unit means moving it later is an edit
plus a restart, never a re-init, and a kanea init re-run will not
overwrite it. When api_addr is declared and --listen was
not passed, init skips the listen question and renders no listen flags into
the unit - a unit that repeated the file's answer would turn the file off.
# /etc/kanea/kanea.hcl
bind {
api_addr = "192.168.1.10:8600"
api_tls = "self-signed" # acme | self-signed | provided | plaintext
# api_domain = "kanea.home.example" # required by acme; names a self-signed cert
# api_cert = "/etc/kanea/api.pem" # provided only; always with api_key
# api_key = "/etc/kanea/api.key"
}
| Field | Notes |
|---|---|
api_addr | The host:port to bind. Required by every other field: TLS with nothing to serve it on is a parse error. |
api_tls | acme, self-signed, provided or plaintext - the same vocabulary services use. What each one gives you is in Exposing the API and dashboard. Unset resolves at the daemon: a declared pair means provided, loopback means plaintext, anything beyond loopback refuses. |
api_domain | Required by acme (an IP cannot hold an ACME certificate). For self-signed it names the certificate, and it becomes required when api_addr binds every interface. |
api_cert, api_key | Your own PEM pair, for provided only. Always together. |
edge_http, edge_https | Accepted, warned about, and not read. They are part of the fuller design sketch; setting one today does nothing but add its name to the startup warning. |
Every contradiction is a parse error rather than a silent preference:
| If you write | It says |
|---|---|
Only one of api_cert/api_key | they go together |
A pair, or any TLS field, with no api_addr | TLS with no api_addr to serve it on |
provided with no pair | provided needs api_cert and api_key |
plaintext beside a pair | plaintext beside a TLS pair drops the pair; remove one |
acme or self-signed beside a pair | it issues its own certificate; remove the pair |
acme with no api_domain | an IP cannot hold an ACME certificate |
self-signed on 0.0.0.0 with no api_domain | binds every interface; set api_domain so the certificate has a name |
Any other api_tls value | use one of acme, self-signed, provided, plaintext |
A worked example: the dashboard on a public name
The most common thing anyone wants from this stanza - a real certificate on a
real name, replacing the self-signed pair kanea init provisions:
# /etc/kanea/kanea.hcl
bind {
api_addr = "0.0.0.0:8600"
api_tls = "acme"
api_domain = "kanea.apps.example.com"
}
The dashboard is then at:
https://kanea.apps.example.com:8600
https://kanea.apps.example.com (no port) goes to
:443, which is kanea-edge, a different process that has never
heard of the dashboard. The dashboard is served by kanead itself and
is not a service, so no route to it can exist and
the edge has nothing to forward to.
What you get is not a 404. On a node with no exposed services the edge holds no
certificates at all, so the TLS handshake fails outright -
SSL_ERROR_INTERNAL_ERROR_ALERT in Firefox,
tlsv1 alert internal error from curl - which reads like a
broken certificate rather than a wrong port. Check the port before you debug the
certificate.
Three things have to be true for that certificate to issue, and only the first is in this file:
-
An A record for
kanea.apps.example.compointing at the node. The name is validated by the CA connecting to it, so it must resolve before the first attempt. -
--acme-emailon the daemon. It is a flag, andkanea initdoes not write it into the unit, so it goes in a drop-in. Without it nothing is ever requested from a CA, and the log says so at every start. Note there is noacmestanza in this file: writing one is accepted, ignored, and warned about by name, which is a confusing way to discover the flag exists. -
Port 80 reachable, with
kanea-edgerunning. The edge answers/.well-known/acme-challenge/on :80 even though the certificate is for kanead's listener, so a stopped edge means no dashboard certificate.
Binding loopback does not put the dashboard behind the edge.
api_addr = "127.0.0.1:8600" makes it reachable only from the node
itself; nothing bridges :443 to it, because a route's upstream is always a
service VIP.
Until the first issuance the handshake fails by name, rather than serving
something weaker. The listener binds a few seconds before the certificate arrives,
so a restart has a brief window where the dashboard is hard-unreachable, and a node
where ACME cannot issue stays that way with the real reason only in
journalctl -u kanead. That is deliberate: a wrong certificate is worse
than a refused connection.
Wanting the bare name with no port means giving the dashboard :443 and
moving the edge with kanea edge --https :8443. Two processes cannot
share a port, so that is the whole decision: every service you deploy then answers
on :8443 instead. It is a reasonable trade on a node whose apps live
elsewhere, and a bad one on a node that hosts anything.
storage: host-volume allowlist
A host volume mounts a directory that already exists on the node, and
whether it may be mounted is not the spec's decision. The allowlist is empty by
default, so host volumes do nothing until this stanza names a parent.
# /etc/kanea/kanea.hcl
storage {
allowed_host_paths = ["/srv/kanea", "/mnt/media"]
}
The check runs after symlink resolution, on the object rather than the
spelling. Three things are refused at startup: a prefix that is not absolute, a
prefix that does not exist, and "/" itself - which "allows the
entire filesystem; list the directories you actually intend to share". An explicit
empty list is legal and means the same as no stanza: no host volumes.
A spec then mounts with storage "x" { type = "host" path = "/srv/kanea/x" }.
A path outside every prefix fails the alloc. A path that does not exist is refused
too, so a typo cannot become a silently empty volume - unless the spec sets
create = true, which makes Kanea create it, still only inside a prefix
you allowed. Creating is not owning: Kanea never chowns a host volume.
device and socket: passthrough grants
A job spec names a grant, never a path. There is no field in the spec for a device node or a socket path: not a validated one, none at all, because the node holds the mapping and the spec only asks for it by name.
# /etc/kanea/kanea.hcl
device "gpu" {
nodes = ["/dev/dri/card0", "/dev/dri/renderD128"] # a list; required
allow = ["media"] # projects that may claim it; required
# mode = "rw" # r, w, m; default "rw"
}
socket "containerd" {
path = "/run/kanea/containerd.sock" # a single string; required
allow = ["ops"]
}
| Attribute | On | Notes |
|---|---|---|
| The block label | both | The grant name a spec references. Must be a DNS-1123 label, or no spec could name it. Duplicates within a kind are refused. |
nodes | device | A list of absolute device paths. Required and non-empty. |
path | socket | A single string. Required, absolute, no .., not /. |
allow | both | Required and non-empty: the projects that may claim the grant. Each must be a valid project name. There is no projects attribute - this is it. |
mode | device | Optional cgroup permissions, some combination of r, w and m. Defaults to "rw"; m (mknod) only if you write it. |
A spec claims one with device "dri" { grant = "gpu" } or
socket "rt" { grant = "containerd" mount_path = "/run/containerd.sock" }.
Paths are resolved and type-checked per alloc, not at load, so a grant naming
a device that is absent or is not a device fails that alloc with a message
naming the grant, rather than stopping the daemon. A transcoder silently running
without its GPU would look healthy and do the wrong thing, so it fails instead. A
grant the node does not hold is refused with a list of the ones it does.
A container holding the runtime socket can start other containers without the
hardening defaults. There is no containment story and none is claimed:
nosuid,noexec,nodev restrict the filesystem entry, not the protocol
spoken over it. The control is that granting it happens on the machine, in a file
no spec author can write, and names one project.
dns: upstream resolvers
The internal resolver answers <service>.<project> names
itself and forwards everything else. By default it forwards where the host forwards
- every nameserver in /etc/resolv.conf, in the file's own order,
read once at startup. This stanza pins that list instead, which is what you want on
a node whose resolv.conf is DHCP's to rewrite:
# /etc/kanea/kanea.hcl
dns {
upstreams = ["1.1.1.1", "10.0.0.53:5353"] # a bare address gets :53
}
Each entry must be an address or a host:port pair, checked when the
file parses. An empty list is refused by name: a stanza meaning "no
upstreams" would silently turn every external name into SERVFAIL, so if that is
what you want, remove the stanza. Allocs never see this list - their
resolv.conf names Kanea's resolver alone, and it forwards on their
behalf from the host's own namespace.
images: the default pull policy
Where images may come from for a service that does not say (R33). The one case this exists for is a node that is not supposed to reach a registry at all - air-gapped, or with images preloaded by an offline bundle - where the default's slow pull-and-fail blames a registry the node was never going to contact:
# /etc/kanea/kanea.hcl
images {
pull_policy = "never" # "if-not-present" (default) | "never"
}
Precedence is the file's usual one: --image-pull-policy on
kanead wins and says so in the log, this stanza is next, and
if-not-present when neither - which is what every node did
before the field existed. A service's own task.pull_policy wins
over all of it. always is refused here: it means
per-service auto-update, and turning that on for every service on a node is
not a default anybody asked for.
images.gc: cleaning up old images
Nothing else ever deletes an image, so every deploy, auto-update pin and rollback leaves its predecessor behind. The image GC collects them - on by default, and this block tunes or disables it:
# /etc/kanea/kanea.hcl
images {
gc {
enabled = true # --image-gc off on kanead also disables it
interval = "12h" # sweep cadence; floor 10m
min_age = "48h" # younger images are never touched; floor 1h
keep = 2 # newest N per in-use repository, as rollback material
}
}
An image is deleted only when all three rules agree: nothing
references it (no service's declared image, pinned digest, rollback
target or init step, and no running alloc), it is older than
min_age, and it is not among the newest keep of a
repository something still uses. A node whose default
pull_policy is "never" refuses to collect at all:
preloaded images cannot be re-pulled, so deleting one is unrecoverable.
kanea images lists what the node holds with each image's
in-use verdict, and kanea images --clean runs one sweep now
(admin, audited).
variables: node-wide spec defaults
Values every spec on this node can reference as ${name}, so one node's
domain or storage root does not have to be repeated in every file:
# /etc/kanea/kanea.hcl
variables {
domain = "home.lan"
media_root = "/srv/kanea/media"
}
Precedence is node under spec under caller: a spec's own variables
block wins over the node's, and --var on the command line wins over
both. Values are strings, numbers or bools - not lists or objects. A
definition may reference the built-ins but never a sibling, so there is deliberately
no ordering or cycle story. The R2 built-ins and service are reserved
names.
GET /v1/vars serves this stanza to any authenticated caller,
which is exactly why nothing secret may live in it. Secrets are
secret: references, resolved per alloc into a tmpfs file; see
Secrets.
Applying a change
The whole file is read once at startup, so there is one apply step and it is the same for every stanza:
sudo systemctl restart kanead
systemctl daemon-reload is not part of this. That is for
changes to the systemd units, which this file is not. Running it after
editing kanea.hcl is harmless but does nothing.
What a restart costs
Less than people expect, and it is worth knowing before you hesitate over a production node:
-
Running workloads keep running. The unit carries
KillMode=processprecisely so a control-plane restart does not take the containers with it; their shims outlive it. -
Ingress keeps serving.
kanea-edgeis a separate process with noAfter=relationship in either direction, and it serves the last route snapshot it read from disk whether or not kanead is up. - What pauses is change, for the second or two it takes to come back: deploys, scaling decisions, certificate renewal, the API and the dashboard. See the process list for the same split across every daemon.
Malformed is fatal by design - no keep-last-good, no partial load -
so a file with a typo means kanead refuses to come back, with the
parse error naming the file and line in
journalctl -u kanead. Fix the file and restart again. If you need
the node up first and the config sorted out afterwards, add
--config off to the unit's ExecStart
(systemctl edit kanead), which starts the daemon as though the file
did not exist, so you can fix the file with the dashboard and the API back up.
There is no offline validator for this file today: the restart is the check. On a node you cannot afford to have refuse a start, keep a known-good copy beside it before editing.
Confirming it took effect
A few lines in journalctl -u kanead answer nearly every "the file
seems to do nothing" question:
| Log line | Means |
|---|---|
server config loaded | The file was found, trusted and parsed. It carries the path. |
server config carries stanzas this version does not read | Something in the file is being ignored, and the message names it. A misspelled stanza shows up here rather than as an error. |
dns upstreams from server configapi/dashboard listener from server config | That stanza won and is in effect. These are the positive confirmations, and their absence is as informative as their presence. |
… is not consulted (--flag wins) | A flag on the unit is overriding that half of the file - there is one of these for --dns-upstream, --image-pull-policy, --listen, --allowed-host-paths and --passthrough-config. This is the single most common cause of a stanza appearing to do nothing; see precedence. |
Flags, the file, and precedence
Every half of the file has a flag that overrides it, and the rule is the same throughout: an explicit flag wins, and the daemon says so in its log. Nothing is ever merged - whichever source supplies a setting supplies all of it, so you never end up with an address from one place and its certificate from another.
| Stanza | Flag that wins | Turn it off with | If neither is set |
|---|---|---|---|
bind | --listen (plus --listen-cert/--listen-key) | --listen none | Unix socket only |
storage | --allowed-host-paths | --allowed-host-paths off | No host volumes |
device, socket | --passthrough-config <file> | --passthrough-config off | No grants |
dns | --dns-upstream | --dns-listen off (disables the resolver) | The host's resolv.conf |
images | --image-pull-policy | n/a - there is always a policy | if-not-present |
variables | none - file-only | --config off | No node variables |
| the whole file | --config <file> | --config off | /etc/kanea/kanea.hcl if it exists |
The disable words are not the same. It is --listen none but
off everywhere else. none was chosen for the listener
because it matches what kanea init asks you to type at the prompt.
--config off disables all six stanzas, not just the security
ones. The file is never even stat-ed, so a present and perfectly good
dns or variables stanza goes with it. And
variables is the one half with no flag of its own, which is why the
daemon probes the file whenever --config does not say
off.
A flagged --passthrough-config file gets the same trust check as
kanea.hcl itself: arriving by argv does not make a policy file more
trustworthy. Note also that --allowed-host-paths does not merge
with the stanza; it replaces it, and the daemon logs
server config storage stanza is not consulted when both are present.
A complete file
Every stanza this version reads, in one file:
# /etc/kanea/kanea.hcl: the node's, never the repository's. # Read once at startup. Edit, then: sudo systemctl restart kanead bind { # where the API and dashboard listen api_addr = "192.168.1.10:8600" api_tls = "self-signed" # acme | self-signed | provided | plaintext } storage { # parents that `host` volumes may mount from allowed_host_paths = ["/srv/kanea", "/mnt/media"] } dns { # pin resolvers; omit to follow the host's resolv.conf upstreams = ["1.1.1.1", "9.9.9.9"] } variables { # node-wide spec defaults; never secrets domain = "home.lan" media_root = "/mnt/media" } images { # the node default; a task's own pull_policy wins pull_policy = "if-not-present" gc { # unused-image cleanup; on by default interval = "12h" min_age = "48h" keep = 2 } } device "gpu" { # claimed as: device "dri" { grant = "gpu" } nodes = ["/dev/dri/renderD128"] allow = ["media"] } socket "containerd" { # root-equivalent for whoever holds it path = "/run/kanea/containerd.sock" allow = ["ops"] }
Names and DNS
Kanea deals in two separate name spaces, and almost every DNS question comes from them being mistaken for one. One is internal, resolved by Kanea, and needs nothing from you. The other is public, resolved by the world's DNS, and is the only part where you create records.
| Internal | Public | |
|---|---|---|
| Looks like | web.shop.kanea | web.shop.apps.example.com |
| Resolved by | Kanea's own resolver | Your DNS provider |
| Answers with | The service's virtual IP | The node's public address |
| Reaches | The allocs directly, east-west | The edge on :443, then the allocs |
| Records you create | None. It is automatic | One wildcard, usually |
| Who can use it | Allocs on this node | Anything that can reach the node |
Internal names, between services
Every alloc gets a generated /etc/resolv.conf pointing at Kanea's
resolver, with a search path that makes the short forms work:
nameserver 10.244.0.1 search shop.kanea kanea options ndots:1 timeout:2 attempts:2
So from a service inside project shop:
| You write | It resolves | When to use it |
|---|---|---|
db | db.shop.kanea | A service in the same project. |
cache.data | cache.data.kanea | A service in another project. |
db.shop.kanea | itself | When you would rather be explicit. |
The address that comes back is the service's virtual IP, and the connection is
load-balanced across its healthy allocs at connect() by the datapath
- there is no proxy in the path and no client-side round-robin to get wrong.
The resolver lives at the node CIDR's .1
(10.244.0.1 by default, moving with --node-cidr), and it
is an anchor address, never a wildcard.
A cross-project name resolves for anyone, because DNS is not the security
boundary. The datapath is, and it is deny-by-default between projects: the
connection is dropped unless the callee lists the caller in
allow_from. This presents as a name that resolves fine and a
connection that times out, which reads like a network fault rather than a policy
decision. If a cross-project call hangs, check allow_from on the
service being called before you look at DNS.
The base domain
An exposed service's public name is
<service>.<project>.<base-domain>, so
--base-domain apps.example.com gives the web service in
project shop the name web.shop.apps.example.com. A spec
can name its own domains instead, with expose { domains = [...] }; the
base domain is what it falls back to when it does not.
kanea init does not set this, and the unit is rewritten
--base-domain, --acme-email and
--tls-default are daemon flags that init never renders into
kanead.service - it writes only seven, and these are not among
them. Without a base domain a service has no generated FQDN at all, so this is
the step between "deployed" and "reachable by name".
Set them in a drop-in, not by editing the unit: kanea init
rewrites kanead.service on every run, so a direct edit is lost the
next time you regenerate units, while a drop-in survives.
sudo systemctl edit kanead and add:
[Service] ExecStart= ExecStart=/usr/local/bin/kanea agent --data-dir /var/lib/kanea \ --log-dir /var/log/kanea/allocs --network ebpf \ --node-cidr 10.244.0.0/24 --cluster-cidr 10.244.0.0/16 \ --edge-group kanea-edge \ --base-domain apps.example.com --acme-email you@example.com
The empty ExecStart= is required: systemd appends to that directive
otherwise, and a unit with two ExecStart lines fails to start. Copy
the existing line from systemctl cat kanead and add to it, rather
than retyping it from here - your node's flags may differ. Then
sudo systemctl daemon-reload && sudo systemctl restart kanead.
The records you create
One wildcard covers every service, present and future, which is why it is the usual answer:
| Record | Type | Points at | Covers |
|---|---|---|---|
*.apps.example.com | A | The node's public IPv4 | Every generated service name |
*.apps.example.com | AAAA | The node's public IPv6 | The same, for v6 clients |
kanea.example.com | A | The node's address | The dashboard, if you gave it a name |
shop.example.com | A | The node's public IPv4 | One custom domains entry |
The AAAA record is about the node's own public address and has nothing to do
with the opt-in container IPv6 feature: the edge binds :443 on whatever
the host has, so a node with public v6 serves v6 clients whether or not allocs have
v6 addresses.
The dashboard is not a service and takes no expose block, so it is not
covered by the wildcard unless you happen to name it under the base domain. Give it
a name with bind.api_domain and point a
record at the node; without one it is reached by IP, and its certificate is the
self-signed pair kanea init provisioned. Remember that you reach it at
the listener's port - https://kanea.example.com:8600
- because the edge, which owns :443, has no route to it.
Nothing needs a record per alloc, and nothing needs a record that changes when you
scale or redeploy: the edge routes by Host header and the datapath
handles the rest, so the DNS side is static once it is set up.
Certificates and the challenge
Which ACME challenge Kanea can use decides what you have to arrange, and the two are not interchangeable:
| HTTP-01 (default) | DNS-01 | |
|---|---|---|
| Needs | Port 80 reachable from the internet, and the name already resolving to the node | An RFC 2136 nameserver accepting TSIG-signed updates |
| Configure | Nothing beyond --acme-email | --acme-dns-server and the TSIG flags |
| Wildcard certificates | Impossible - there is no single host to reach | Yes |
| Works on a private node | No | Yes |
HTTP-01 is the default and needs no configuration: the edge serves the challenge
path on :80, and kanead reaches its own edge first to
confirm the challenge is actually being served before it asks the CA to try
(--acme-verify-url). The record has to exist and point at the node
before issuance, so create DNS first, deploy second.
DNS-01 is RFC 2136 dynamic update, with TSIG:
--acme-dns-server ns1.example.com:53 --acme-dns-zone example.com # default: the challenge name's parent --acme-dns-tsig-key kanea-updates --acme-dns-tsig-secret secret:shared/tsig # a reference, never the literal --acme-dns-tsig-algorithm hmac-sha256.
There is no provider integration - no Cloudflare, Route 53 or
Azure DNS plugin. RFC 2136 is the whole DNS-01 story, which suits BIND, Knot and
PowerDNS, and means a provider that only offers a proprietary API needs a zone you
run yourself, or a CNAME delegation of _acme-challenge into one.
The TSIG secret is required whenever --acme-dns-server is set,
and must be a secret: reference: a literal is refused by name, and so
is an absent one, because unsigned dynamic updates would let anyone on the network
answer a challenge.
Certificates are per service until a node has 20 generated names, at which point it collapses them into one wildcard for the base domain - but only if a DNS-01 solver is configured, since HTTP-01 cannot validate a wildcard. The number comes from Let's Encrypt's limit of 50 certificates per registered domain per week, and every redeploy that changes a name spends one.
Custom domains are never collapsed. A name you declared in
expose { domains } is somebody else's zone, and Kanea has no standing
to ask a CA for *. of it. Use --acme-directory staging
while you are getting this working: the staging CA has far looser limits, and its
certificates are not trusted, which is the point.
Private networks and homelabs
No public name, no port 80 from the internet, and still real HTTPS. Two routes, and the difference is whether a browser has to be taught to trust anything:
-
The node's own CA (
--tls-default self-signed). Point a wildcard record on your local resolver - a Pi-hole, a router, an internal BIND - at the node, then install the CA once per device withkanea ca show > kanea-ca.crt. No CA to reach, no rate limit to spend, works entirely offline. The cost is one trust-store install per device, and devices you do not administer cannot be taught. - Real certificates over DNS-01. If you own a public domain and can run RFC 2136 updates against it, the node gets publicly-trusted certificates for names that resolve only on your LAN. Nothing has to be reachable from the internet, because the challenge is answered in DNS rather than over HTTP. Every device trusts it with no setup.
Either way the record points at the node's LAN address, which is ordinary
split-horizon DNS: the name resolves internally and either does not resolve
externally or resolves elsewhere. Kanea does not care which, because it never
resolves its own public names - the edge routes on the Host
header it is given.
Prefer not to deal with names at all? Publish a port
and reach the service at http://<node>:8096. That path needs no
DNS and no certificate, and ip_restriction bounds who may use it.
Job spec reference
Specs are HCL v2. If you have written Nomad job files the shape will be familiar,
though the vocabulary is Kanea's. Every rule below is enforced when the file is parsed
(errors carry file, line and column) and kanea plan is where you see
them.
A service needs an image and nothing else. This is a complete, valid spec:
spec_version = 1
project "demo" {}
service "web" {
project = "demo"
task "app" {
image = "nginx:1.27-alpine"
}
}
Everything on this page is additive to that.
Top level
A spec file contains spec_version and any number of
project, service and storage blocks, plus an
optional variables block of shared values. Multiple
files applied together are parsed as one set, and file order is irrelevant: a
service may reference a project declared in another file.
| Field | Type | Notes |
|---|---|---|
spec_version | number | Currently 1. Future spec revisions are gated on this field, which is how an upgrade can tell an old file from a new one. |
project "<name>" | block | Zero or more. Name is a DNS-1123 label. |
service "<name>" | block | Zero or more. |
storage "<name>" | block | Zero or more. May also be declared in the server config. |
variables | block | Optional. Shared values the rest of the spec references as ${name} (R30). |
variables
Declare a value once, reference it anywhere as ${name}, or as a bare
identifier where HCL takes an expression, like count. The node may
supply defaults from a variables stanza in its own
/etc/kanea/kanea.hcl; the spec's block wins on a collision, and
pipeline-supplied values (${GIT_SHA_SHORT} and friends) sit above both.
variables {
domain = "shop.example.com"
replicas = 3
}
service "web" {
project = "shop"
count = replicas
expose {
domains = ["${domain}", "www.${domain}"]
}
}
Values are strings, numbers or bools: a list or object is refused by name, and so
is redeclaring a name or shadowing a built-in. A variable's value may reference node
variables and built-ins, never another spec variable. Variables are not
secrets: the node's stanza is served to any signed-in caller over
GET /v1/vars, so a secret stays a secret: reference; a
variable whose value contains one is fine.
project
A named group of services, and the isolation boundary for network policy, secrets, containerd namespaces and DNS.
| Field | Type | Notes |
|---|---|---|
description | string | Free text, ≤ 512 chars. Shown in the dashboard. |
git | block | Optional GitOps source for this project. |
notifications | block | Optional channels and filters. |
git
| Field | Type | Notes |
|---|---|---|
url | string | Required. Clone URL: https://, ssh://, scp-form or a local path - http:// and git:// are refused (the deploy credential would travel in cleartext / nothing authenticates the content), and credentials in the URL's userinfo are refused in favour of auth_ref. |
branch | string | Defaults to the repository's default branch. |
path | string | Directory within the repository holding specs, e.g. .kanea/. |
auth_ref | secret ref | Deploy key or token. Project-scoped like every other reference. |
webhook_secret_ref | secret ref | HMAC key for the push webhook. |
poll_interval | duration | How often to poll when no webhook arrives. |
require_approval | bool | Sync marks the project; a human promotes. |
A synced spec that declares a different project is refused. It is the same boundary that scopes secrets, and it is the only thing between "can push to one repo" and "owns every service on the node".
The webhook never deploys from the request either. It marks the project as dirty and the sync loop re-reads the source over Kanea's own credential, so a forged body can at most cause a legitimate sync to happen sooner.
notifications
Up to five channel blocks; telegram, slack (Discord accepts
the same shape on its /slack endpoint), ntfy,
smtp, webhook; plus filters:
| Field | Type | Notes |
|---|---|---|
on | list(string) | Event globs, e.g. ["deploy.failed", "scale.*"]. Validated against the known event vocabulary at parse time; a filter that could never match is a spec error, not a silence. |
severity | string | Floor: info, warning, error. Composes with on as an AND. |
telegram | block | chat_id, token_ref |
slack | block | url_ref: an incoming-webhook URL is a credential in path form, so it is referenced, never inlined |
ntfy | block | url, token_ref |
smtp | block | host, port, from, to, username, password_ref |
webhook | block | url, secret_ref (HMAC signature over the body) |
Egress is checked at dial time against every resolved address, and redirects are refused. A hostname is not a destination.
storage
A named volume backend. Declared here or in the server config; services mount it by
name with a volume block.
type | Fields | Notes |
|---|---|---|
local | - | A directory on the node, managed by Kanea. |
host | path, create | A directory the operator already owns, or, with create = true, one Kanea creates inside an allowed prefix. See R15: it does nothing until an operator allowlists its parent in /etc/kanea/kanea.hcl's storage { allowed_host_paths = […] } stanza (or --allowed-host-paths, which wins). |
nfs | server, export, options | |
smb | server, share, auth_ref, options | |
s3 | bucket, endpoint, auth_ref, mode | mode = "ro" selects mountpoint-s3, "rw" selects s3fs. |
auth_ref is R5-scoped like every other credential reference: every service
mounting the storage must be allowed to read it (its own project or shared/),
because mounting is reading. And an endpoint must be https:// -
the mount helper sends the resolved credential to that address on every request.
Every file operation costs an object-store round trip; around 30 ms. Listing or
creating a 200-file directory takes tens of seconds, no driver implements
truncate (s3fs silently no-ops it), and a FUSE call against a dead
backend blocks uninterruptibly for tens of seconds. Use them for bulk, read-mostly
data. Never for a hot path or many small files.
service
| Field | Type | Default | Notes |
|---|---|---|---|
project | string | - | The owning project. |
description | string | - | Free text, ≤ 512 chars. |
count | number | 1 | Replicas. Must sit inside scaling's min/max if that block is present. |
hardening | string | compatible | The service's posture (R13). restricted is the strong profile in one word: the task must declare a non-zero user, runs with no capabilities at all, and a capability grant beside it is refused. Init steps are deliberately unchanged - chown as root in a step, run the task restricted. |
depends_on | list(string) | - | Start ordering. See R10. |
task | block | - | Exactly one in v1. |
build | block | - | Build from source instead of pulling. |
network | block | - | Ports and ingress policy. |
expose | block | - | Public exposure and middleware. |
health_check | block | - | Zero or more, each labelled. |
volume | block | - | Zero or more mounts. |
scaling | block | - | Autoscaling policy. |
update | block | rolling | Rollout strategy. |
restart | block | always | Restart policy. |
task
| Field | Type | Notes |
|---|---|---|
image | string | Optional only if a build block is present (R8). Digests are supported and preferred. |
command | list(string) | Overrides the image entrypoint. An argument array, never a shell string (R12). |
args | list(string) | Keeps the image's entrypoint and replaces its arguments (R12, v1.97). Beside command, the argv is command then args - the Kubernetes pair. args = [] is refused; omit the field instead. |
capabilities | list(string) | Grants added to the baseline set; "none" starts from nothing instead (R13). |
env | map | Values may be secret: references or ${service.…} interpolations. |
resources | block | cpu in MHz (1000 = one core), memory in MiB. An omitted limit is unbounded: all cores / all allocatable memory. pids overrides the per-alloc process cap (256 when omitted; the cap is always on). |
registry_auth_ref | string | secret: reference to a docker config.json used to pull the image. Project-scoped like every other reference (R5). |
pull_policy | string | Where the image may come from (R33): if-not-present (default), never, or always. Omit it to take the node's default. See below. |
read_only_rootfs | bool | Mounts the container's root filesystem read-only. Volumes and files still mount as declared, so a writable /tmp is a volume away. kanea plan warns when it is off. |
device | block | Requests a host device the operator granted, by name (R17). See below. |
socket | block | Requests a host unix socket the operator granted, by name (R18). See below. |
init
Setup that has to finish before the workload starts: a schema migration, a
chown of a freshly-created volume, a config render. A service
may declare any number of init blocks; they run in
declaration order, one at a time, to completion, and the task is
created only once the last has exited zero (R32).
Each step shares the alloc: the same network namespace (so
${service.…} and internal DNS resolve exactly as they do for
the task, which is what makes a wait-for-database step possible), the same
volumes at the same mount paths, and the same secrets. Everything else it
declares for itself, and nothing is inherited from task.
service "api" {
# The strong posture (R13): the chown happens in the root init step,
# so the task itself runs as 999 with no capabilities at all.
hardening = "restricted"
# Runs as root to fix a directory the task will own as 999.
init "fix-perms" {
image = "busybox:1.36"
command = ["chown", "-R", "999:999", "/data"]
capabilities = ["CAP_CHOWN"]
timeout = "1m"
}
init "migrate" {
image = "registry.example.com/shop/api-migrate:1.4"
command = ["/bin/migrate", "up"]
env = {
DATABASE_URL = "secret:shop/database-url"
DATABASE_HOST = "${service.postgres.host}" # a real dependency edge
}
timeout = "5m"
}
task "app" {
image = "registry.example.com/shop/api:1.4"
user { uid = 999 }
}
}
| Field | Type | Notes |
|---|---|---|
image | string | Required. Unlike a task an init container has no build block, so there is no pipeline to produce one. |
command | list(string) | Overrides the image entrypoint. An argument array, never a shell string (R12's rule). |
args | list(string) | Keeps the step image's entrypoint and replaces its arguments (R12, v1.97), exactly as on a task. |
env | map | secret: references and ${service.…} interpolations work exactly as they do on a task. |
resources | block | Its own. An omitted limit is unbounded (R11); nothing is inherited from the task. |
user | block | Its own numeric identity (R23). Independent of task.user on purpose: running as root to chown a directory the task owns as 999 is the canonical use. |
capabilities | list(string) | R13's baseline and permitted set, same rule as a task's. |
registry_auth_ref | string | This step's pull credential, project-scoped like every other reference (R5). |
pull_policy | string | R33's, minus always: there is one pinned-image field and it belongs to the task. |
timeout | string | Bounds this step. Omit it for no timeout. A floor, not a deadline: the kill lands within one reconcile pass. |
A sequence re-runs from the first step whenever the alloc running it is created - a deploy, a crash restart, a spec change - so init steps must be idempotent. That is the same standing requirement the netns, the secrets tree and the staged host binds carry: an alloc is built, never resumed.
It runs on the first alloc; every other alloc of the service
waits for it and then starts with no sequence of its own. That is what
makes a schema migration expressible: three replicas, one migration.
Without it a first deploy of count = 3 creates three allocs
in one pass and runs three copies of the migration at once, against the
same database.
The gate is the first alloc's record leaving init
at the same spec hash, so a deploy re-runs the sequence once and
the others wait again. A waiting alloc says what it is waiting for on
kanea ps. If the first alloc fails its sequence the
service does not start, and the wait names it: the alternative is
failing every replica for one migration's mistake.
The cost is per-alloc volumes. Local storage gives each alloc its
own directory, so a step that prepares this alloc's volume
- the classic chown - now prepares the first
alloc's and no other. kanea plan warns when a service with
count > 1 declares both init steps and a local volume.
Steps that work on shared state are unaffected.
env_group
An environment declared once and taken by the services that want
it. variables substitutes
${name} into text you write, so sharing an environment with
them still means writing every key in every service; a group is injected.
env_group "common" {
LOG_LEVEL = "info"
REGION = "eu-central-1"
}
env_group "db" {
DATABASE_HOST = "${service.postgres.host}"
DATABASE_URL = "secret:shop/database-url"
}
service "api" {
project = "shop"
env_from = ["common", "db"]
task "app" {
image = "registry.example.com/shop/api:1.4"
env = { LOG_LEVEL = "debug" } # the service's own env wins
}
}
Groups apply in the order env_from lists them, then the
task's own env on top. It is opt-in per service rather
than project-wide deliberately: environment is baked into a container, so
a shared value changing rolls every service that takes it, and a blast
radius that size should be something a spec states rather than something
a service inherits by living in a project.
Not once per spec, and the difference is visible. The
${service.…} namespace is scoped to a project, so the
db group above resolves to postgres.shop.kanea
for a service in shop and to a different address for one in
another project. It also means the dependency edge lands on the
taking service, which is what makes it start behind
postgres. And a secret: in a group is scoped
against every consumer: a group carrying shop's
credential is refused for a service in another project.
file
Content Kanea places in the container, instead of baking it into an image or mounting a host volume and putting the file there yourself.
service "web" {
file "nginx" {
path = "/etc/nginx/conf.d/app.conf"
source = "./nginx.conf" # read at parse, embedded in the record
}
file "pgpass" {
path = "/etc/app/pgpass"
mode = "0400"
content = "db:5432:app:${secret.shop[\"database-password\"]}"
}
}
content is an ordinary HCL string, so a whole config file
goes in a heredoc; this is the shape most files actually take.
file "app-config" {
path = "/etc/app/config.yaml"
content = <<-EOT
server:
addr: ":8080"
upstream: "${service.api.host}:${service.api.port.http}"
database:
dsn: "postgres://app:${secret.shop[\"database-password\"]}@db:5432/app"
log:
# a literal dollar-brace the app expands itself
format: "$${level} $${msg}"
EOT
}
Three things there are worth naming. ${service.…} and
${secret.…} interpolate in the same file, and it is the
second that decides the rest: because one line names a secret, the whole
file is written 0400 on a tmpfs of its own and the record
keeps the reference, never the value. Content of its own that looks
like interpolation is escaped $${…}, which is HCL's rule
rather than Kanea's. And file blocks are the one place a
<<-EOT body reads better than a source
path: a short config that belongs to the spec does not need a second file
beside it.
| Field | Type | Notes |
|---|---|---|
path | string | Required. Absolute, clean, ..-free. Refused under /dev, /proc, /sys, at /etc/resolv.conf, and where a volume or socket already mounts. |
content | string | Inline. ${service.…} and ${secret.…} interpolate; a literal dollar-brace is $${…}. |
source | string | A path beside the spec, read at parse and embedded. Relative only, no symlink in any component, mutually exclusive with content. Refused where the parser has no directory: the dashboard's spec editor and MCP, because that parse happens inside kanead as root and reading a path there would be an arbitrary file read. |
mode | string | Octal. Default 0644, or 0400 for a file interpolating a secret. An execute bit is refused: a file block delivers configuration, not a program. |
| - | - | Size. 64 KiB per file and 128 KiB per service, post-substitution (PRD §21). A Desired record is CDC-replicated and the change log holds a second copy of every value it ships, so bulk bytes here are paid for on every apply, forever. |
Init containers must be idempotent. That is a requirement, not advice: a half-run sequence is abandoned rather than resumed, and the sequence re-runs from the first step whenever the alloc running it is created - a deploy, a crash restart, a spec change - for the same reason the network namespace and the secrets are rebuilt.
There is deliberately no device, socket,
health_check, expose, scaling,
count, depends_on, network or
build field on an init block. An init container
is a step, not a service, and the absence is the refusal.
A step's output goes to its own log: kanea logs shop/api -c
migrate. Attempts append to one transcript per alloc, and each
attempt opens with a separator line naming the step and when it started,
so a tail says where the previous attempt ends. A step that runs and
exits non-zero, or outlives
its timeout, fails the alloc and spends the restart budget (R29), so a
broken migration stops after attempts rather than hammering a
database forever; a step that could not be pulled or created is
retried every pass without spending it, because nothing ran.
pull_policy
Where an image may come from (R33), on a task or on an
init block. Three values:
| Value | Meaning |
|---|---|
if-not-present | Use the node's content store, pull only what is absent. The default, and what every spec did before this field existed. |
never | Use the content store or fail the alloc, naming the policy. For air-gapped nodes and ones whose images are preloaded: a pull attempt there is a slow failure that blames a registry the node was never going to contact. |
always | Re-resolve the tag and, when the digest has moved, pin it and roll every replica through the update policy. Task only. |
always does not mean re-pull at every container
creation. That would let two replicas of one spec run different bytes,
which is exactly what “a deploy is a spec-hash mismatch” exists
to prevent. It lowers to update { auto = true }, so the tag is
polled, a moved digest is pinned, and every replica rolls together through
max_parallel, min_healthy and the health check.
That also means it inherits auto-update's refusals: it cannot sit beside a
build block, needs a tag rather than a digest, and cannot be
combined with update { auto = false }. The unattended cadence
is update { interval } (default 6h, floor 5m), but
re-applying the spec forces an immediate re-check: push new bytes
under the same tag, kanea apply, and the digest is
re-resolved within about a minute instead of waiting out the interval.
Omitting pull_policy takes the node's default, which an
operator sets in /etc/kanea/kanea.hcl:
images {
pull_policy = "never" # or --image-pull-policy, which wins
}
Files are mounted read-only with nosuid,noexec,nodev,
and after volumes - so a file at a path inside a volume wins, which
is what "declare a file at a path" has to mean.
${secret.<scope>.<name>} resolves when the spec
is parsed to an opaque placeholder. The stored record keeps the
reference, exactly as it does for an environment variable, and
the value is substituted on the node when the container is created - so
a rotation lands at the next replacement, and the credential is not in
the state database, in a backup archive, or in
GET /v1/services. A file carrying one is written
0400 onto a tmpfs of its own, owned by the workload's user.
Use the bracket form for any name that is not a bare identifier, which
is most of them: ${secret.shop["database-password"]}.
source can be read
From the directory beside the spec for kanea run, and out
of the commit for a synced repository. The dashboard's spec editor and
the MCP spec tools parse text rather than files - inside the
daemon - so they refuse source by name and want
content. That refusal is the feature: a parser that opened
files there would be reading the node's filesystem as root on behalf of
whoever is signed in.
Content is capped at 64 KiB per file and 128 KiB per service: a service record is replicated in full on every deploy.
device and socket
A GPU for transcoding, a USB dongle, or the container runtime's socket for
a watchtower-style updater. Both blocks name a grant and never a
path: the operator defines grants on the node, and the spec asks for one
by name. There is no field to write a device path into, which is the
property the whole model rests on: a spec cannot request
/dev/mem because there is nowhere to say it.
service "jellyfin" {
task "app" {
image = "jellyfin:10.9"
device "dri" {
grant = "gpu" # defined by the operator, not here
}
}
}
service "watchtower" {
task "app" {
image = "watchtower:1.7"
socket "runtime" {
grant = "containerd"
mount_path = "/var/run/docker.sock"
}
}
}
| Field | Type | Notes |
|---|---|---|
grant | string | Required, on both blocks. Names a device/socket grant in the node's server config, /etc/kanea/kanea.hcl. |
mount_path | string | socket only, required. Where the socket appears in the container. A device appears at its host path. |
read_only | bool | socket only. Restricts the filesystem entry, not the protocol spoken over it. |
The operator side lives in the node's server config,
/etc/kanea/kanea.hcl (full reference:
Node configuration → device and socket),
probed automatically at daemon start:
no flag, no unit editing, and a kanea init re-run never
touches it. The default is that the file does not exist: a spec asking
for a grant on a node with none fails its alloc rather than starting
without it. Grants name the projects that may claim them:
# /etc/kanea/kanea.hcl; the node's, never the repository's
bind { # the API/dashboard listener (v1.61)
api_addr = "192.168.1.10:8600"
api_tls = "self-signed" # acme | self-signed | provided | plaintext
}
variables { # node-wide spec-variable defaults (v1.63), never secrets
domain = "home.lan"
}
device "gpu" {
nodes = ["/dev/dri/renderD128"]
allow = ["media"]
}
socket "containerd" {
path = "/run/kanea/containerd.sock"
allow = ["ops"]
}
Setting it up, end to end:
- Create the file (init already made the directory):
sudo install -o root -g root -m 0644 /dev/null /etc/kanea/kanea.hcl, then write the grant blocks above, and, if you use host volumes, thestorageallowlist stanza in the same file. It must stay root-owned and writable only by its owner; kanead refuses a group- or world-writable policy file. sudo systemctl restart kanead: the file is read once, at startup, never polled. Running workloads and ingress both survive the restart; see applying a change.- Verify: the startup log carries
server config loadedwith the path, and awarnline naming the passthrough consequence when grants are present.kanea psshows the alloc converging.
A socket grant is root on this node for whoever holds it. A container with the runtime socket can start other containers without the hardening defaults. There is no containment story and none is claimed: the control is that granting it happens on the machine, in a file no spec author can write, and names one project.
network
network {
port "http" { container = 8096 }
publish "http" {
host = 8096 # http://<node>:8096
mode = "http" # "http" (default) | "tcp"
ip_restriction { allow = ["192.168.0.0/16"] }
}
policy { allow_from = ["analytics/collector"] }
}
port "<name>":containeris the port inside the container. The name is whatexposeandhealth_checkrefer to, and what${service.x.port.<name>}resolves.publish "<port name>": a node port the edge binds for this service, with or without a domain. The label names theportabove; there is deliberately no field for a container port number, so a published port can never forward somewhere the service did not declare. See R21.mode = "tcp"relays bytes for Postgres, a game server, or anything that is not HTTP. It keepsip_restrictionand nothing else: arate_limitorheadersblock on one is a plan error rather than a control that is silently dropped. On a tcp listener the upstream sees the edge's address, not the client's, soip_restrictionis the whole mitigation and it is enforced at accept time.mode = "udp"relays datagrams (game servers, DNS, syslog): one session per client source address, pinned to a backend by rendezvous hash so a conversation survives an edge restart. Because one spoofed datagram can open a session, two bounds assume forgery: at most 8 live sessions per source IP, and at most 4 KiB back to the client until it sends a second datagram - the proof its address is real. Both refusals are counted inkanea_edge_udp_refused_total.- Which node ports a spec may claim belongs to the node, not the spec
(
--publish-ports, unprivileged by default). A repository anyone can push to must not be able to take :22. See R22. policy.allow_from: fully-qualified"<project>/<service>"peers permitted to reach this service. See R14.
expose
North-south exposure: the edge proxy, TLS, and the middleware chain. The block may
repeat: each expose is one complete route with its own domains, port,
TLS and middleware, so one service can serve its UI and its API on different names
and ports. Only the first block may omit domains, and blocks that
declare auth must declare the same auth.
| Field | Type | Notes |
|---|---|---|
domains | list(string) | Omitted → <service>.<project>.<base_domain>. No two services may claim the same domain, counting generated ones. |
port | string | Which declared network { port } the route proxies to. Omitted → the port named http, or the sole declared port. Explicit beats every convention; naming an undeclared or udp port is a plan error. |
tls | block | mode: acme, self-signed, provided or plaintext; name selects one of the node's provided grants. Omit the block and the node's --tls-default decides. A mode names a source, never a path. |
ip_restriction | block | allow, deny: lists of CIDRs. Empty allow means the world; deny wins. |
rate_limit | block | requests, window, per, burst. per is ip, service, or header:<name>. |
headers | block | request_set, request_remove, response_set, response_remove. |
The chain evaluates in a fixed order regardless of declaration order:
Host match → IP allow/deny → rate limit → header transforms → upstream
The headers block rejects any attempt to set or remove the
X-Forwarded-* set, or the hop-by-hop headers. Those carry the client
identity that IP restriction, rate limiting and the audit log are all keyed on: a
spec able to rewrite X-Forwarded-For would be forging the thing every
other control trusts.
health_check
health_check "http" {
type = "http" # http | tcp | exec
path = "/healthz" # http only
port = "http" # port name, required for http and tcp
interval = "10s"
timeout = "2s"
failures = 3
}
An exec check runs inside the task's container and takes
command as an argument array, never a shell string: the same rule
as task.command, for the same reason.
Alloc health is only ever written by a probe. A service without a
health_check block has every alloc unhealthy forever, and that is
correct: Kanea asks whether a check is configured before it asks whether one passed.
It matters because min_healthy in the update block is
meaningless without one.
volume
volume "data" {
storage = "local-ssd" # a declared storage block
mount_path = "/var/lib/data"
read_only = false
size = "10GiB" # a budget, not a quota: see below
}
Mount failures fail the alloc loudly. A volume that silently is not there is worse than one that stops the deploy.
size declares a budget (R31). Kanea measures what the volume
actually uses on a slow background schedule, shows it in
kanea volume list, and emits volume.over_budget when it
is crossed, and volume.under_budget when it comes back. It is
deliberately not a quota and nothing enforces it: no filesystem quota
mechanism exists on the node, and nfs, smb and
s3 could not carry one if it did. Declaring it on an s3
volume is a plan error, because that driver is never walked: one request per
directory, on a schedule, is not a measurement worth taking. Changing a
size never redeploys the service: a budget is not baked into a
container.
scaling
scaling {
min = 2
max = 10
metric "cpu" { target = 70 } # percent of resources.cpu
metric "rps" { target = 500 } # requests per second, per alloc
metric "p95_latency_ms" { target = 800 }
cooldown = "2m"
}
cpucomes from containerd cgroup metrics and is a percentage of the alloc's declared limit: the number the workload can actually use.rpsandp95_latency_mscome from the edge for exposed services, and from the datapath's own map counters for east-west traffic.countmust fall inside[min, max], otherwise the autoscaler would immediately contradict the spec.
update and restart
update {
strategy = "rolling" # rolling | replace
max_parallel = 1
min_healthy = "30s"
auto = true # follow task.image's tag (R19); off by default
interval = "6h" # how often to re-resolve it; minimum 5m
deadline = "10m" # how long a new digest has to prove itself
}
restart {
attempts = 5
backoff = "10s,30s,1m,5m"
}
max_parallel bounds allocs that are down, not replacements in
flight: anything already unavailable spends the budget first, so a failing deploy
halts rather than marching through every replica. min_healthy applies only
to allocs this deploy has already replaced.
restart bounds crash restarts: each one waits the next
backoff entry (the last repeats), and an alloc that exhausts
attempts is failed and left alone until a deploy or
kanea restart - either resets the count, because the budget belongs
to the spec hash that spent it. The budget counts only crashes the daemon
watched (v0.31.0): an alloc found already stopped at startup - a node reboot, a
power loss - is recovered immediately without spending an attempt, however many the
counter shows, and without waiting out a backoff armed before the outage. An alloc
already marked failed stays failed across a reboot: that verdict came from
crashes the platform saw, and the fix is the same deploy or restart as ever.
auto is the watchtower feature, as policy rather than as a container
holding your runtime socket. Kanea re-resolves the tag the service already
declares, and when the digest behind it moves it pins the new one: the tag stays in
the spec, the digest is what runs. That is an ordinary deploy nobody typed: it rolls
with max_parallel, waits min_healthy and is gated by the
service's health check, because it goes through the same machinery every other deploy
does. If the new digest does not converge within deadline, the previous
one is re-pinned and the service goes back to what worked.
It is refused on a service whose image is already a digest (nothing to
follow) and on one with a build block (the pipeline owns that image).
Private registries need task.registry_auth_ref. Both outcomes emit
events: image.updated and image.update_failed.
build
build {
context = "./web"
dockerfile = "Containerfile" # optional; auto-detected when omitted
target = "registry.example.com/shop/web" # optional; omitted, the node's
# internal registry fills in
# <registry>/shop/web
tag = "${GIT_SHA_SHORT}"
cache_repo = "registry.example.com/shop/web-cache"
registry_auth_ref = "secret:shop/registry"
}
- Builds run on rootless BuildKit, unprivileged end to end.
- An omitted
targetmeans the node's internal registry. Every node runs a loopback-only build registry (--registry, default127.0.0.1:5100;offdisables it, and--buildkit offimplies that), and the target is composed on the node as<registry>/<project>/<service>, sobuild { context = "." }is a complete build configuration: no external registry, no push credential. Pushes to it are credential-gated per boot; pulls are loopback plain HTTP, which containerd already speaks. An explicit target behaves exactly as before, andcache_repomay not point at the internal registry. ContainerfileandDockerfileboth work, andContainerfilewins when both exist.contextanddockerfileare paths inside the checkout: absolute and..forms are refused, and the context must be a real directory, never a symlink.targetandcache_repomust parse as image references, andtagmay not contain a comma, an=or whitespace - buildctl reads those flags as comma-separated option lists, so a comma is option injection, not a name.- The push credential is materialised as a
config.jsonfor the build and never enters the build context. - A service with a
buildblock and notask.imageis legitimate: the reconciler skips it until the first build pins a digest. - Builds are serialised, and refused when full rather than queued forever.
Isolation is collective, so a second concurrent build would share the first's budget.
queuedis a real state and shutdown cancels what is still waiting.
Interpolation
Three kinds of ${…} appear in a spec, and they resolve at different times.
| Form | Resolved | Notes |
|---|---|---|
${VAR} | Parse time | From the spec's variables block, the node's variables stanza in /etc/kanea/kanea.hcl, and built-ins such as KANEA_PROJECT. Node under spec under caller (R30). |
${GIT_SHA_SHORT} | Checkout time | Survives parsing as a literal reference: its value does not exist until a commit is checked out. |
${service.<n>.host} | Alloc start | Becomes <n>.<project>.kanea. Validated at plan time; resolved as a DNS name, never an IP, so load-balancer reprogramming cannot break it. |
${service.<n>.port.<p>} | Alloc start | The named port's frontend port. |
secret:<path> | Alloc start | Not interpolation: a reference the reconciler resolves. Project-scoped (R5). |
Removing a service from a spec
Deleting a service block and re-applying does not stop the
service. An apply replaces the desired state of the services it names and leaves
everything else alone, so that kanea run web.hcl can never delete what
db.hcl declares. Removing a service is therefore a deliberate act, and
there are two ways to perform it.
| One service | Everything a spec dropped | |
|---|---|---|
| Command | kanea remove shop/web | kanea run --remove-orphans app.hcl |
| Scope | the service you name | the spec's project blocks |
| Dry run | - | kanea plan --remove-orphans app.hcl |
--remove-orphans makes the spec authoritative for the projects it
declares a project block for: anything stored in one of them that the
file no longer declares is deleted, in the same atomic batch as the applies. Projects
the spec does not mention are never touched, so a spec for shop cannot
reach into data.
$ kanea plan --remove-orphans app.hcl
- destroy shop/legacy-worker (count 1, image legacy:v3)
volumes - scratch local /scratch rw
(the mount goes; the volume's data is NOT deleted)
~ update shop/web
image web:v1 -> web:v2 (rolls allocs)
Plan: 2 change(s) - 1 update, 1 destroy; 1 replace running allocs.
Run `kanea run` to apply.
$ kanea run --remove-orphans app.hcl
[the same block]
Apply? [Y/n]
applied shop/web
removed shop/legacy-worker
1 service(s) removed. Volume data was not deleted.
--image
kanea run app.hcl shop/web --remove-orphans is refused by name. A
selector sends part of the spec, so the apply cannot claim to be the whole
of any project - and if it did, it would read as "delete everything in
shop except web". --image is refused for the
same reason from the other end: it declares one service and no project at all.
Note that flags must come before file arguments:
kanea run --remove-orphans app.hcl, not the other way round.
What a prune destroys, and what survives
| Goes | Stays |
|---|---|
The containers, and the alloc records behind kanea ps | Volume data, local and network |
| The service's virtual IP (released, and another service may take it) | Certificates, including issued ACME ones |
| Its routes, and its network volume mounts | Secrets under the project |
| The restart generation and any pinned image digest | Its internal datapath id, which is keyed by name |
Volume data is never deleted, which cuts both ways and is why the command says
so every time. A service pruned by mistake comes back with its data by re-applying it,
because the directory is derived from project/service/index and nothing
removes it. And a prune frees no disk: reclaiming it is a manual
rm under the volume directory, deliberately, since Kanea has never
deleted a volume it did not create and R15 says as much for host paths.
Each removal emits service.removed and is named in the audit log, so a
prune is not a quiet operation even when nobody is watching the terminal.
A GitOps sync is unaffected: a synced repository still applies additively, and removing a service from a synced spec leaves it running until someone prunes or stops it.
Validation rules
The rules a spec is checked against, all enforced at parse or plan time with file/line/column diagnostics. They are numbered because the error messages cite them, and the numbering is the PRD's; the cards below cover the ones most people meet, not every rule the parser knows.
Names are DNS-1123 labels. Project and service names: lowercase alphanumeric and -, alphanumeric at both ends, ≤ 63 characters. Parse errors abort the run.
Variables. Flat ${VAR} interpolation in any attribute, from the spec's variables block, the node's variables stanza and built-ins: node under spec under caller. Reserved names (the built-ins, service) cannot be declared, values are primitives, and a definition never references a sibling.
Secrets are referenced, never inlined. Resolved at alloc start. Primary injection is a tmpfs file at /run/kanea/secrets/<alloc>/<name>; env vars are supported and documented as weaker.
kanea plan is a real dry run and shows the create/change/destroy diff before anything is applied.
Secret references are project-scoped. A service may name secret:<own-project>/… or secret:shared/… and nothing else. Cross-project references are rejected. Git, registry, storage and notification credentials follow the same scoping.
spec_version = 1 is declared by every file; future revisions are gated on it.
Health check types are http, tcp and exec. exec runs inside the container and takes an argument array, never a shell string.
The minimal service is image-only. At least one of task.image or build must be present. When both are, the pipeline-built digest wins and task.image is the pre-first-build value.
Service references are same-project in v1, validated at plan against the whole applied set: the referenced service and port must exist, and file order is irrelevant. Resolved as DNS names, never IPs. Cycles are rejected, with the cycle shown in the diagnostic.
Dependencies gate starts, not stops. depends_on and every implicit reference edge mean a dependent will not start until its dependencies are healthy. If a dependency degrades afterwards, dependents keep running and events are emitted: no cascading stops.
A declared limit is enforced; an omitted one is the node's capacity. Omit resources and the alloc gets all cores and all allocatable memory, bounded by the workload parent cgroup, not by a per-alloc number nobody typed. A default pids.max applies regardless; resources.pids declares a service's own cap (256 when omitted), and changes it like any other limit. A declared memory.max breach OOM-kills the alloc, emits an event, and the restart policy applies. Scaling on cpu or memory (percent-of-limit) requires the corresponding limit and is refused at plan without it. Declared values are the admission units counted against the node's workload budget. Functions keep small defaults (cpu = 100, memory = 64): the wasm sandbox's caps are promises.
task.command is an argument array. The first element must be non-empty; later ones may be empty, because some programs use that meaningfully; redis-server --save "" is how you disable snapshots. task.args keeps the image's entrypoint and replaces its arguments (v1.97): alone, argv is the image's ENTRYPOINT then args, resolved on the node from the image itself; beside command, argv is command then args - the Kubernetes pair. Its elements may all be empty (they are arguments, not a program), but args = [] is refused: the record cannot carry "declared empty" apart from "absent", and command is the spelling for "run the bare entrypoint". Both apply to init blocks too; a function takes neither.
task.capabilities adds to a baseline, and "none" takes the baseline away. Every runc alloc starts from the baseline set; CHOWN, DAC_OVERRIDE, FOWNER, FSETID, KILL, SETGID, SETUID, so PUID-style images start without a capability line. Binding :80 needs no capability at all: every alloc netns sets ip_unprivileged_port_start=0, which is why NET_BIND_SERVICE is no longer baseline. The effective grant is the union of baseline and declared list; ["none"] is the full drop-ALL posture, ["none", "CAP_NET_RAW"] exactly one grant. Declarable beyond the baseline: SETPCAP, SETFCAP, NET_BIND_SERVICE, NET_RAW (never baseline; the datapath's identity is the IP), SYS_CHROOT, MKNOD, AUDIT_WRITE. One honest caveat on SETFCAP: file capabilities are xattrs, so a binary given them inside an rw host volume keeps them on the host for whoever executes it later; no-new-privileges contains the grant inside the container, not the file. Privilege-equivalent ones; SYS_ADMIN, SYS_MODULE, SYS_PTRACE, BPF, PERFMON and friends; are rejected at parse time: granting them would be the privileged escape hatch v1 deliberately does not have. Effective capabilities go into the bounding, effective and permitted sets, never inheritable or ambient. hardening = "restricted" on the service names the strong posture: a non-zero user required, drop-ALL enforced, any declared grant refused - and init steps unchanged, because the canonical shape is a root init step that chowns and exits before a restricted task.
allow_from only ever adds reachability. Each entry is a fully-qualified "<project>/<service>"; the datapath's ingress rules only ever union, so an entry can never weaken the project default-deny. There is no wildcard: "analytics/*" is a parse error, because naming the peer is the point. Same-project entries are accepted and redundant.
host volumes are operator-gated. The path is validated as absolute, clean and ..-free at parse time, but whether it may be mounted is not the spec's decision. kanead refuses any path outside storage.allowed_host_paths in the server config /etc/kanea/kanea.hcl (or --allowed-host-paths, which wins when set), whose default is empty. The check is applied after symlink resolution, and the directory must already exist - unless the storage block sets create = true, which makes Kanea create it. The default has not moved: without that flag a missing path is still refused, because creating one on demand turns a typo into a volume that is silently empty. Creation is allowlist-gated first - the nearest existing parent is resolved and checked before anything is written - so a create outside a permitted prefix leaves no directory behind. Creating a directory is not owning it: Kanea still never chowns a host volume.
expose fails closed. A service may only be exposed if it declares a port, and the upstream port must be unambiguous; declared with port = "<name>", or named http, or the only one declared. The block may repeat, each one a complete route validated independently; only the first may omit domains, and blocks that declare auth must agree. Every domain is validated as a hostname, and no two services may claim the same one, counting generated FQDNs. Middleware is checked here too: CIDRs must parse, rate_limit needs a positive requests and a valid window, and headers may not touch the hop-by-hop or X-Forwarded-* sets.
task.device names a grant, not a device. There is no field for a device path, so a spec cannot ask for one. Parse time checks only that the grant name is a DNS-1123 label. The node refuses a grant it does not have, a grant whose allow list does not name the requesting project, and a path that is no longer a character or block device: checked after symlink resolution, at every alloc start. The device appears at its host path, and the grant carries the cgroup permissions (rw by default, never m unless written). A failed grant fails the alloc; it never starts without what it asked for.
task.socket is R17 for unix sockets, and is privilege delegation. mount_path is validated as absolute, clean and ..-free, may not sit under /dev, /proc or /sys, and may not collide with another socket or a volume. The bind carries nosuid, noexec and nodev. None of that makes it safe and none of it is meant to: a container holding the container runtime's socket can create containers without the hardening defaults, so the grant is equivalent to root on the node. The server config is the only control over it, which is why it is project-scoped and empty by default.
init blocks run to completion before the task. In declaration order, one at a time; the task is created only once the last has exited zero. Each shares the alloc's network namespace, volumes and secrets, and declares its own image, command, args, env, resources, user, capabilities and timeout - nothing is inherited from task. There is no device, socket, health_check, expose, scaling, count, depends_on, network or build field: an init container is a step, not a service. Names are DNS-1123 labels, unique within a service, and compose into a container id and a log file. A sequence runs once per service, not once per alloc (v1.92): on alloc 0, with every other alloc waiting on that alloc's record leaving init at the same spec hash, which is what makes "three replicas, one migration" expressible. A local volume is one directory per alloc, so a step preparing this alloc's own volume prepares the first's alone; kanea plan warns on that shape. They must be idempotent: a half-run sequence is abandoned rather than resumed, and it re-runs whenever the alloc running it is created. A step that ran and failed, or outlived its timeout, spends the restart budget (R29); one that could not be pulled or created is retried every pass without spending it, because nothing ran.
pull_policy says where an image may come from. if-not-present (the default), never (content store only - the air-gapped case) or always. always is not a per-create re-pull, which would let two replicas of one spec run different bytes: it lowers to update { auto = true }, so a moved digest is pinned and every replica rolls together, and it inherits that rule's refusals. It is refused on an init block, because there is one pinned-image field and it belongs to the task. The node supplies the default through /etc/kanea/kanea.hcl's images stanza; an omitted policy means "the node decides".
An env_group is declared once and taken by a service. A top-level block names environment variables; a service opts in with env_from = ["common", "db"], and the groups apply in that order with the task's own env on top. Opt-in per service rather than project-wide, because environment is baked into a container and a shared value changing rolls every service that takes it. A group is evaluated once per consuming service: ${service.…} is project-scoped, so one group taken from two projects resolves differently, and the dependency edge lands on the service that took it. A secret: inside a group is scoped against every consumer, so a group carrying one project's credential is refused for a service in another.
A file block puts content in a container, and a secret in it stays a reference. path is absolute, clean and ..-free, refused where a volume or socket already mounts and under the system paths a volume destination is; the bind is read-only with nosuid,noexec,nodev, and an execute bit in mode is refused. Files mount after volumes, so one inside a volume's path wins. Exactly one of content and source; source is read where the spec is parsed and is refused where there is no directory (the dashboard editor, MCP), because a parser reading files there would be reading the node's filesystem as root. ${secret.<scope>.<name>} resolves to an opaque placeholder at parse and is substituted on the node at container create, so the value never enters the state database, a backup, or the API; such a file is written 0400 on a tmpfs of its own. Content is SpecHash material - editing a config file rolls the service, which is the feature - and is capped at 64 KiB per file and 128 KiB per service.
update.auto follows the tag the service declares. Off by default. Kanea re-resolves task.image's tag every interval (default 6h, minimum 5m) and pins the digest behind it when it moves; the declared tag is never overwritten, because it is what the next poll re-reads. The pinned digest is server-owned and survives kanea apply: except when you edit image or turn auto off, which both hand authority back to the spec. An apply does re-arm the poll, though: the tag is re-resolved within about a minute of one, so a re-pushed tag lands with a push and an apply rather than an interval later. A failed update reverts to the digest that was running if the new one has not converged within deadline: converged means healthy where a check block exists, and running without crash-looping where it does not. Refused on a digest-pinned image and on a service with a build block.
Full example
Everything above, in one file. This parses and validates against Kanea's own parser.
# shop.hcl: everything for one project spec_version = 1 project "shop" { description = "E-commerce storefront stack" git { url = "https://github.com/example/shop-deploy.git" branch = "main" path = ".kanea/" auth_ref = "secret:shop/github-deploy-key" } notifications { slack { url_ref = "secret:shop/slack-webhook" } on = ["deploy.failed", "service.unhealthy", "scale.*"] severity = "warning" } } storage "local-ssd" { type = "local" } service "postgres" { project = "shop" task "db" { image = "postgres:16-alpine" env = { POSTGRES_PASSWORD = "secret:shop/postgres-password" } resources { cpu = 1000 memory = 1024 } } network { port "pg" { container = 5432 } } volume "data" { storage = "local-ssd" mount_path = "/var/lib/postgresql/data" } health_check "tcp" { type = "tcp" port = "pg" interval = "10s" } } service "web" { project = "shop" description = "Storefront frontend" count = 3 depends_on = ["postgres"] build { # No target: the node's internal registry builds, pushes and deploys # this service without any external registry. context = "./web" } task "app" { image = "registry.example.com/shop/web:latest" env = { NODE_ENV = "production" DATABASE_URL = "secret:shop/database-url" DATABASE_HOST = "${service.postgres.host}" DATABASE_PORT = "${service.postgres.port.pg}" } resources { cpu = 500 memory = 512 } } network { port "http" { container = 3000 } } expose { domains = ["shop.example.com", "www.shop.example.com"] tls { mode = "acme" } ip_restriction { deny = ["198.51.100.7/32"] } rate_limit { requests = 100 window = "1m" per = "ip" burst = 20 } headers { response_set = { Strict-Transport-Security = "max-age=63072000; includeSubDomains" } response_remove = ["Server", "X-Powered-By"] } } health_check "http" { type = "http" path = "/healthz" port = "http" interval = "10s" timeout = "2s" failures = 3 } scaling { min = 2 max = 10 metric "cpu" { target = 70 } metric "rps" { target = 500 } cooldown = "2m" } update { strategy = "rolling" max_parallel = 1 min_healthy = "30s" } restart { attempts = 5 backoff = "10s,30s,1m,5m" } }
Sample stacks
Five complete stacks, from a static site to a Kafka cluster. Every one parses and
validates against Kanea's own parser: copy the file, change the names and domains,
and start with kanea plan, which will tell you about anything the copy
broke before anything runs.
A static site
The smallest production-shaped thing: two replicas behind Let's Encrypt, with a
health check. The check is not decoration: it is what a rolling deploy waits on
(min_healthy has nothing to measure without one), what anything
declaring depends_on this service waits for, and the difference
between a replica that stops answering showing up in kanea status
and staying a mystery.
spec_version = 1 project "web" {} service "site" { project = "web" count = 2 task "nginx" { image = "nginx:1.27-alpine" resources { cpu = 200 memory = 128 } } network { port "http" { container = 80 } } expose { domains = ["example.com", "www.example.com"] tls { mode = "acme" } } health_check "http" { type = "http" path = "/" port = "http" interval = "10s" timeout = "2s" failures = 3 } }
This works the moment example.com resolves to the node and ports 80/443
reach it: the ACME HTTP-01 flow needs nothing else. From here the natural next step
is replacing the stock image with your own: add a build
block and the pipeline builds and pins a digest on every push.
kanea plan site.hcl kanea run site.hcl --wait=60s kanea status web/site
Two apps and a database
A frontend and an API sharing PostgreSQL and Redis. This is the shape most
multi-service deployments take, and it exercises the machinery that matters:
depends_on gates the apps until their backends are healthy
(which is why postgres and redis have checks) and the
${service.…} references resolve to internal DNS names at alloc start,
so nothing here hard-codes an address. One secret, referenced from two services,
never written in the file.
spec_version = 1 project "paste" { description = "Pastebin: frontend, API, database, cache" } storage "db-data" { type = "local" } service "postgres" { project = "paste" task "db" { image = "postgres:16-alpine" env = { POSTGRES_DB = "paste" POSTGRES_USER = "paste" POSTGRES_PASSWORD = "secret:paste/db-password" } resources { cpu = 1000 memory = 1024 } } network { port "pg" { container = 5432 } } volume "data" { storage = "db-data" mount_path = "/var/lib/postgresql/data" } health_check "up" { type = "tcp" port = "pg" interval = "10s" } } service "redis" { project = "paste" task "cache" { image = "redis:7-alpine" command = ["redis-server", "--save", ""] resources { cpu = 250 memory = 256 } } network { port "redis" { container = 6379 } } health_check "up" { type = "tcp" port = "redis" interval = "10s" } } service "api" { project = "paste" count = 2 depends_on = ["postgres", "redis"] task "app" { image = "registry.example.com/paste/api:1.4.2" env = { DB_HOST = "${service.postgres.host}" DB_PORT = "${service.postgres.port.pg}" DB_USER = "paste" DB_PASSWORD = "secret:paste/db-password" REDIS_ADDR = "${service.redis.host}:${service.redis.port.redis}" } resources { cpu = 500 memory = 512 } } network { port "http" { container = 8080 } } expose { domains = ["api.paste.example.com"] tls { mode = "acme" } rate_limit { requests = 300 window = "1m" per = "ip" burst = 50 } } health_check "http" { type = "http" path = "/healthz" port = "http" interval = "10s" timeout = "2s" failures = 3 } update { strategy = "rolling" max_parallel = 1 min_healthy = "30s" } } service "web" { project = "paste" count = 2 depends_on = ["api"] task "app" { image = "registry.example.com/paste/web:1.4.2" env = { API_URL = "http://${service.api.host}:${service.api.port.http}" } resources { cpu = 250 memory = 256 } } network { port "http" { container = 3000 } } expose { domains = ["paste.example.com"] tls { mode = "acme" } } health_check "http" { type = "http" path = "/" port = "http" interval = "10s" timeout = "2s" failures = 3 } }
Create the secret first, then deploy; order within the file never matters:
kanea secret put paste/db-password # value on stdin kanea run paste.hcl --wait=90s
Details worth stealing: redis-server --save "" is a command as an
argument array with a meaningfully empty argument (R12; a shell string could not
say that); the rate limit lives only on the API route, because the frontend serving
assets at API rates would be rate-limiting your own pages; and postgres and redis
declare no expose and no publish, so they are reachable
from this project's services and from nothing else on the network.
Jellyfin with local media
A media server is the stack where the node itself gets a say: the library is a
directory the operator already owns, and hardware transcoding needs a GPU device.
Both cross the line a job spec cannot cross alone: a host volume does
nothing until its path is allowlisted on the node (R15), and the device
block names a grant, never a path (R17).
spec_version = 1 project "media" {} storage "config" { type = "local" } storage "library" { type = "host" path = "/srv/media" } service "jellyfin" { project = "media" task "app" { image = "jellyfin/jellyfin:10.9.11" device "dri" { grant = "gpu" # hardware transcoding; the grant is defined on the node } resources { cpu = 4000 memory = 4096 } } network { port "http" { container = 8096 } publish "http" { host = 8096 # http://<node>:8096, LAN only ip_restriction { allow = ["192.168.0.0/16"] } } } volume "config" { storage = "config" mount_path = "/config" } volume "media" { storage = "library" mount_path = "/media" read_only = true } health_check "http" { type = "http" path = "/health" port = "http" interval = "15s" timeout = "5s" failures = 3 } }
The node's half, in the server config: without it the spec is valid and the alloc fails, loudly:
# /etc/kanea/kanea.hcl; the node's, never the repository's
storage {
allowed_host_paths = ["/srv/media"]
}
device "gpu" {
nodes = ["/dev/dri/renderD128"]
allow = ["media"]
}
A failed grant fails the alloc rather than starting without the GPU, because a
transcoder silently falling back to software looks healthy and does the wrong thing.
The library is mounted read_only (a media server has no business
writing to it) while /config is an ordinary local volume Kanea
manages. The publish block binds :8096 on the node for the LAN, with
the edge enforcing the CIDR allowlist; for access from outside, add an
expose block with a domain instead of
widening the CIDR.
SMB and S3 volumes
The same volume block, backed by things that are not on the node: a
file browser over a NAS share and a read-only bucket. The credential for either
driver is one secret whose value is <user>:<secret>
(username and password for SMB, access key and secret key for S3) resolved at mount
time into a 0600 file, never onto a command line. Omit
auth_ref entirely for a public bucket or an open share.
spec_version = 1 project "files" {} # A share on the NAS. The secret's value is "<username>:<password>". storage "nas" { type = "smb" server = "192.168.1.20" share = "documents" auth_ref = "secret:files/nas" } # A bucket, read-only. The secret's value is "<access-key>:<secret-key>". storage "archive" { type = "s3" bucket = "household-archive" endpoint = "https://minio.internal:9000" # omit for AWS S3 auth_ref = "secret:files/archive" mode = "ro" # mountpoint-s3; "rw" selects s3fs } service "filebrowser" { project = "files" task "app" { image = "filebrowser/filebrowser:v2.32.0" user { uid = 1000 gid = 1000 } resources { cpu = 500 memory = 256 } } network { port "http" { container = 80 } } expose { domains = ["files.example.com"] tls { mode = "acme" } } volume "documents" { storage = "nas" mount_path = "/srv/documents" } volume "archive" { storage = "archive" mount_path = "/srv/archive" read_only = true } health_check "http" { type = "http" path = "/health" port = "http" interval = "15s" timeout = "5s" failures = 3 } }
Create both secrets first:
kanea secret put files/nas # username:password on stdin kanea secret put files/archive # access-key:secret-key kanea run files.hcl
The user block does double duty here (R23, R24): the process runs as
1000:1000, and the volumes inherit that ownership; no
uid/gid on the volume needed. On a mounted filesystem
there is nothing to chown, so the ownership travels in the mount
options instead, which is exactly why it works on a share the NAS controls. The
inheritance is resolved at parse time, not on the node, so a spec means the same
thing everywhere. (host and nfs volumes are the
exception: those drivers cannot carry ownership, so declaring it on one is a plan
error, and inheritance skips them.)
mode on the S3 storage selects the driver ("ro" is
mountpoint-s3 and the default, "rw" is s3fs) and the
storage reference's warning applies in full: an object
store is not a filesystem, so keep it for bulk, read-mostly data. A custom
endpoint points at MinIO or any S3-compatible store; omit it for AWS.
An nfs export is the same shape with server and
export. For all of them, a mount that cannot be established fails the
alloc loudly rather than starting the service beside an empty directory, and a
mount that dies later is supervised and remounted.
A Kafka cluster (KRaft)
Three brokers, no ZooKeeper. Each broker is its own count = 1
service rather than one service with count = 3, and that is
the load-bearing decision: a Kafka broker has an identity (a node id and an
advertised name that must be stable across restarts) and replicas of one service
are deliberately interchangeable. Three services give each broker its own DNS name,
its own data volume, and its own line in the quorum.
spec_version = 1 project "kafka" { description = "Three-broker KRaft cluster, no ZooKeeper" } storage "kafka-data" { type = "local" } service "kafka-1" { project = "kafka" task "broker" { image = "apache/kafka:4.0.0" env = { KAFKA_NODE_ID = "1" KAFKA_PROCESS_ROLES = "broker,controller" KAFKA_LISTENERS = "PLAINTEXT://0.0.0.0:9092,CONTROLLER://0.0.0.0:9093" KAFKA_ADVERTISED_LISTENERS = "PLAINTEXT://kafka-1.kafka.kanea:9092" KAFKA_CONTROLLER_LISTENER_NAMES = "CONTROLLER" KAFKA_LISTENER_SECURITY_PROTOCOL_MAP = "PLAINTEXT:PLAINTEXT,CONTROLLER:PLAINTEXT" KAFKA_CONTROLLER_QUORUM_VOTERS = "1@kafka-1.kafka.kanea:9093,2@kafka-2.kafka.kanea:9093,3@kafka-3.kafka.kanea:9093" KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR = "3" KAFKA_LOG_DIRS = "/var/lib/kafka/data" CLUSTER_ID = "MkU3OEVBNTcwNTJENDM2Qg" } resources { cpu = 1000 memory = 2048 } } network { port "client" { container = 9092 } port "controller" { container = 9093 } } volume "data" { storage = "kafka-data" mount_path = "/var/lib/kafka/data" } health_check "up" { type = "tcp" port = "client" interval = "15s" timeout = "5s" failures = 3 } } service "kafka-2" { project = "kafka" task "broker" { image = "apache/kafka:4.0.0" env = { KAFKA_NODE_ID = "2" KAFKA_PROCESS_ROLES = "broker,controller" KAFKA_LISTENERS = "PLAINTEXT://0.0.0.0:9092,CONTROLLER://0.0.0.0:9093" KAFKA_ADVERTISED_LISTENERS = "PLAINTEXT://kafka-2.kafka.kanea:9092" KAFKA_CONTROLLER_LISTENER_NAMES = "CONTROLLER" KAFKA_LISTENER_SECURITY_PROTOCOL_MAP = "PLAINTEXT:PLAINTEXT,CONTROLLER:PLAINTEXT" KAFKA_CONTROLLER_QUORUM_VOTERS = "1@kafka-1.kafka.kanea:9093,2@kafka-2.kafka.kanea:9093,3@kafka-3.kafka.kanea:9093" KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR = "3" KAFKA_LOG_DIRS = "/var/lib/kafka/data" CLUSTER_ID = "MkU3OEVBNTcwNTJENDM2Qg" } resources { cpu = 1000 memory = 2048 } } network { port "client" { container = 9092 } port "controller" { container = 9093 } } volume "data" { storage = "kafka-data" mount_path = "/var/lib/kafka/data" } health_check "up" { type = "tcp" port = "client" interval = "15s" timeout = "5s" failures = 3 } } service "kafka-3" { project = "kafka" task "broker" { image = "apache/kafka:4.0.0" env = { KAFKA_NODE_ID = "3" KAFKA_PROCESS_ROLES = "broker,controller" KAFKA_LISTENERS = "PLAINTEXT://0.0.0.0:9092,CONTROLLER://0.0.0.0:9093" KAFKA_ADVERTISED_LISTENERS = "PLAINTEXT://kafka-3.kafka.kanea:9092" KAFKA_CONTROLLER_LISTENER_NAMES = "CONTROLLER" KAFKA_LISTENER_SECURITY_PROTOCOL_MAP = "PLAINTEXT:PLAINTEXT,CONTROLLER:PLAINTEXT" KAFKA_CONTROLLER_QUORUM_VOTERS = "1@kafka-1.kafka.kanea:9093,2@kafka-2.kafka.kanea:9093,3@kafka-3.kafka.kanea:9093" KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR = "3" KAFKA_LOG_DIRS = "/var/lib/kafka/data" CLUSTER_ID = "MkU3OEVBNTcwNTJENDM2Qg" } resources { cpu = 1000 memory = 2048 } } network { port "client" { container = 9092 } port "controller" { container = 9093 } } volume "data" { storage = "kafka-data" mount_path = "/var/lib/kafka/data" } health_check "up" { type = "tcp" port = "client" interval = "15s" timeout = "5s" failures = 3 } }
The quorum voters and advertised listeners are written as literal internal DNS
names (<service>.<project>.kanea), not
${service.…} references: deliberately. A reference is also a start
dependency (R10), and three brokers referencing each other is a cycle
kanea plan rejects (R9). Kafka's quorum is designed to form as
the peers come up, so the ordering edge is not wanted; the literal names are exactly
the strings the interpolation would have produced, minus the edge. Note the shared
storage "kafka-data" block is safe: each service's volume gets its own
directory beneath it, so the brokers never share a log dir.
A client in the same project bootstraps with all three names:
kafka-1.kafka.kanea:9092,kafka-2.kafka.kanea:9092,kafka-3.kafka.kanea:9092;
a service in another project also needs allow_from naming it on each
broker (R14). The cluster is deliberately not exposed north-south: Kafka's protocol
hands clients the advertised names to dial, and those resolve only inside the node;
publishing :9092 would let an external client connect once and then fail on
redirect. Generate your own CLUSTER_ID
(kafka-storage.sh random-uuid) rather than shipping the example's.
kanea run kafka.hcl --wait=120s kanea ps --project=kafka kanea logs kafka/kafka-1 --tail=50 # watch the quorum form
CLI reference
One binary, one command tree. The CLI talks to kanead over a unix socket
(--socket overrides it on every client command) and the daemon commands
are the ones systemd runs for you.
The socket is root-owned, so client commands run under sudo, or without
it, after joining the kanea group init creates and logging in
again: sudo usermod -aG kanea <user>. Membership is root-equivalent,
exactly like docker's group, and is never granted by Kanea itself.
Services are addressed as project/service throughout, or as a bare service
name with --project.
Every client command also takes a remote endpoint instead of the socket:
--url / $KANEA_URL, --token /
$KANEA_TOKEN and --ca-cert / $KANEA_CA_CERT.
That is how a laptop or a CI runner drives a node it is not sitting on; see
Working with a remote node for the setup and the pipeline
examples.
The CLI also installs through Homebrew, on Linux and on
macOS. kanea plan parses and validates a job spec with file-and-line
diagnostics before anything dials a socket, so authoring needs no daemon at all
- and since v1.82 the rest of the CLI is not tied to the socket either: with
KANEA_URL and a token it drives a remote node completely. A Mac is a
first-class client; it is just never the node, which is Linux. See
Working with a remote node.
Setup
kanea init
Interactive first install: preflight checks, configuration, the master-key ceremony,
the systemd units, and, since v0.5, the rest of the way to a working platform: it
asks for the dashboard's listen address (loopback by default; none keeps
the API socket-only), starts kanead, creates the first admin account over
the local socket, and ends with a summary of what it built; the dashboard URL, your
account, the internal DNS address and the subnet layout. Run it once, as root.
sudo kanea init [--data-dir=/var/lib/kanea] [--log-dir=/var/log/kanea/allocs]
[--unit-dir=/etc/systemd/system] [--prefix=/usr/local/lib/kanea]
[--network=ebpf|netns] [--containerd=…|external] [--buildkit=…|off]
[--node-cidr=10.244.0.0/24] [--cluster-cidr=10.244.0.0/16]
[--node-cidr6=…] [--cluster-cidr6=…] [--service-cidr6=…]
[--listen=127.0.0.1:8600|none] [--listen-cert=…] [--listen-key=…]
[--reserve=256M] [--admin-user=…] [--timeout=2m]
[--bundle=…] [--no-install] [--no-start] [--skip-checks] [--skip-units]
A full walk-through, with an annotated run and what a re-run keeps, is in
Installing → Setting up the node. Note there is no
--service-cidr here: the v4 service pool is fixed at
10.201.0.0/16 and is never rendered into the unit; change it by
editing kanead.service.
A non-loopback --listen is served over TLS, and since v0.23 init
provisions it: with no explicit --listen-cert/--listen-key
pair it mints a ten-year self-signed pair at /etc/kanea/api.crt and
api.key (once; a pair already there is left alone) and points the unit
at it. Bring your own with the flags, or take the listener over with the
bind stanza below. The one refusal left is an unspecified host
(0.0.0.0, :8600): a certificate's SAN needs a host to name.
--admin-user plus a piped password makes it scriptable;
--no-start writes the files and stops, which is the pre-v0.5 behaviour.
Re-running is safe: an existing master key and an existing account are left alone.
Since v1.61 the listener can live in the server config instead: a
bind { api_addr = … api_tls = … } stanza in
/etc/kanea/kanea.hcl, where api_tls is
acme, self-signed, provided or
plaintext; the same modes services use (the full field set is in
Exposing the API and dashboard). When it is
declared and --listen was not passed, init skips the listen question
and renders no listen flags into the unit: the file owns the listener, and moving
the API and dashboard later is an edit to the file plus
systemctl restart kanead, never a re-init.
The ceremony prints the master key and requires you to type it back; if that fails it is discarded and nothing is written. Without that key every backup this node ever makes is unreadable. Have somewhere to record it before you start.
kanea doctor
kanea doctor [--data-dir=…] [--prefix=…] [--network=ebpf|netns] [--containerd=…]
[--buildkit=…|off] [--node-cidr=…] [--cluster-cidr=…] [--service-cidr=…]
[--offline]
Verifies the node: dependencies and their versions against the pinned matrix, the containerd socket, bpffs and the cgroup2 mount, kernel version, cgroups v2 and slice placement, the effective memory floor, subnet overlap, the state database's permissions, the build socket, disk headroom and clock synchronisation. It also names known interference, like a host firewall dropping on the hooks alloc traffic crosses (docker, ufw, firewalld). Safe to run any time.
A failing check exits non-zero; warnings do not. A check this user may not
perform - the containerd socket, the bpf pin root and the state database are
all root-owned - is reported SKIP rather than FAIL,
and the run says once at the end how many need root. "I could not look" is not
"it is broken", and reporting it as one sends an operator to reinstall a healthy
node. --offline skips the one network probe, which is a reachability
check for the pinned artefacts rather than a requirement - on an air-gapped
node it is the expected answer, not a problem.
kanea firewall
kanea firewall [--manager=ufw|firewalld|nft|iptables] [--all]
[--prefix=…] [--node-cidr=…] [--cluster-cidr=…] [--network=ebpf|netns]
Prints the host-firewall rules this node's workloads need, derived from this node's
own cluster CIDR and resolver address rather than from an example. Two allowances,
because alloc traffic crosses two hooks a firewall owns: forward, on its way
off the node, and input, to reach the internal resolver - a query to the
resolver is a new inbound connection to the host on a veth. With no argument it
prints for the manager that appears to own the ruleset; --all prints
every one.
It prints and never applies. Kanea owns exactly the kanea
nftables table and writes nothing outside it: a rule placed in a manager's ruleset is
flushed away by that manager on its next reload, so applying one would be a fix that
silently stops being applied. A published port needs its own inbound allow, which
this command cannot know without asking the daemon.
kanea install
kanea install [--list] [--dry-run] [--only=containerd,runc,…] [--force]
[--bundle=…] [--containerd=external] [--arch=amd64|arm64]
[--prefix=…] [--conf-dir=…] [--data-dir=…] [--run-dir=…]
[--unit-dir=…] [--skip-units]
[--node-cidr=…] [--cluster-cidr=…] [--node-cidr6=…] [--cluster-cidr6=…]
Places the pinned host components: containerd, runc, rootless
buildkitd and the wasmtime shim. kanea init runs this for
you, so you normally only reach for it directly to inspect
(--list, --dry-run) or to repair one component
(--only, --force). Versions and SHA-256 hashes are
compiled into the binary and never fetched. Full notes are in
Installing → Host components.
kanea bundle
kanea bundle create [--arch=amd64|arm64] [-o=…] [--dir] [--no-images] [--containerd=…]
Builds an offline bundle of the host components on a machine that has a network, for
a node that does not. --dir writes a directory instead of a tarball and
--no-images omits the image components (leaving the binaries alone).
The bundle carries no hashes of its own: it is verified against the installing
node's binary. See Air-gapped nodes.
kanea ca
kanea ca show [--out=…] [--quiet] # the CA certificate in PEM kanea ca info # its subject, fingerprint and validity
show prints the node's own CA certificate, which is what devices need
in order to trust self-signed services and the dashboard:
kanea ca show > kanea-ca.crt. There is deliberately no
kanea ca rotate and no route that returns the CA key: rotation
means re-trusting every device, and a command for it would imply it is cheap.
kanea upgrade
kanea upgrade [--check] [--version vX.Y.Z] [--require-signature] [--no-fetch] [--skip-backup] [--dry-run] [--timeout=2m]
One command, both halves: it fetches the latest release (or --version),
verifies it; sha256 against the release's checksums.txt always, the
cosign keyless signature over that file when cosign is installed, a loud note when
it is not (--require-signature turns both soft endings into refusals);
installs it atomically over its own path, then takes a pre-upgrade
backup, restarts kanea-edge and then kanead in that order,
runs any state migrations, and waits for health. Already at the target version means
nothing to download, so running it twice is safe by construction.
--no-fetch restarts onto whatever binary is already installed
(the orchestration half alone) and --check only reports the running,
installed and latest versions. Air-gapped nodes get their binary from the
offline bundle flow; the fetch refuses with that pointer rather than hanging.
One thing it deliberately never does: rewrite systemd units. When release notes
say the units changed, re-run sudo kanea init after upgrading
(idempotent: the master key, accounts and settings are kept) then
sudo systemctl daemon-reload.
Deploying
kanea plan
kanea plan app.hcl [more.hcl …] [selector …] kanea plan --image=nginx:1.27-alpine --name=web --project=demo [--count=1]
A real dry run: one block per service, one row per resource that would be added,
changed or removed, the resulting workload budget, and every validation rule.
--remove-orphans adds the - destroy lines
kanea run --remove-orphans would act on, computed from the same scope,
so the plan and the run cannot disagree. Multiple files are parsed as one set, so
file order does not matter. This is where you find out about a cycle, a
cross-project secret, a duplicate domain or a rate limit that would fail open.
It also warns, never fails, on the weak defaults: an effective uid of 0, a
writable root filesystem, and an image following a moving tag without opting into
auto-update - the shapes hardening = "restricted" and
read_only_rootfs exist to replace.
$ kanea plan app.hcl
~ update shop/web
image web:v1 -> web:v2 (rolls allocs)
env + DB_URL, ~ LOG_LEVEL (rolls allocs)
volumes + cache s3:backups /var/cache rw 10 GiB budget (rolls allocs)
~ data local /data rw -> local /srv/data rw
expose + api.example.com (tls acme, port 8080)
publish - 8443/tcp port http
scaling min 1, max 5, rps 100 -> min 1, max 10, rps 100
+ create shop/worker (count 2, image worker:v3)
volumes + queue local /queue
files + /etc/app.conf (412 B, content 3f2a1b9c, mode 0644)
check http /healthz on port 8080 every 10s, timeout 2s, 3 failure(s)
Plan: 2 change(s) - 1 create, 1 update; 1 replace running allocs.
Run `kanea run` to apply.
Rows marked (rolls allocs) replace running containers, the ones
without are applied to the service as it is or republished to the edge. That is the
difference between an edit that costs a rolling restart and one that costs nothing,
and it is the number the summary line counts. It is derived from the same rule the
reconciler uses to decide a deploy, so the plan and the daemon cannot disagree about
which is which.
An environment value, because it may be a secret-env:
reference, and a config file's content, because it carries secret
placeholders. You get the keys, the paths, and a short digest of the content:
enough to see that something changed, never enough to leak it into a terminal
scrollback or a pasted issue.
kanea run
kanea run app.hcl [more.hcl …] [selector …] [--wait=60s] [--yes|-y] kanea run --image=nginx:1.27-alpine --name=web --project=demo [--count=1]
Applies the spec. kanea apply is an alias: same flags, same
behaviour. --wait is how long to wait for allocs to reach running
before returning; 0 returns immediately. Running it twice with an
unchanged spec does nothing: a deploy is a spec-hash mismatch, not an invocation.
It shows the plan and asks first. The block above is printed by the same
code kanea plan uses, so what you confirm is exactly what you were
shown, and then:
Apply? [Y/n]
Enter applies. Anything that is not y or yes aborts and
nothing is sent. --yes (or -y) skips the question.
The prompt appears only when stdin is a terminal. A piped or redirected stdin
(every CI job, every pipeline, every kanea run … < /dev/null)
applies without asking, exactly as it did before the prompt existed. No
existing automation needs --yes added to it.
An apply is additive. A service the spec no longer declares keeps running,
because kanea run on one file must not delete what another file
declares. --remove-orphans opts out of that, for the projects this spec
declares a project block for:
kanea run --remove-orphans app.hcl
Anything stored in one of those projects that the spec does not declare is deleted,
in the same atomic batch as the applies, so a rename never leaves both or neither in
place. It is refused with a selector (kanea run app.hcl shop/web
--remove-orphans) and with --image: a selector sends part of the
spec and --image declares no project, so neither can honestly claim to
be the whole of anything.
What a prune destroys is worth reading once.
--image creates a bare service and cannot update a real one. It
builds a record from project, name, count and image alone, so applying it over a
service that declares ports, expose, env, volumes, a health check or scaling would
delete them. It refuses when that would happen, naming what would be lost;
kanea deploy is the verb for changing the
image of a service that already exists.
A selector scopes both commands to part of the file: kanea run
app.hcl shop/web applies one service, shop alone a whole
project, and several selectors union. An argument that exists on disk is a spec
file; only a non-existent one is read as a selector, and one that is neither is
refused by name. The whole file is still parsed and validated (a selector never
changes what a spec means, only how much of it is sent) and every selector must
match at least one service in the file. An apply is additive either way: services
not in the request are never touched, so a scoped run cannot delete anything.
kanea deploy
kanea deploy [--project P] <[project/]service> <image> [--wait 60s] [--no-wait]
Point an existing service at a new image and leave the rest of its spec alone. It
reads the record, changes the image and writes the whole record back, because there
is no route that sets an image on its own and round-tripping is what stops a deploy
dropping a field. An init step declaring the task's previous image
moves with it (v0.31.1) - a migration on the app's own image must not run
yesterday's bytes against today's application - and the output names the steps
that followed; a step on any other image is untouched. Waits for the new image to
be running by default, which is what makes it usable as a pipeline's last step.
This is the CI verb; see Deploying a new image.
kanea stop
kanea stop [--project=p] <[project/]service> [--rm]
Scales to zero. --rm also deletes the service declaration, behind
the same confirmation kanea remove asks; a piped stdin is never
prompted, so scripted removals run unchanged.
kanea remove
kanea remove [--project=p] [--yes] <[project/]service>
Deletes one service declaration (alias: rm). Asks
[y/N] on a terminal; --yes/-y skips.
The containers, alloc records, VIP, routes and mounts go; volume data is
kept, so re-applying the spec brings the service back with its data.
kanea start
kanea start [--project=p] <[project/]service> [count]
stop's counterpart: scales a stopped service back up. The daemon does
not remember the pre-stop count (a stopped record says zero) so it starts one
replica unless a count is given, or an autoscaled service's own floor (the scale
route refuses a count outside the declared bounds). A service already running is
left exactly as it is: start is idempotent, never a second spelling
of scale.
kanea restart
kanea restart [--project=p] <[project/]service>
Rolls the service's allocs through its update policy: a generation bump, the same route the dashboard uses, not a second path into the runtime. It is also the way out of an exhausted crash-restart budget: the bump is a new spec hash, and the restart count belongs to the hash that spent it.
Inspecting
kanea ps
kanea ps [--project=p] [--service=s] [-a]
The alloc table: id, service, state, health, restarts, age, address.
A removed alloc leaves no record (only failed-and-still-declared ones
persist to explain themselves), so a stopped service is invisible here;
-a adds what is declared but not running: services scaled to
zero (stopped) and slots the reconciler has not created yet
(pending).
kanea describe
kanea describe [--project=p] <[project/]service>
One service in full: the declared spec beside what is actually true;
image and its pinned digest under auto-update, routes from every
expose block and published port, volumes and grants, the
alloc table with health verdicts, a stats snapshot, and the service's
recent events. Stats and health render absent as absent (-),
never as zero: a missing metric and an idle service are different facts.
The alloc table's REASON column says why each alloc last
stopped; OOMKilled, Signalled,
Error, Completed, or why one that never ran did
not: ImageFailed, VolumeFailed,
GrantFailed, NetworkFailed,
CreateFailed, StartFailed. An OOM kill is read
from the alloc's cgroup, never guessed from the exit code (a forced stop
exits 137 exactly like a memory kill does) so a stop reports as
Signalled, and a kill against a service that declared no
memory limit names the node's collective ceiling rather than a number
nobody typed. The reason shows whatever the alloc's current state is: a
running row carrying OOMKilled means it is up
now and was killed for memory last time.
kanea status
kanea status [--project=p] [[project/]service]
Health, recent events, current and desired counts, and the scaling picture.
kanea logs
kanea logs [--project=p] <[project/]service> [-f] [--tail=N] [--alloc=ID] [-c NAME] [--previous]
Merged across allocs by default; --alloc narrows to one.
--tail shows the last N lines before following. -c
reads an init container's log instead of the
task's, by its block name. --previous reads the log files a
stopped or removed service left on disk - a torn-down alloc keeps no
record, but its log file survives, so the last output before a stop, a
crash-loop's end, or a removal is still readable. Log drains are
non-blocking with drop counters: a slow reader can never stall a
workload's write().
kanea exec
kanea exec [--project=p] [--alloc=ID] [--user=UID] [-it] <[project/]service> -- <command…>
A debug shell inside an alloc. Admin-only, and audited whether or not the session establishes: "someone tried to open a shell on production" is worth keeping either way, so the attempt and the requested command are both recorded.
- The
--is required. The command crosses the wire as separate arguments rather than one joined string, because every rule for splitting a string back into arguments is wrong for something somebody will eventually pass. -itallocates a terminal and forwards stdin. A shell needs it.--usertakes a numeric uid only. Resolving a name would mean reading the container's own/etc/passwd, and a container-controlled file deciding which uid the control plane runs a process as is not a thing to build.
kanea ui
kanea ui [--addr=…] [--open]
Prints the dashboard URL; --open launches a browser.
Scaling and builds
kanea scale
kanea scale [--project=p] <[project/]service> <count>
Writes the desired count and returns; the reconciler converges. This is the same route the autoscaler uses, which is why manual and automatic scaling can never disagree about mechanism.
kanea build
kanea build [--project=p] <[project/]service> [--deploy=true] [--follow=true]
Triggers the service's build pipeline. --deploy rolls the built digest out
on success; --follow streams the build log. Builds are serialised: a
second one is queued, and refused rather than blocked when the queue is
full.
kanea project
kanea project sync <project> # re-read the git source now kanea project builds <project> [--service=s] [--limit=N]
kanea images
kanea images [--clean] [--json]
The node's containerd images, per project, with each image's size, age
and whether anything still references it. --clean runs one
image GC sweep now (admin-only, audited) and reports what it removed and
reclaimed; the rules it deletes under, and the gc block that
tunes them, are in the node config
reference. On a node whose default pull_policy is
"never" the sweep answers with a refusal naming why:
preloaded images cannot be re-pulled.
kanea volume
kanea volume list [--json]
Every storage resource with the mounts using it nested underneath: driver, where the bytes live, measured usage, declared budget and mount state. It is what to reach for when a disk is filling up - a flat list would repeat an NFS export's address once per service and still not say they were the same export.
Usage is sampled in the background, so a volume reads - until
it has been measured, and s3 volumes are never walked. That dash is
an absence, not a zero: a volume nobody has looked at and an empty one are
different facts. A local volume appears once per alloc, because that is how many
of it there are - each with its own contents, and each judged against the
budget separately.
There are no other subcommands, deliberately. A volume exists because a spec declares it, so creating or deleting one here would be a second way to change desired state that the reconciler would immediately undo.
Secrets and accounts
kanea secret
kanea secret put [--from-file=path] <project>/<name> # value on stdin kanea secret ls [<project>] kanea secret rm <project>/<name>
Not for an operator, not over the API, and not for an AI agent at any tool tier.
ls lists names. The value goes in and is only ever resolved into a
running alloc.
kanea user and kanea token
kanea user add [--role=admin|viewer] <name> kanea user ls kanea user rm <name> # also revokes the account's sessions kanea user revoke-sessions <name> # sessions only; the account stays kanea token create [--role=viewer] [--expires-in=720h] <name> kanea token ls kanea token rm <id>
Accounts live in the Store, not in a config file. Tokens default to never expiring;
--expires-in takes a Go duration. The first admin is created by
kanea init itself; OIDC and LDAP identities never appear in
user ls: they are ephemeral, a session and nothing else.
Backup and restore
kanea backup create [--reason="on-demand"]
kanea backup list
kanea backup verify <archive-id>
kanea restore --from s3://bucket/prefix [--snapshot=ID] [--target=path]
[--s3-endpoint=…] [--s3-region=…] [--s3-access-key=…] [--s3-path-style]
[--master-key=path] [--data-dir=…]
verify reads the archive and checks its hashes and authentication tags
without restoring anything: an archive that cannot be verified is one you find out
about now rather than during an outage.
The command stages the restore; it is performed at the next daemon start, before anything opens the Store. That is the interface rather than a safety check: the API has no method that restores at all, and there is no restore button in the dashboard, because a restore replaces everything on the node and belongs at a terminal.
Daemons
Normally systemd runs these. The flags are here because systemctl cat will show them to you.
kanea agent
kanea agent [--config=/etc/kanea/kanea.hcl|off] [--data-dir=…] [--log-dir=…]
[--volume-dir=…] [--socket=/run/kanea/kanead.sock]
[--containerd=…] [--network=ebpf|netns] [--node-cidr=…] [--cluster-cidr=…]
[--service-cidr=10.201.0.0/16] [--bpf-dir=/sys/fs/bpf/kanea]
[--dns-listen=…|off] [--dns-upstream=…] [--registry=127.0.0.1:5100|off]
[--allowed-host-paths=…|off] [--passthrough-config=…|off]
[--listen=…|none] [--listen-cert=…] [--listen-key=…]
[--base-domain=…] [--tls-default=acme|self-signed|provided|plaintext]
[--tls-certs-config=…] [--acme-email=…] [--edge-group=kanea-edge]
[--publish-ports=1024-65535] [--secrets-providers-config=…]
[--dashboard=true] [--autoscale=true] [--log-level=info]
[--backup-s3-endpoint=…] [--backup-s3-region=…] [--backup-s3-access-key=…]
[--oidc-client-id=…] [--ldap-url=…] [--acme-dns-tsig-key=…]
--listen beyond loopback requires --listen-cert and
--listen-key. Unset, the listener comes from the server config's
bind stanza when one is declared (below); --listen none
forces socket-only regardless of the file. Credential-shaped options such as
--oidc-client-secret, --ldap-bind-password and
--acme-dns-tsig-secret take secret: references, never
literals.
--ldap-url enables directory logins beside local accounts and OIDC:
ldaps:// (or ldap:// with StartTLS forced; there is no
insecure option), a user search under --ldap-user-base-dn with
--ldap-user-filter, and group-to-role mapping through
--ldap-admin-groups/--ldap-viewer-groups, deny-by-default.
A local account with the same name always wins, and the login rate limit runs before
any bind reaches the directory.
The operator-owned settings live in the server config,
/etc/kanea/kanea.hcl: which directories host volumes may
come from, which devices and sockets are granted to which projects, where the API
and dashboard listen, which resolvers the internal DNS forwards to, and node-wide
spec variables. It is read once at startup, refused if anyone but its owner could
have written it, and absent by default. The complete reference - every stanza,
every field, the precedence rules and how to apply an edit - is
Node configuration.
Every one of those has a flag that overrides it, and the disable words differ
(--listen none, but off elsewhere). The full precedence
table is in Node configuration → Flags and
precedence.
kanea init renders exactly seven of these flags into
ExecStart: --data-dir, --log-dir,
--network, --node-cidr, --cluster-cidr,
--edge-group, and --listen (with its cert pair, and only
when the bind stanza does not own the
listener). The three *6 flags join them only when IPv6 is enabled, so
a v4-only unit stays byte-identical.
Everything else on this page is set by editing the unit
(systemctl edit kanead) or by the server config. That includes two
that look like they should be init's, because init has flags by the same name and
uses them only for the install: --containerd and
--buildkit. A non-default containerd socket or buildkit address
needs a unit edit.
kanea edge
kanea edge [--routes=/run/kanea-edge/routes.json] [--certs=…] [--http=:80]
[--poll=…] [--drain=…] [--memory-limit=128MiB] [--log-level=info]
Its own process and its own systemd unit, with no After=kanead.service:
north-south traffic surviving a control-plane restart is the entire reason it is
separate. --drain is how long in-flight requests get on shutdown.
MCP
kanea mcp [--socket=/run/kanea/kanead.sock] [--verbose]
A stdio MCP server for an AI agent running on the node: 24 tools in read, mutate and
destructive tiers. The same server is available over streamable HTTP at
/mcp on the API listener for agents running anywhere else.
Setting it up for Claude Code, opencode or Codex is
AI agents (MCP), which has the client configuration for both
transports.
Tools reach the platform only by making requests against the API's own handler, so an agent is never more privileged than the credential it was given. Tiers are advertised as well as enforced, and the advertisement fails closed. A refusal comes back as a tool result rather than a protocol error, because the model is what has to react to it.
kanea version
Prints the version stamped in at build time. kanea upgrade compares it against what the running daemon reports.
AI agents (MCP)
Kanea ships a Model Context Protocol server, so an agent can look at the node and operate it through the same API you do. It speaks both transports: stdio, for an agent running on the node itself, and streamable HTTP on the API listener, for one running anywhere else.
There is nothing to install. The server is the same kanea binary, and
the tools reach the platform only by making requests against the API's own handler
- which is what makes the next section a statement about capability rather
than about good intentions.
A tool's only verb is "send this request", so nothing in the MCP server can hold a Store, a secrets store or an auth store, and no tool can do something the token behind it could not do at the REST API. Which tools an agent can even see follows from that token's role, and there is no secrets tool at any tier - not a redacted one, none.
Which transport
| stdio | Streamable HTTP | |
|---|---|---|
| Command / endpoint | kanea mcp | https://<node>:8600/mcp |
| Agent runs | On the node | Anywhere |
| Authenticates by | Unix socket access | A bearer token |
| Needs | root, or kanea group membership | An API listener (--listen or a bind stanza) |
| Role | Always admin: the socket is root-equivalent | Whatever the token has |
The practical rule: if your editor runs on the node, use stdio. If you are on
a laptop and the node is a server - which includes every Mac, since
kanead is Linux-only - use HTTP. Prefer HTTP when you want the
agent to be read-only, because that is the transport where you can choose a role.
stdio, on the node
kanea mcp [--socket=/run/kanea/kanead.sock] [--verbose]
It talks to kanead over the unix socket and speaks MCP on stdin and
stdout. --verbose logs protocol activity to stderr; stdout is
the protocol channel, and a single stray line on it corrupts the stream in a way
that presents as a client which connects and then does nothing.
Your editor runs as you, not as root, so kanea mcp fails to reach the
daemon unless you have joined the group kanea init created:
sudo usermod -aG kanea $USER # then log out and back in
Membership is root-equivalent, exactly like docker's group. An agent on this transport is therefore always an admin, with every tool available to it. If that is more than you want to hand over, use HTTP with a viewer token instead.
Claude Code
claude mcp add kanea -- kanea mcp
Add --scope project to write it into the repository's
.mcp.json instead of your own config, so everyone working on that
project gets it. The file form is the usual one:
{
"mcpServers": {
"kanea": {
"command": "kanea",
"args": ["mcp"]
}
}
}
opencode
In opencode.json at the project root, or your global
~/.config/opencode/opencode.json:
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"kanea": {
"type": "local",
"command": ["kanea", "mcp"],
"enabled": true
}
}
}
Codex
codex mcp add kanea -- kanea mcp
Which writes to ~/.codex/config.toml; the equivalent by hand is:
[mcp_servers.kanea]
command = "kanea"
args = ["mcp"]
Streamable HTTP, from anywhere
The same server is mounted at /mcp on the API listener -
deliberately not under /v1, because that path is this API's versioned
surface and MCP carries its own version in the initialize handshake. It
exists only when the node has an API listener at all: a socket-only node
(--listen none) has no HTTP transport to offer.
Mint a token for the agent, and choose its role deliberately:
sudo kanea token create --role viewer agent-readonly # can look, cannot touch sudo kanea token create --role admin agent-deploy # can deploy, scale and stop sudo kanea token create --role viewer --expires-in 720h agent-30d
The secret is printed to stdout once and is not recoverable, so
… > token.txt captures exactly it while the commentary goes to stderr.
kanea token ls shows what exists and kanea token rm <id>
revokes one.
Claude Code
claude mcp add --transport http kanea https://192.168.1.10:8600/mcp \ --header "Authorization: Bearer $KANEA_TOKEN"
opencode
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"kanea": {
"type": "remote",
"url": "https://192.168.1.10:8600/mcp",
"enabled": true,
"headers": { "Authorization": "Bearer <token>" }
}
}
}
Codex
[mcp_servers.kanea]
url = "https://192.168.1.10:8600/mcp"
bearer_token_env_var = "KANEA_TOKEN"
A JSON entry with a url and no type is a
configuration error, not a remote server. Claude Code reads a typeless entry
as stdio and skips it with a message saying so; opencode needs
"type": "remote" for the same reason. Copying an
mcpServers block from another client is where this usually happens.
If the node serves a self-signed certificate - which it does by
default, whether from kanea init's provisioned pair or the node CA
- a client that does not trust it will fail the TLS handshake before MCP is
ever spoken. Install the CA with kanea ca show on the machine running
the agent, or point the client at a name covered by an acme
certificate.
The tools
Twenty-five tools in three tiers. The tier an agent gets is its credential's
role, and the listing is filtered, not merely refused: a viewer token does not
see the mutating tools in tools/list at all, so the model cannot
propose calling one.
| Tier | Needs | Tools |
|---|---|---|
| Read 13 tools | viewer | list_projects, get_project, list_services, get_service, list_allocs, get_logs, get_events, get_node_stats, get_service_stats, list_pipelines, list_storage, list_backups, get_audit* |
| Mutate 9 tools | admin | plan_spec, apply_spec, scale_service, restart_service, stop_service, deploy_service, run_pipeline, create_backup, test_notification |
| Destructive 3 tools | admin, plus an explicit confirm | restore_backup, delete_service, delete_project |
* get_audit is the one read tool a viewer cannot actually use: the
audit log is admin-only at the API, so the tool is listed but the call comes back
403. That is the tier system being advisory and the API being authoritative, in the
one place where they disagree.
A destructive tool called without confirm: true is refused with a
result telling the model to call it again only if the operator has explicitly
asked for this, and that it cannot be undone. The gate is enforced in the MCP
server rather than the API, because it is not an authorization rule - an
admin is allowed to delete a project and the API will let them. It is a rule about
agents: a destructive action has to be arrived at deliberately rather than by
pattern-matching a tool name, and a human reading the transcript afterwards can see
that it was. A refusal comes back as a tool result rather than a JSON-RPC
protocol error, because a protocol error is handled by the client library and never
reaches the model, which is the one party that has to react to it. An
unknown tool is the exception, and deliberately so: that is a bug in the
client, not a refusal for a model to reason about.
Every call lands in the audit log with the token's id, exactly like a REST call,
because it is one. The dashboard's audit tab, GET /v1/audit and
the get_audit tool all show it - which means an agent can read
its own trail, and so can you.
Deciding what to hand over
- Start with a viewer token. It covers the whole diagnostic loop - what is deployed, what is failing, the logs, the events, the stats, the audit trail - which is most of what an agent is actually useful for, and it cannot change anything.
-
Give admin only when you want the agent deploying.
apply_spec,scale_serviceanddeploy_serviceare real deploys, with the same rolling-update rules and health gates a human's would have; nothing about a change arriving from an agent makes it a different kind of change. - Prefer HTTP with a scoped token over stdio when the distinction matters. The stdio transport authenticates by socket access, and socket access is root-equivalent, so there is no read-only version of it.
-
Set an expiry.
kanea token createwarns when a token never expires, for good reason:--expires-in 720hcosts nothing and bounds the blast radius of a leaked editor config. -
Secrets are unreachable by construction, so do not plan around them. There
is no tool that reads one at any tier; a spec references
secret:<project>/<name>and the agent never sees a value.
Working with a remote node
The CLI reaches kanead two ways. Locally it uses the unix socket, and
your credential is membership of the kanea group. From anywhere else it
uses the node's HTTPS listener, and your credential is a bearer token. Everything
below is the second one - a laptop, or a CI runner that has just built an image.
| Local | Remote | |
|---|---|---|
| Reaches the daemon by | the unix socket | https://node:8600 |
| Credential | kanea group membership | a bearer token |
| Set up with | usermod -aG kanea $USER | KANEA_URL, KANEA_TOKEN, KANEA_CA_CERT |
| Get the CLI from | the install script | Homebrew, or a release archive |
| Authorized as | admin, always | the token's role |
A remote CLI is not a new privilege. It presents the same token the REST API and the MCP server have always accepted, and the role decides exactly what it decided there: a viewer token can look and cannot deploy. What was missing until v1.82 was a client that spoke it, not a permission.
Setting it up
1. Get the CLI
On a laptop, or in a CI image, Homebrew is the easy channel - it needs no root and no node:
brew tap m18h/kanea brew trust m18h/kanea # brew ≥ 6 refuses formulae from untrusted third-party taps brew install kanea
For a container base without Homebrew, take the release archive; it is one static binary and needs nothing installed beside it. The verify-by-hand steps are the same.
2. Mint a token, on the node
sudo kanea token create --role admin ci --expires-in 720h
The secret goes to stdout once and is not recoverable, so
… > token.txt captures exactly it while the commentary goes to
stderr. Deploying needs --role admin; a viewer token can read
everything and change nothing, which is the right choice for a job that only
inspects. kanea token ls shows what exists and
kanea token rm <id> revokes one, immediately.
kanea token create warns when a token never expires, for good reason:
--expires-in 720h costs nothing and bounds how long a leaked CI
variable is worth anything. Rotating is one command on each side.
3. Trust the node's certificate
A Kanea node serves its own CA's certificate or a self-signed one unless you gave it
an acme name, so a fresh client does not trust it and the TLS handshake
fails before anything else happens. Export the CA once:
sudo kanea ca show > kanea-ca.crt # on the node export KANEA_CA_CERT=/path/to/kanea-ca.crt
KANEA_CA_CERT takes a file path or the PEM itself, because a CI
system hands secrets to a job as values and writing them to a file first is a step
that gets skipped. Both of these work:
KANEA_CA_CERT=/etc/kanea/ca.crt KANEA_CA_CERT="$(cat kanea-ca.crt)" # or a CI variable holding the PEM
If the node's certificate comes from Let's Encrypt (an acme
bind stanza), skip this entirely: the
public roots already trust it. There is deliberately no
--insecure-skip-verify; the whole reason the inline form above
exists is that a skip flag is what people reach for when the honest path is awkward.
4. Point the CLI at the node
export KANEA_URL=https://kanea.apps.example.com:8600 export KANEA_TOKEN=$(cat token.txt) kanea ps kanea status kanea logs -f shop/web
Every client command takes these, as environment variables or as
--url, --token and --ca-cert. A flag you
actually pass beats the environment; an explicit --socket keeps a
command local even with KANEA_URL exported, so a node's own shell is
never redirected by a variable someone set for another tool.
Deploying a new image
kanea deploy [--project P] <[project/]service> <image> [--wait 60s] [--no-wait]
kanea deploy points an existing service at a new image and
leaves everything else exactly as declared. It reads the service's record,
changes the image, and writes the whole record back - which matters, because
there is no API route that sets an image on its own, and round-tripping the record
is what makes it impossible for a deploy to drop a field it does not know about.
One deliberate exception rides along (v0.31.1): an init step whose
image is exactly the task's previous one is pointed at the new image too, so a
migration that runs the app's own image deploys in lock-step with the app.
Byte-for-byte equality on the declared reference, nothing cleverer; a step on
its own image never moves. The same rule applies when a GitOps build deploys
its result, so a step that starts on the app's image stays on it build after
build.
kanea deploy shop/web ghcr.io/acme/web@sha256:9f2a…
It waits for the new image to be running and fails if it does not, so a pipeline
goes red at the deploy rather than the next morning. --no-wait returns
as soon as the change is accepted. Deploying the image a service already declares is
reported and does nothing, so re-running a pipeline on an unchanged commit is not an
error.
Prefer a digest to a tag. A tag can move under you, so two allocs replaced a minute apart can run different code; a digest is the thing that was built.
kanea run --image
That flag builds a service from nothing - project, name, count, image -
so applying it over an existing service used to delete its ports, expose block,
env, volumes, health check and scaling, silently, and you found out when
traffic stopped arriving. It now refuses and names what would be lost. Use it to
create a bare service; use kanea deploy to change one that exists.
In a pipeline
The CLI ships as a container image, ghcr.io/m18h/kanea, which is
what a runner wants: linux/amd64 and linux/arm64 in
one manifest list, tagged vX.Y.Z and latest. Pin the
version - that is what tags are for. See
the container image for what is in it and how to
verify it.
GitLab:
deploy:
stage: deploy
image:
name: ghcr.io/m18h/kanea:vX.Y.Z # pin the version
entrypoint: [""] # GitLab runs the job script as the container command
variables:
KANEA_URL: https://kanea.apps.example.com:8600
script:
- kanea deploy shop/web "$CI_REGISTRY_IMAGE@$IMAGE_DIGEST"
# KANEA_TOKEN: masked project variable
# KANEA_CA_CERT: a file variable, or the PEM as a masked variable
GitHub Actions:
- name: Deploy
# or run the image directly: docker run --rm -e KANEA_URL -e KANEA_TOKEN \
# ghcr.io/m18h/kanea:vX.Y.Z deploy shop/web "$IMAGE"
env:
KANEA_URL: https://kanea.apps.example.com:8600
KANEA_TOKEN: ${{ secrets.KANEA_TOKEN }}
KANEA_CA_CERT: ${{ secrets.KANEA_CA_PEM }}
run: kanea deploy shop/web "ghcr.io/acme/web@${{ steps.build.outputs.digest }}"
unknown command "sh" - the image's entrypoint is
kanea and GitLab passes the job script to the container as its
command; use the map form above with entrypoint: [""].
KANEA_TOKEN is set but no endpoint - you set the token and not
the URL. … has no credential without a token - the secret was
not exported and arrived empty, which is why an empty variable counts as unset
rather than as a token. … certificate from an unknown authority
- set KANEA_CA_CERT. refusing to send a token … over
plain HTTP - the endpoint is http:// and not loopback,
which would put the token on the wire in clear text.
What stays on the node
Two commands refuse a remote endpoint, and both are about the machine rather than the platform:
-
kanea upgraderestarts this host'skaneadandkanea-edgeunits and installs a binary over this host's. Reading a remote daemon's version and then restarting local services is the worst outcome available here, so it is refused by name. Upgrade over ssh, or on the node. -
kanea mcpserves MCP over stdio, and its credential is the local socket. A remote agent already has the node's own/mcpendpoint with a bearer token, so a second spelling would add nothing.
kanea init, install, bundle and
doctor never take an endpoint at all: they act on the host they run on.
Everything else works remotely, including kanea exec (over
wss://, with the same token) and kanea logs -f.
The alternative: let the node pull
A token in CI is not the only way. Kanea's own design answer is the other direction: the node watches a git repository, and CI commits the digest rather than calling out.
- Your pipeline builds and pushes the image, then commits the new digest into the
spec repository the project's
gitblock names. - The push fires the webhook (
POST /v1/webhooks/git/<project>, authenticated with GitLab'sX-Gitlab-Tokenor GitHub'sX-Hub-Signature-256against the project'swebhook_secret_ref). - Kanea re-reads the repository over its own credential and applies what it finds. Nothing is deployed from the request, so a forged delivery cannot choose an image.
It needs no token in CI and leaves a git history of every deploy; it costs a spec
repository and a commit per release. kanea project sync <project>
forces the read immediately rather than waiting for the poll. Which one to prefer is
a question of whether you would rather your deploys be a push or a record.
Troubleshooting
Where to look when something is wrong, in the order that usually finds it: the workload first, then the daemons underneath it, then the node. Everything in this section is read-only and safe to run on a live node.
Checking services
Start with what Kanea believes is true:
kanea ps -a # every alloc; including stopped services and pending slots kanea status shop/web # health, recent events, current vs desired counts kanea describe shop/web # the full picture: spec, routes, volumes, allocs, stats, events
Three columns in ps carry most of the signal. State:
pending means the reconciler has not created the slot yet; usually an
image still pulling, or a dependency that is not healthy. Health: a
- means the service declares no health check, which is a different fact
than failing one; a check-free service is never reported unhealthy, only running or
not. Restarts: a climbing count is a crash loop; when it stops climbing the
restart budget is exhausted and the alloc has been failed and left alone
(see below for the way out). A node reboot moves
no counter: allocs found dead at boot are recovered without spending the budget.
Then the processes underneath. A standard install runs four units:
systemctl status kanead kanea-edge kanea-containerd kanea-buildkit
kanead: the control plane. Workloads and north-south traffic both survive it being down; what stops is change: deploys, scaling, certificate renewal, the API and dashboard.kanea-edge: the ingress proxy. Deliberately independent ofkanead(noAfter=in either direction): it serves the last route snapshot it read from disk whether or not the control plane is up.kanea-containerd: Kanea's own containerd, socket at/run/kanea/containerd.sock. Restarting it does not stop running containers (KillMode=process: shims outlive it), but nothing can be created or probed while it is down. Absent when the node adopted an existing daemon with--containerd external.kanea-buildkit: the rootless build daemon. Only builds need it.
Finally the node itself: kanea doctor verifies dependencies and their
pinned versions, the containerd socket, bpffs and the cgroup2 mount, slice placement
and the effective memory floor, the build socket, disk headroom and clock sync, and
it names known interference, like a firewall FORWARD-drop policy (docker, ufw) eating
east-west traffic. Safe to run any time.
Viewing logs
Workload logs stream through the CLI:
kanea logs shop/web -f # merged across allocs, follow kanea logs shop/web --tail=200 # the last 200 lines first kanea logs shop/web --alloc=<id> # one alloc only (ids from kanea ps)
On disk they are one file per alloc under /var/log/kanea/allocs/
(<alloc-id>.log); a replaced alloc starts a fresh file under its new
id. Drains are non-blocking with drop counters, so a burst of logging can be dropped
but can never stall the workload's write(): if lines are missing under
load, that is the drop counter doing its job, not a lost file.
Daemon logs go to stderr, which under systemd means the journal:
journalctl -u kanead -e # control plane: reconciler, deploys, certificates, backups journalctl -u kanea-edge -e # ingress: TLS, routing, published ports journalctl -u kanea-containerd -e # runtime: image pulls, task create failures journalctl -u kanea-buildkit -e # the build daemon journalctl -u kanead -f --since "15 min ago"
A deploy that goes wrong is usually legible in kanead's journal; a task
that will not create at all (a missing shim, a device the node did not grant) often
explains itself one level down in kanea-containerd's.
Build logs are their own stream: kanea build shop/web --follow
live, kanea project builds shop for history, and one file per run under
/var/lib/kanea/builds/.
Slices and resources
Everything Kanea runs sits in one of two cgroup slices, and that split is the
resource-isolation story (architecture):
kanea.slice holds the control plane (all four units above) with a
kernel-guaranteed memory floor (MemoryMin, default 256 MiB;
raise it with kanea init --reserve on a node that runs builds), and
kanea-workloads.slice holds every alloc under a collective
ceiling of total RAM minus that reserve.
systemctl status kanea.slice # the control plane, with its live memory number systemctl status kanea-workloads.slice # every alloc as a child cgroup systemd-cgls kanea-workloads.slice # the tree, one cgroup per alloc systemd-cgtop # live CPU/memory per slice
The ceiling is computed and applied by kanead at startup (it depends on
how much memory the node has, which a unit file cannot know) so read it from the
cgroup filesystem, which is the ground truth either way:
cat /sys/fs/cgroup/kanea.slice/memory.min # the floor cat /sys/fs/cgroup/kanea-workloads.slice/memory.max # the ceiling cat /sys/fs/cgroup/kanea-workloads.slice/memory.current
Per-alloc cgroups live one level down
(/sys/fs/cgroup/kanea-workloads.slice/<alloc>/) with their own
memory.max, memory.current and cpu.max. A
declared resources limit is enforced exactly there; an omitted one reads
as max: unbounded within the collective ceiling, by design, never a
filled-in default.
The floor and the OOM score adjustments live in the unit files, not in the Go code:
a Kanea started outside its units runs without the guarantee, and the first time
the node is under memory pressure the kernel picks whatever is largest, which is
usually kanead. kanea doctor checks slice placement and
the effective floor for exactly this reason; journalctl -k | grep -i
oom shows what the kernel actually chose.
Common situations
-
A service is crash-looping.
kanea logsfor why. Once therestartbudget is exhausted the alloc is failed and left alone on purpose:kanea restart shop/webclears it, because the restart count belongs to the spec hash that spent it and a restart is a new one. So does deploying a fix. Only crashes the daemon watched count: after a power loss or reboot, every alloc found stopped comes back on its own with the budget untouched (v0.31.0; earlier versions charged one attempt per outage and could leave services failed at boot). -
A deploy did nothing. A deploy is a spec-hash mismatch, not an invocation:
re-running an unchanged spec is a no-op by design.
kanea planshows the diff the daemon would see: an empty one means the spec really is what is running. -
The site is unreachable but the service is healthy. Work outward:
kanea describeshows the routes the service declares, thensystemctl status kanea-edge. The edge reads its world from two files (/run/kanea-edge/routes.jsonand its certificate bundle) so a route missing from that snapshot is a publishing problem inkanead's journal, and a route present in it is a proxying problem in the edge's. -
Browsers show a certificate error. Deliberate fail-closed behaviour, not
breakage: a domain whose certificate is not issued yet refuses the handshake, and a
providedcertificate that stops resolving serves plaintext; neither ever silently falls back to a weaker certificate.kanead's journal has the ACME or resolution failure. -
Containers cannot reach each other.
kanea doctorfirst: a foreign FORWARD-drop firewall policy (docker, ufw) is a finding it names. Then remember policy is deny-by-default: a cross-project call needsallow_fromon the callee. -
The
bindstanza is ignored: the dashboard stays on localhost. An explicit--listenalways beats the file, and an init run from before the stanza existed rendered--listen 127.0.0.1:8600into the kanead unit.systemctl cat kanead | grep -- --listenconfirms it; remove the flag (and--listen-cert/--listen-keyif present) fromExecStart, thensystemctl daemon-reloadand restart. A re-run ofkanea initwill not put it back, withbind.api_addrdeclared it renders no listen flags at all. -
kaneadwill not start on a fresh cloud node. If the journal says/etc/resolv.conf lists only loopback resolvers, the node is older than v0.23.2 and is hitting a fixed bug: systemd-resolved's127.0.0.53stub is the only nameserver a stock Debian or Ubuntu server has, and it was being discarded, leaving no upstreams and a refusal.sudo kanea upgradeis the fix - the stub is a perfectly good upstream, sincekaneadforwards from the host's own namespace. Adnsstanza pins resolvers if you want to, but it is an override, never a requirement. -
The server config seems to do nothing. Three things, in order.
Did it load?
journalctl -u kanead | grep 'server config'showsserver config loadedwith the path, or nothing at all if the file is absent. Is a stanza being ignored? The same grep showscarries stanzas this version does not readwith the offending names - that is where a misspelled stanza lands, since only unknown attributes inside a read stanza are errors. Is a flag beating it? Look foris not consulted: an explicit flag on the unit always wins, and--config offdisables the file entirely. Remember the file is read once, so an edit needssystemctl restart kanead. -
Builds fail with
exec: "buildctl": executable file not found in $PATH. The binary is installed -kanea doctorwill confirm it is at its pinned version - butkaneadcannot find it, becausebuildctllives in Kanea's own bin dir and systemd's defaultPATHdoes not include it. Nodes whose units were written before this was fixed need the units regenerated:sudo kanea init(idempotent) thensudo systemctl daemon-reload && sudo systemctl restart kanead.kanea doctornames it under thebuildkitcheck. Upgrading alone will not fix it:kanea upgradedeliberately never rewrites units. -
The dashboard fails the TLS handshake, or the browser says the certificate is
broken. Check the port first. The dashboard is
kanead's own listener, so it answers onbind.api_addr's port (https://<name>:8600); the bare name goes to:443, which iskanea-edge, and on a node with no exposed services the edge holds no certificates at all, so the handshake dies withSSL_ERROR_INTERNAL_ERROR_ALERTrather than a 404. If the port is right, the other cause is a certificate that has not issued yet: the listener refuses the handshake rather than serving something weaker, andjournalctl -u kanead | grep 'api listener certificate installed'says whether it arrived. See the worked example. -
A service has no public name, or the name does not resolve. Three
separate things have to be true, and they fail differently.
kanea describeshows the domains the service actually claims: if it shows none, the node has no--base-domain- whichkanea initnever sets, so it is a drop-in on the unit, not a spec change. If it shows a name that does not resolve, the DNS record is yours to create; a wildcard covers every service at once. If it resolves but the certificate never issues, the record has to point at the node before issuance for HTTP-01, and--acme-emailhas to be set or nothing is requested at all. -
A cross-project call resolves but times out. That is policy, not DNS.
Names resolve for everyone because DNS is not the security boundary; the datapath
is, and it denies between projects by default. The callee needs
allow_fromnaming the caller. A timeout rather than a refusal is what a dropped packet looks like, which is why this reads as a network fault. -
The autoscaler stopped scaling. Its circuit breaker trips on purpose after
repeated failed actions, and says so in the journal, on the dashboard, and as
kanea_circuit_breaker_openin/v1/metrics. The trip survives a daemon restart by design: restartingkaneadis not a way around it; fixing the cause is.kanea scalestill works meanwhile: the breaker pauses the automatic decisions, not the route they travel.