Operations & Production
This page covers running the Stone-Age.io Platform in production: what state you're protecting, the availability model, backup and recovery, upgrades, and how component versions relate to each other.
The short version: the platform's operational posture follows directly from the plane split. The Data Plane gets high availability the NATS-native way — clustering at the hub, leaf-node autonomy at the edge. The Control Plane gets something simpler and, for its role, better: a small, easily-backed-up SQLite file and a recovery time measured in minutes.
1. What You're Protecting
Operational state lives in different places, with different owners and different backup stories. Know which is which before designing your routine:
| State | Where it lives | Loss impact | Protected by |
|---|---|---|---|
| Identity hierarchy — NATS Operator key, NATS Accounts, Nebula CAs, user/Thing credentials | Control Plane (pb_data) |
Severe. The NATS Operator key is the root of the entire chain of trust — losing it orphans every Account and credential it signed. | This page (§3). |
| Inventory & contracts — Orgs, Things, Thing Types, Locations, schemas, rules-adjacent config | Control Plane (pb_data) |
High, but recoverable — re-entry is tedious, not impossible. Also recoverable from a GitOps workspace. | This page (§3), plus stone pull workspaces. |
| Live state — Digital Twin KV, other KV buckets, JetStream streams | NATS servers (JetStream storage) | Depends on the bucket. The reported twin (twin) repopulates as devices report again. The desired twin (twin_desired) does not — operators wrote it, nothing re-derives it, and losing it loses every setpoint. Neither does any other bucket a tenant created, nor its contents. Stream retention is a buffer, not an archive. |
JetStream replicas (replicas: 3 on clustered NATS), stream mirrors. |
| Historical telemetry | Your Layer 3 TSDB | Your call — it's BYO. | Your TSDB's own backup tooling. |
| Edge leaf config | nats-leaf.conf + creds on the edge box |
None. Regenerate with agent -leaf-config — see Leaf Nodes. |
Nothing needed. |
The takeaway: pb_data is the crown jewels. It's also a single directory — one SQLite database plus pb_data/storage, which holds every uploaded file (Thing and Location photos, floor plans, organization logos, avatars) — which makes protecting it straightforward. The upload half grows with the inventory and rides in every backup; stone_age_database_size_bytes counts only the database, so watch the directory's disk usage too (Health & Metrics §3).
Backups contain secrets. A Control Plane backup includes the NATS Operator key, every org's Nebula CA private key, and credential material — including every issued
.credsfile and Nebula host config in plaintext, since at-rest encryption covers the minting keys and not issued credentials (Configuration §2.2). Treat backup artifacts with the same care as the live database: restrict the S3 bucket, encrypt at rest, and don't leave downloaded copies on workstations. Consider--encryptionEnv(see Configuration §4) to encrypt app settings at rest.
2. The Availability Model
The platform deliberately puts high availability where it matters and fast recovery where HA would be wasted complexity.
Data Plane: HA by construction
The runtime path — telemetry, commands, rules, live dashboards — never depends on a single process:
- NATS clusters horizontally. Run three or five
nats-servernodes and the bus survives node loss; JetStream streams and KV buckets withreplicas: 3survive it durably. This is stock NATS operations — their docs cover it well. - Leaf nodes keep sites autonomous. A WAN or hub outage doesn't stop site-local devices, rules, or stream processors. The KV buckets a site declares for sync reconverge when connectivity returns; anything else published while islanded reaches the hub only if a stream at the leaf holds it. See Leaf Nodes.
- Rule engines and stream processors scale horizontally, and their durable state is in replicated KV. What they hold in memory — rule-router's throttle windows, for one — is per instance and is lost on restart.
Control Plane: recovery-oriented, not failover-oriented
The Control Plane is a low-traffic metadata store, and its outage is far less dramatic than it sounds:
| While the Control Plane is down... | Status |
|---|---|
| Device telemetry, commands, live dashboards' data | ✅ Unaffected — pure Data Plane |
| Layer 1 rules, Layer 2 processors, Layer 3 ingestion | ✅ Unaffected |
| Already-issued NATS/Nebula credentials | ✅ Keep working — auth is verified by the cluster, not by PocketBase |
| Console login, entity management | ❌ Paused |
| Provisioning new Orgs / Things / credentials | ❌ Paused |
| Agent bootstrap and credential sync | ⏸️ Paused — the Agent retries, and an already-provisioned site keeps running on the credential and leaf config it holds |
Nothing in that bottom half is latency-critical. So instead of running an HA database topology to protect a metadata store, the platform's answer is aggressive backups plus a short, rehearsed restore path (§3–4). With scheduled native backups, S3 offsite copies, and filesystem snapshots, realistic time-to-recovery is minutes — which, for a service whose outage pauses provisioning but not production traffic, is the right trade.
2.1 Where the NATS server runs
The table above assumes the classic split: Control Plane in one process, nats-server in another. stone-age serve --nats runs the NATS server inside the Control Plane process instead, from the same nats.conf that nats export writes (Getting Started §3).
That is a real deployment option, not a development toy, but it changes the first row of that table. There are three rungs and you move between them by editing config:
| What runs | A Control Plane restart costs | Use when | |
|---|---|---|---|
| 1. Embedded | One process | A brief total bus outage | Small installs; the fabric can blink during an upgrade |
| 2. Embedded + external | Control Plane + one nats-server, clustered |
Devices reconnect; fabric stays up | Upgrade windows start to hurt |
| 3. Fully external | Control Plane + a NATS cluster | Nothing | HA, or scaling the bus independently |
At rung 1, the Data Plane's independence from the Control Plane is suspended — restarting stone-age restarts the bus. That is the whole trade, and it is why --nats is off by default.
At rung 2 the independence comes back for planned work, with one wrinkle worth knowing before you rely on it. Measured against a two-node cluster with the embedded node stopped:
| Core NATS pub/sub, device connections | Survive |
| JetStream KV reads and writes (R1) | Survive |
| JetStream management — create/delete stream, consumer, bucket | Stalls until the node returns |
Two nodes means RAFT quorum is two, so losing one leaves the JetStream meta group without a leader. What stops working is creating and modifying streams and buckets — Control Plane work, and the Control Plane is what is down. Telemetry keeps flowing throughout.
Two nodes is not high availability. It buys independence for planned upgrades and nothing more; losing the external node at rung 2 is worse than running at rung 1. Real fault tolerance needs three voting nodes — rung 3.
Devices must know both URLs. A device holding a single URL pointed at the embedded node still drops when it restarts. Give clients both, or the benefit disappears in practice.
Moving 2 → 3 is the one step that isn't purely additive; see §5.5. The reasoning behind all of this is recorded in ADR 0001.
3. Backups
Use the layers together: native scheduled backups as the authoritative artifact, S3 for offsite, pb (pb-cli) for scripting and rehearsal, and ZFS snapshots for instant local rollback.
3.1 Native scheduled backups
PocketBase — and therefore the platform binary — ships with backup support built in. A backup is a consistent zip of the entire pb_data directory, taken safely while the server runs.
Configure it as the SuperUser in the embedded admin UI (/_/ → Settings → Backups):
- Schedule — a cron expression (e.g.
0 2 * * *for nightly at 02:00). - Max kept — how many backups to retain before the oldest is pruned.
- Storage — local disk by default, or an S3-compatible bucket (endpoint, bucket, region, credentials). With S3 configured, every scheduled backup lands offsite automatically — no extra tooling.
This alone satisfies the baseline: nightly, consistent, offsite, auto-pruned.
3.2 Scripted backups with pb-cli
pb-cli (pb) is a generic PocketBase CLI that drives the same backup API from scripts — useful for pre-upgrade snapshots, extra offsite copies, and restore rehearsal. Backup operations require SuperUser auth.
# One-time setup
pb context create prod --url https://platform.acme.io
pb auth --collection _superusers
# On-demand, e.g. from cron or a pre-upgrade hook
pb backup create --name "pre-upgrade-$(date +%Y%m%d-%H%M)"
# Pull a copy off the platform host entirely
pb backup download "pre-upgrade-20260610-0900" /mnt/backup-vault/
# Prune: keep the five newest, delete the rest (careful!)
pb backup list --output json \
| jq -r 'sort_by(.modified) | reverse | .[5:] | .[].key' \
| xargs -I {} pb backup delete {} --force
pb also moves backups between environments (backup upload + backup restore), which is how you rehearse recovery and stage upgrades against real data — see §5.3.
3.3 Filesystem snapshots (ZFS)
We recommend running the Control Plane with pb_data on its own ZFS dataset with automatic snapshots (sanoid, zfs-auto-snapshot, or your distro's equivalent):
zfs create tank/stone-age
# point the binary at it: ./stone-age serve --dir /tank/stone-age/pb_data
- Near-zero cost, near-instant rollback. Frequent snapshots (every 5–15 minutes) cost almost nothing and
zfs rollbackrestores the whole directory in seconds — the fastest possible answer to "the upgrade went sideways" or "someone deleted the wrong org." - Replication for offsite.
zfs send | zfs recvto a second box gives you a warm standby of the data directory with no application awareness needed. - One caveat: a snapshot of a running database is crash-consistent, not application-consistent. SQLite in WAL mode recovers cleanly from that in practice, but the native backup zip remains the authoritative restore artifact — snapshots are the convenience layer on top, not a replacement.
3.4 What this routine does not cover
By design, per the table in §1: JetStream/KV contents (protect with replicas and mirrors at the NATS layer), your TSDB (its own tooling), and edge state (self-healing). One more worth keeping: a periodic stone pull workspace in git is a human-readable, diffable record of your tenant configuration — not a substitute for backups (it carries no secrets or identity material), but a fine complement for auditing and selective re-creation.
4. Recovery
Restore in place
For "bad change, wind it back" scenarios:
pb auth --collection _superusers
pb backup list
pb backup restore nightly-20260609 # confirms before acting; the server restarts itself
Or equivalently: admin UI → Settings → Backups → restore. Or, if the damage is filesystem-level and you're on ZFS: zfs rollback tank/stone-age@<snapshot> and restart the service.
Rebuild from nothing
Total host loss. You need: the platform binary (or the means to build it) and any backup artifact.
- Prepare the new host.
- Install the
stone-agebinary on it. - Recover
pb_data. Use whichever of these applies:- Restore the ZFS replica.
- Unzip a native backup into place.
- Start the binary empty, then run
pb backup uploadfollowed bypb backup restoreagainst it.
- Run
./stone-age servewith your existingconfig.yamlandSTONE_AGE_*env vars. - Re-point DNS and the reverse proxy at the new host.
The NATS cluster needs no changes — it kept running the whole time, and every credential it validates was signed by keys that are back in place. The restored Control Plane reconnects on the System Account and resumes propagating changes in real time, exactly as described in Architecture §2.
Verify after any restore
curl -s localhost:8090/api/ready | jqfirst. One request answers most of this list: whether the schema imported, whether an operator exists, whether the NATS server still trusts this database’s operator, and whether the binary is older than thepb_datayou just restored. Every warning or failure carries the command that fixes it. See Health & Metrics.- Console login works (Platform Operator user) and the NATS Status: Connected indicator is green.
- Create a throwaway Thing in a test org — confirms the provisioning hooks and the System Account connection end-to-end.
- Sites are attached again. There is no console screen for this; ask the hub over a tenant's own NATS connection — the site-connectivity widget recipe in Leaf Nodes §7, or
$SYS.REQ.ACCOUNT.PING.CONNZfrom thenatsCLI. That reading comes from the hub, so it is a live fact rather than a replayed one.
Rehearse this. A backup you've never restored is a hypothesis, not a backup. The pb migration flow in §5.3 doubles as a restore drill — do it on a schedule, not just before upgrades.
5. Upgrades
5.1 How upgrades work
The platform binary embeds its schema and runs migrations automatically: replace the binary, restart, and any pending schema migrations apply at startup. Because the UI and schema are compiled into the same artifact, the Control Plane upgrades atomically — there is no window where the UI, API, and schema disagree.
A migration file is what makes a schema or API-rule change reach your deployment. The embedded
schema.jsonis applied when a database is first created; an existingpb_datakeeps its collections and its rules until amigrations/schema_update_*.goaccompanies the change. So "the fix is in the new binary" is only true if the release shipped the migration — which matters most for authorization changes, since the API rules are the platform's permission layer. See Authorization §7.
The procedure:
- Read the release notes. Pre-1.0, breaking changes can occur. They are called out per release and have so far been minimal.
- Back up. Run
pb backup create --name "pre-upgrade-vX.Y.Z", or take a ZFS snapshot, or both. Thirty seconds of discipline that makes the rollback below trivial. - Swap the binary.
- Restart the service. Migrations run, then the server comes up.
- Verify — the same checklist as §4.
If it went wrong, roll back in this order:
- Stop the service.
- Put the previous binary back.
- Restore the pre-upgrade backup, or run
zfs rollback. - Start the service.
Migrations are forward-only. Rollback is always old binary + restored data, never the new binary against old data.
5.2 Pre-1.0 expectations
Until 1.0, treat minor versions as potentially breaking and pin what you deploy. In practice breaking changes have been small and migration-handled, but the contract is explicit: read the notes, back up first. After 1.0, standard semver discipline applies.
5.3 Rehearse on staging with real data
pb makes a realistic dress rehearsal cheap:
# Copy production state to staging
pb context select production && pb auth --collection _superusers
pb backup create --name "rehearsal-source"
pb backup download rehearsal-source ./rehearsal.zip
pb context select staging && pb auth --collection _superusers
pb backup upload ./rehearsal.zip --name "from-prod"
pb backup restore from-prod
# Now run the NEW binary against staging and watch the migrations apply
If the upgrade misbehaves, you found out on staging — against your actual schema and data shape, not a toy fixture.
5.4 Upgrading the other components
The Control Plane is the only component with a database and migrations. Everything else upgrades by binary swap, in any order, because the interfaces between components are protocols, not shared code (§6):
rule-router, stream processors, Telegraf — restart with the new binary; they reconnect to NATS and resume. Durable state is in NATS; what they hold in memory (rule-router's throttle windows, for one) is lost on restart.- Agents — swap and restart; designed around reconnection and fail-soft behaviour. On a site gateway, restarting the agent bounces the leaf server too if
nats.server_configis set; leave it unset and a separately supervisednats-serverkeeps the bus up across the upgrade, which is usually what you want on a live site. - NATS and Nebula — stock upstream upgrade procedures; the platform places no constraints beyond theirs (see §6).
5.5 Moving the NATS server out of the Control Plane
Going from serve --nats to a standalone nats-server (§2.1). Almost all of it is free: the NATS Operator JWT, every account, and every user credential live in the Control Plane database and are re-derived, not migrated. Devices keep the credentials they already hold.
JetStream data is the exception. Streams, consumers, and the KV buckets holding Digital Twin state live in the embedded server's store directory. An R1 stream on the embedded node dies with that node. Plan for it before you start.
The easiest version of this migration is the one where there is nothing to move:
If you expect to grow out of embedded mode, don't put JetStream on the embedded node in the first place. Add the external node early, keep
jetstream: {}out of the embedded config, and this section becomes a config edit.
Otherwise, replicate before you drain:
-
Add the external node.
- Give both NATS configs a matching
clusterblock. Use the samenamein both, and point each one's routes at the other. - Restart both servers.
- Confirm the route formed.
nats server listmust show two servers in the cluster.
- Give both NATS configs a matching
-
Raise replicas on anything you intend to keep. Nothing is safe to drain until it is replicated.
- Set
replicas: 2on every stream on the embedded node:nats stream update <name> --replicas 2. - Set
replicas: 2on every KV bucket on the embedded node. - Confirm each one reports the new peer as current, not catching up:
nats kv status <bucket>.
- Set
-
Drain the embedded node. This moves JetStream leadership off the node, rather than losing it abruptly.
- Step down its raft leadership:
nats server raft step-down.
- Step down its raft leadership:
-
Stop the embedded server.
- Drop
--nats, or setnats.embedded: false. - Point
nats.server_urlat the external node. - Restart the Control Plane.
- Drop
-
Return replicas to their intended value. A single remaining node cannot hold
replicas: 2. Do one of the following:- Set the replica count back to
1on every stream and bucket you changed in step 2. - Add the third node now, then set the replica count to
3.
- Set the replica count back to
-
Verify before you delete anything.
- Confirm devices reconnect and twins update.
- Confirm
nats stream reportshows every stream present, with the expected message counts. - Remove the old store directory. Do this last, and only once the two checks above pass.
Rehearse this on a copy first (§5.3). Steps 2 and 3 are where data is lost if the peer was not actually current, and "it looked fine" is not the same as a message count that matches.
6. Component Version Compatibility
Stone-Age.io is a set of independent binaries, so the compatibility question is really: what does each component actually depend on? The answer is deliberately narrow — components couple to protocols and the collections schema, never to each other's code.
| Component | Depends on | Compatibility notes |
|---|---|---|
| Stone Age Console (embedded UI) | Ships inside the stone-age binary |
Always in lockstep with the schema by construction. No version skew is possible. |
stone CLI |
PocketBase REST API + the platform's collections schema; NATS protocol | The REST API is stable upstream PocketBase. The collections schema is the platform's own contract — additive changes don't break the CLI; pre-1.0 breaking schema changes are flagged in release notes. |
| Agent | PocketBase auth API (bootstrap only) + NATS protocol | After bootstrap it's a pure NATS client. |
rule-router |
NATS subjects + KV only | Knows nothing about PocketBase. Versioned independently. |
| Stream processors, Telegraf, TSDB, Grafana/Perses | NATS subjects only | Fully platform-agnostic. The subject contract (Thing Types) is the only interface. |
nats-server |
The exported NATS Operator and resolver config (Getting Started §3) | Any modern NATS 2.x with JetStream and JWT/operator-mode auth. Follow upstream support guidance. |
nebula |
Certificates issued by the org CAs | Stock upstream; the platform only mints standard Nebula certs and configs. |
Practical guidance:
- Upgrade the Control Plane first when a release touches the schema — clients (
stone, the Agent) tolerate additive changes, and the release notes call out anything that isn't. - Layer 1–3 components don't care about platform releases at all. Their contract is the subject namespace, which is yours to keep stable — see Connectivity.
- Pre-1.0, pin versions across
stone-age,stone, and the Agent, and move them together when the notes mention schema changes. The Agent releases from its own repository on its own tags, so its version does not track the platform's. Post-1.0, additive-only within a major version is the rule.
7. Production Checklist
A condensed pre-flight list for taking a deployment to production:
- [ ] TLS everywhere outward-facing — HTTPS in front of the Control Plane;
wss://on the NATS WebSocket listener; TLS on client/leaf ports per the NATS docs. - [ ] Backups scheduled (admin UI cron) with S3 offsite configured — and a restore actually rehearsed (§4).
- [ ] The readiness probe wired into whatever runs the process —
GET /api/ready, which returns503only when something is genuinely broken and200for warnings, so it is safe to put in front of a load balancer. Point your scraper atGET /metricswhile you are there. Both are unauthenticated by default and neither needs a NATS connection or a session;metrics.tokencloses the second one, and a proxy closes either (Health & Metrics). - [ ]
pb_dataon its own dataset/volume, ideally ZFS with automatic snapshots (§3.3). - [ ] App-settings encryption enabled via
--encryptionEnv(Configuration §4). This covers SMTP, S3 and OAuth2 secrets — and nothing else. - [ ] Column encryption set separately —
nats.encryption_keyandnebula.encryption_keyinconfig.yamlare what encrypt the minting keys: the NATS operator, account and user seeds and signing keys, and the Nebula CA and host private keys. They are empty by default, they cannot be applied retroactively to rows that already exist, and losing a key loses the material it protected. A checklist that ticks--encryptionEnvand stops has left every tenant’s CA private key in plaintext. Neither covers issued credentials —creds_fileandconfig_yamlstay plaintext by construction, so disk encryption and encrypted backups carry that part (Configuration §2.2). - [ ] Audit retention configured deliberately —
audit.retentiondefaults keep everything forever (Configuration §2). Remember the log is Platform-Operator-only: no tenant role can read it. The plain "who changed this, and when" question is now self-served by the tenant-facingactivityfeed; what still reaches you is any request needing the values a change carried (Authorization §5). - [ ] NATS account limits reviewed — the shipped defaults are 100 connections, 5000 subscriptions, 5 GiB of JetStream disk and 64 MiB of JetStream memory per organization. Size disk to what each plan is sold with, keep memory small, and set connections well above real load. They are stamped at provisioning, so decide before creating tenants: an existing account changes only by a Platform Operator editing its record (Configuration §2).
- [ ] NATS clustered (3+ nodes) with
replicas: 3on JetStream streams and KV buckets that matter. - [ ] Credential expiry reviewed — the NATS Users and Nebula Hosts lists flag anything expiring within 30 days, and anything already expired. Nebula host certificates always carry an expiry (
validity_years); NATS user JWTs carry one only if it was set. Expired device credentials fail silently and, because a fleet is usually provisioned in a batch, they tend to fail together. Regenerate before the date, not after the outage. Nebula’s expiry is the worse one, because Nebula is the out-of-band path: reissuing a host certificate needs the Control Plane you were trying to reach over the mesh in the first place, so an expired fleet takes away the route you would have used to fix it. -
[ ] Nebula expiry alerted, not just visible — the list badges above only help someone already looking at the right screen, and a CA minted with a ten-year validity lapses long after everyone who knew about it stopped thinking about it. The Control Plane reports Nebula certificate expiry without anyone logging in:
GET /api/readycarries anebula_cert_expirycheck. It warns, and deliberately never fails — readiness failing means "stop sending this node traffic", and a lapsed device certificate is no reason to take the console out of a load balancer.-
GET /metricspublishesstone_age_certificate_expiry_seconds{kind}(nebula_ca/nebula_host), the soonest expiry of that kind as an absolute Unix timestamp, alongsidestone_age_certificates,_expiredand_expiringcounts. Alert relative to now rather than on a stored countdown, so the horizon lives in the alert:stone_age_certificate_expiry_seconds{kind="nebula_host"} - time() < 30 * 86400 stone_age_certificate_expiry_seconds{kind="nebula_ca"} - time() < 90 * 86400
Host certificates count only
active = truerows, so a decommissioned device’s lapsed certificate does not page anyone. There is no per-organization label:/metricsis open by default, and a tenant name beside a certificate inventory is free reconnaissance. Set an alert on the CA series in particular — every host certificate chains to it.Give the CA a wider horizon than a host: 90 days, not 30. A host certificate is reissued in a moment; a CA can only be rotated, which is a staged procedure with a wait in the middle of it, and 30 days is less than that procedure comfortably needs. The console,
_expiringand the readiness check all use 90 days for a CA and 30 for a host for this reason — hence the two expressions above rather than one. - [ ] A decommissioning path agreed — know before you need it that clearingactiveon a Thing is the control that actually cuts a device off, that it applies to a site gateway exactly as to any other device, and what it does: new logins refused, every issued session token killed, the linked NATS identity suspended (revoked, nothing reissued), and the linked Nebula host blocklisted — which takes effect as peers pick up their next config, since Nebula has no CRL. Reactivating issues a new.credsthe device must be given (Authorization §4.2). Deactivate, never delete: deleting a Thing cascades to neither identity, so its credential and certificate stay live. - [ ] A tenant-suspension path agreed — clearingactiveon an organization withdraws its NATS account: the account claim is deleted, so every device, edge agent and browser in that tenant disconnects at once, and it stays withdrawn across restarts. It is Platform Operator only and reversible — no credential is revoked or re-minted, so re-ticking the box restores the account and every existing.credsconnects again. It is deliberately narrow: it does not touch Nebula, does not lock the console, and does not end anyone's session, so a suspended tenant still signs in and reads what it owns. The operator and system organizations refuse it. - [ ] SuperUser reserved for infrastructure work; day-to-day administration through a Platform Operator user (Getting Started §2). - [ ] Least-privilege role review — walk each org's memberships and confirm nobody holds more than they need.adminis not a junior grant: it is identical toownerin every API rule, including every credential-bearing collection. Most humans wantmember(Authorization). - [ ]./scripts/test-authz.shgreen on the exact commit you're deploying — the API rules are the platform's tenancy enforcement, and the suite is the only thing that checks them against a live server. If the release changed a rule, confirm it also shipped a migration (§5.1). - [ ] Astone pullworkspace in git for reviewable, diffable tenant configuration (Stone CLI §5).
8. Where to Go Next
- First-time setup the checklist assumes: Getting Started.
- Config keys referenced above: Configuration Reference.
- Roles, API rules, and the audit-log boundary: Authorization & Roles.
- The plane split that shapes this whole page: Architecture and Platform Layers.
- Edge resilience during outages: Leaf Nodes.
- The GitOps workspace as a config audit trail: Stone CLI.