Skip to content

Configuration Reference

The Stone-Age.io platform binary (stone-age) is configured through three sources, evaluated in priority order:

  1. Environment variables prefixed with STONE_AGE_ — override everything.
  2. The --config /path/to/config.yaml flag — load a specific config file.
  3. A config.yaml discovered automatically — looked up in ./ first, then /etc/stone-age/.
  4. Hardcoded defaults — used for any key not set above.

The binary runs without a config file at all (defaults are sensible), but most production deployments commit a config.yaml to source control and override hostnames or secrets via environment variables.


1. The config.yaml File

tenancy:
  organizations_collection: "organizations"
  memberships_collection: "memberships"
  invites_collection: "invites"
  invite_expiry_days: 7

nats:
  account_collection_name: "nats_accounts"
  user_collection_name: "nats_users"
  role_collection_name: "nats_roles"
  operator_name: "stone-age.io"
  server_url: "nats://localhost:4222"       # where THIS PROCESS dials
  websocket_urls: []                        # where a BROWSER dials. Not the same thing.
  leaf_url: ""                              # where an EDGE box's leaf remote dials. Not the same thing either.
  jetstream_domain: "hub"                   # this hub's own JetStream domain
  encryption_key: ""                        # encrypts the minting keys at rest (§2.2)
  managed_export_subject: "helpdesk.>"
  log_to_console: false
  default_limits:
    max_connections: 100
    max_subscriptions: 5000
    max_payload: 1048576                    # 1 MiB, nats-server's own default
    max_jetstream_disk_storage: 5368709120  # 5 GiB
    max_jetstream_memory_storage: 67108864  # 64 MiB
  export_collection_name: "nats_account_exports"
  import_collection_name: "nats_account_imports"
  embedded: false                             # run NATS inside this process
  embedded_config: "./nats-config/nats.conf"  # the config --nats loads

nebula:
  ca_collection_name: "nebula_ca"
  network_collection_name: "nebula_networks"
  host_collection_name: "nebula_hosts"
  log_to_console: false
  default_ca_validity_years: 10
  encryption_key: ""        # encrypts the CA and host private_key columns at rest

audit:
  collection_name: "audit_logs"
  log_to_console: false
  retention:
    max_age: ""             # Go duration string, e.g. "720h" for 30 days. "" disables.
    max_records: 0          # 0 disables record-count retention.
    interval: "0 2 * * *"   # Cron schedule for the cleanup job.

readiness:
  interval: 15s
  timeout: 5s

metrics:
  enabled: true
  token: ""   # empty = open, the default

branding:
  dir: ""     # a host directory of overrides; "" uses the embedded defaults

2. Section Reference

tenancy

Controls the multi-tenancy collections: organizations, memberships and invitations. This is platform code (hooks/org_membership.go, hooks/invites.go); it was the pb-tenancy library until that was absorbed.

Key Type Default Purpose
organizations_collection string "organizations" Name of the Orgs collection. Override only if migrating from a non-default schema.
memberships_collection string "memberships" Name of the User↔Org link collection.
invites_collection string "invites" Name of the pending-invite collection.
invite_expiry_days int 7 How long an outstanding invite token remains valid.

There is no tenancy.log_to_console any more. The one thing it did was silence invitation-email failures, which now go through the application logger unconditionally. An old config that still sets it is harmless: the key is not read.

nats

Controls the pb-nats library: NATS account/user/role provisioning, exports/imports management, and the System Account connection.

Key Type Default Purpose
account_collection_name string "nats_accounts" NATS Account collection name.
user_collection_name string "nats_users" NATS User collection name.
role_collection_name string "nats_roles" NATS Role collection name.
operator_name string "stone-age.io" The NATS Operator name stamped into the NATS Operator JWT at first run.
server_url string "nats://localhost:4222" Where the Control Plane connects to NATS as a System Account client. Not the browser address — see §2.1.
websocket_urls string list [] The WebSocket addresses a browser dials, served to the console at runtime by GET /api/client-config. See §2.1.
leaf_url string "" Where an edge box's leaf remote dials this hub — usually port 7422, and often a different hostname from either address above. Served to gateways as hub_leaf_url by GET /api/me/leaf-config. Empty means no edge box can bootstrap. Must start with nats-leaf:// or tls://; the platform refuses to start otherwise.
jetstream_domain string "hub" This hub's own JetStream domain — what an edge addresses the hub as across the link ($JS.<domain>.API). Served to gateways as hub_domain. Must match the jetstream { domain } in the hub's nats.conf. Agents cache it, so changing it later reaches a site only when that site re-runs agent -leaf-config.
encryption_key string "" 32-character key encrypting the NATS minting keys at rest: the operator, account and user seeds and private keys, and the signing keys. Not issued credentials — see §2.2. Empty means plaintext in SQLite.
managed_export_subject string "helpdesk.>" The subject subtree a managed organization exports into the provider's hub account. The matching hub-side import remaps it to <subtree>.<organization code>.>, so the tenant token is baked into the NATS-Operator-signed account JWT and provenance is unforgeable — which means a managed organization needs a code before its export routes anywhere. Must end in .>; the platform refuses to start otherwise. See ADR 0002.
log_to_console bool false Verbose NATS-library logging.
default_limits.max_connections int 100 Max connections for new Org accounts. Headroom, not a fence — see the note below.
default_limits.max_subscriptions int 5000 Max subscriptions for new Org accounts.
default_limits.max_payload int 1048576 Max payload bytes for new Org accounts (1 MiB, nats-server's own default).
default_limits.max_jetstream_disk_storage int 5368709120 JetStream file storage for new Org accounts (5 GiB). What a plan is sold with.
default_limits.max_jetstream_memory_storage int 67108864 JetStream memory storage for new Org accounts (64 MiB). A blast-radius limit — keep it small and the same across plans.
export_collection_name string "nats_account_exports" Account-level Export collection name. See Connectivity §1.
import_collection_name string "nats_account_imports" Account-level Import collection name.
embedded bool false Run a NATS server inside the Control Plane process. Acted on by serve only. Equivalent to --nats.
embedded_config string "./nats-config/nats.conf" The nats.conf that embedded loads — the file nats export writes. Equivalent to --nats-config.

The limits are stamped into each new organization's signed account JWT, at provisioning time only. Changing a number here does not reach an account that already exists; that account's nats_accounts record has to be edited, and that record's update rule is Platform Operator only — the limits are the resource envelope a tenant was sold, so raising them is not a tenant action.

-1 means unlimited. 0 on either JetStream field disables JetStream for the account entirely — and the Digital Twin with it, since the twin is two KV buckets. Never use 0 to mean "no limit".

The two storage numbers do different jobs. Disk tracks what a plan is sold with. Memory is shared: a memory-backed stream competes for the RAM every other tenant on the box needs, so it is the one breach not confined to the account that caused it. Connections and subscriptions are headroom. Their breach mode is a device that cannot connect or a subscription that quietly fails, so set them well above any modelled load and alert as you approach them. Clients behind an edge leaf node do not count toward max_connections: the hub sees one leaf connection per site, so the count is browsers, the stone CLI, services, and devices that dial the hub directly.

embedded does not configure the NATS server. The nats.conf does — ports, JetStream, WebSockets, clustering, TLS. embedded only decides whether the Control Plane runs that config itself or leaves it to a separate nats-server. The two produce an identical server, which is why moving between them is a config change rather than a migration. See Operations §2.1.

One thing must agree across the two files: the port. server_url is where the Control Plane dials, and the port in nats.conf is where the server listens. --nats refuses to start when they differ, because nothing in the process could then reach the server it just started.

2.1 server_url and websocket_urls are different addresses

This is the configuration mistake with the least helpful symptom, so it is worth stating plainly.

Who dials it Typical value
nats.server_url this process, to publish account claims nats://nats:4222 inside a container
nats.websocket_urls a browser, to hold a live console session wss://bus.example.com:9222

Different port, often a different host, and never derive one from the other — a Control Plane publishing to nats://nats:4222 inside a container says nothing about what a browser on someone’s laptop can reach.

websocket_urls is served to the console at runtime by GET /api/client-config rather than compiled in, deliberately: the console is embedded in the binary, so a build-time constant would mean a frontend rebuild per provider — the same problem branding.dir exists to avoid. The endpoint is authenticated; there is no pre-login need for the bus address.

The console resolves the address in three tiers: a per-device override in localStorage, then this key, then a compiled-in ws://localhost:9222. Four rules follow from that:

  • The device override replaces this list; the two are never merged. The NATS client shuffles its server list by default, so a merged list is a pool picked at random rather than a priority order — and the reason a device overrides is to reach its local leaf node instead of the hub. Those are different JetStream domains holding different data under the same bucket names, so merging would make which dataset you are looking at a coin flip per reconnect.
  • Multiple entries mean one cluster. Peers, not failover order. Do not list a hub URL and a leaf URL together.
  • There is no JetStream domain setting, on purpose. The console passes no domain, so $JS.API resolves to the JetStream of whichever server was dialed — hub URL gives the hub, leaf URL gives that leaf’s own domain, which is its Thing's code. The URL already selects the domain, and a separate knob could only disagree with it, failing as an empty bucket list with no diagnosis.
  • An HTTPS page cannot open ws://. Browsers block it outright, so the settings form rejects a plaintext URL rather than saving one that can never connect.

In an environment variable, separate multiple URLs with spaces, not commas — viper splits that value on whitespace, and the platform rejects a comma-joined entry at startup rather than treating it as one malformed URL:

STONE_AGE_NATS_WEBSOCKET_URLS="wss://a.example.com:9222 wss://b.example.com:9222"

2.2 The encryption keys

nats.encryption_key and nebula.encryption_key encrypt the secret columns at rest: the NATS operator, account and user seeds and private keys and the account signing keys, and the Nebula CA and host private_key columns. Both default to empty, which means those values sit in the SQLite file in plaintext.

What these keys do not cover: issued credentials

The keys protect the material needed to mint identities. They do not protect credentials already issued, and cannot:

  • nats_users.creds_file contains the user seed by construction, and the browser reads it straight from the API to open its own NATS connection.
  • nebula_hosts.config_yaml embeds the host key inline, because Nebula requires it there.

So a stolen database with the key held elsewhere yields no ability to mint new identities, and every credential already issued. Carry that threat with disk encryption, encrypted backups and access control on the host. The readiness check says the same thing when encryption is on: "minting keys; issued credentials are plaintext by construction". The platform repository's SECURITY.md describes what a stolen database does and does not yield.

Set these before creating anything real, and back the keys up separately

A row written with a key cannot be read back without it. There is no recovery path: losing the key loses every seed and private key it protected, which means re-provisioning every NATS identity and re-issuing every Nebula certificate in every affected organization. Supply them through STONE_AGE_NATS_ENCRYPTION_KEY and STONE_AGE_NEBULA_ENCRYPTION_KEY rather than committing them to config.yaml.

Equally, turning encryption on for a database that already has rows does not retroactively encrypt them, and turning it off does not decrypt what is already encrypted. Decide at install time.

Generate 32 characters and keep them somewhere that is not the backup of the database they protect:

openssl rand -hex 16    # 32 characters

This is not the same thing as --encryptionEnv

PocketBase’s --encryptionEnv flag encrypts app settings — SMTP passwords, S3 credentials, OAuth2 secrets. It does not touch the NATS and Nebula columns above, because those collections belong to this platform rather than to PocketBase.

A production checklist that ticks --encryptionEnv and stops has left every tenant’s CA private key in plaintext. Both are needed, and they are configured in different places: --encryptionEnv is a CLI flag, these are config.yaml keys.

nebula

Controls the pb-nebula library: CA, network, and host certificate management.

Key Type Default Purpose
ca_collection_name string "nebula_ca" Certificate Authority collection name.
network_collection_name string "nebula_networks" Per-CA network collection name.
host_collection_name string "nebula_hosts" Host certificate collection name.
log_to_console bool false Verbose Nebula-library logging.
default_ca_validity_years int 10 Default validity for newly-generated org CAs.
encryption_key string "" 32-character key encrypting the Nebula CA and host private_key columns at rest. Empty means plaintext in SQLite. The generated config_yaml still embeds the host key in plaintext — see §2.2.

audit

Controls the pb-audit library: audit logging of create, update, delete, and auth events.

Key Type Default Purpose
collection_name string "audit_logs" Where audit records are written.
log_to_console bool false Mirror audit events to stdout. The older spelling audit.log_console is still honoured.
retention.max_age string "" Go duration string (e.g. "720h" = 30 days). Empty disables age-based pruning.
retention.max_records int 0 Max records to keep. 0 disables count-based pruning.
retention.interval string "0 2 * * *" Cron expression for the retention sweep job.

What the audit log holds: every create, update, delete and auth event records changed_fields — the names of the fields that moved. Full before/after values are kept only for an allowlist of collections (auditSnapshotCollections in the platform's main.go: organizations, memberships, users, the inventory and type collections, NATS roles, Nebula networks and email templates). Everything credential-bearing — NATS users and accounts, Nebula hosts and CAs, invites, exports and imports — records field names and no values, so the log is not an archive of every credential ever minted.

Who can read the audit log: audit_logs list and view are @request.auth.is_operator = true. Platform Operators and SuperUsers only — no tenant role, including owner, can read it, and the console's /audit route is gated on the same flag to match. A tenant admin cannot self-serve an audit export. The request has to go through a Platform Operator. What a tenant can read for itself is the activity feed — actor, action, record, timestamp, and no values — which is a separate collection with its own rules and no retention setting here. See Authorization §5.

branding

Key Type Default Purpose
dir string "" A host directory whose branding.json, logo.svg and theme.css override the embedded defaults, served at /branding/*. Empty disables the overlay.

The point of the overlay is that re-skinning the console needs no frontend rebuild — the console is embedded in the binary, so a compiled-in brand would mean one build per provider. Missing files fall back individually. A starting template ships in branding.example/ in the repository.

This brand is the operator's, and an Organization's own logo does not replace it. The sidebar mark and the login screen resolve to branding and nothing else — it is the one fixed landmark, and the brand row is also the link to /, so letting a tenant logo win there made it change on every organization switch, directly above the switcher that had just done it, captioned with the operator's appName. A tenant's logo appears in the org switcher instead, where switching it is the control. The QR label prints neither: a label belongs to one organization, so it carries that organization's code and the record's code and name, and no operator brand.

readiness

The background prober behind GET /api/ready.

Key Type Default Purpose
interval string "15s" How often the checks re-run. The endpoint serves the last snapshot, so this is also how stale an answer can be.
timeout string "5s" Deadline for one full probe. Checks run concurrently, so this bounds the slowest one, not their sum.

The endpoint is always registered — there is no key to disable it. Readiness is a contract with your orchestrator, and a deployment that had quietly turned it off is one nobody could explain later. Closing it to the outside is a proxy's job.

metrics

The Prometheus exposition at GET /metrics.

Key Type Default Purpose
enabled bool true Register the route at all.
token string "" Shared secret. Empty means open, which is the default.

token is accepted as Authorization: Bearer <token> or HTTP Basic with any username, which between them cover every scraper in use. It deliberately does not use PocketBase auth: those tokens are JWTs that expire and no scraper has a refresh flow, so it would take a custom sidecar to read a standard format.

Nothing here is labelled per organization. The Control Plane holds the NATS operator and $SYS and has no credential inside a tenant's account, so per-org data is reduced to row counts — and a tenant label on a row count is a customer name attached to an inventory number, on an endpoint that is open by default.

See Health & Metrics for the checks, the metric names and the alert expressions.

observability (the Agent, not this file)

The edge agent's own /ready and /metrics, in the Agent's config.yaml — not in the Control Plane's. On by default, on loopback.

Key Type Default Purpose
addr string "127.0.0.1:9100" Listen address. Set it empty to create no listener at all; the checks still run and still log either way.
metrics_token string "" As metrics.token above. Set this before moving addr off loopback.
interval string "15s" Probe interval.

Two things to know about that default. It is not a gateway setting — every Agent serves these, because the reason to answer locally is that cmd.health travels over NATS and NATS is the link that breaks. And 9100 is node_exporter's own default port on Linux and FreeBSD, so on a box running both, move one of them. (windows_exporter uses 9182, so Windows is unaffected.)

This is where per-site health is visible in detail, and it keeps answering with the WAN down — which is exactly when you want it. The Control Plane cannot report it: it holds the NATS operator and the $SYS account and has no credential inside any organization's account. (Whether a site's leaf node is attached is a separate, cheaper question, which a tenant answers by asking the hub over its own NATS connection — a dashboard widget, not a console screen.) See Leaf Nodes.


3. Environment Variable Overrides

Every key in config.yaml has an STONE_AGE_* environment variable equivalent. The mapping rule is:

  • Prefix with STONE_AGE_.
  • Replace . (YAML path separator) with _.
  • Uppercase the whole thing.

Examples:

# Override NATS server URL (highest-priority source)
export STONE_AGE_NATS_SERVER_URL="nats://nats.internal:4222"

# Mirror audit events to stdout (useful in development)
export STONE_AGE_AUDIT_LOG_TO_CONSOLE=true

# Give new orgs a bigger JetStream disk allowance (10 GiB)
export STONE_AGE_NATS_DEFAULT_LIMITS_MAX_JETSTREAM_DISK_STORAGE=10737418240

# Keep the encryption keys out of config.yaml
export STONE_AGE_NATS_ENCRYPTION_KEY="$(cat /run/secrets/nats_key)"

# Tighten the invite window
export STONE_AGE_TENANCY_INVITE_EXPIRY_DAYS=2

Environment overrides are the recommended way to inject secrets and per-environment values (dev/staging/prod) without forking the YAML file.


4. CLI Flags

Platform flags

Flag Default Purpose
--config <path> — Load a specific config.yaml.
--nats false serve only: run a NATS server in this process. Sets nats.embedded.
--nats-config <path> ./nats-config/nats.conf serve only: which nats.conf --nats loads. Sets nats.embedded_config.

--nats and --nats-config are registered on the root command, so other subcommands accept them, but only serve acts on them — the rest open the database directly and have no bus to talk to.

PocketBase flags

Because the platform binary embeds PocketBase, the standard PocketBase CLI flags are also available — most notably:

  • --dir <path> — override the data directory (default ./pb_data).
  • --dev — enable verbose logging and SQL statement printing.
  • --encryptionEnv <name> — name of an env var holding a 32-character key used to encrypt app settings at rest.
  • --queryTimeout <seconds> — default SELECT query timeout.

These apply uniformly to all subcommands (serve, migrate, bootstrap, nats export, superuser upsert).


5. Operational Notes

  • Changing operator_name after first run is not safe — the NATS Operator JWT is generated once at first SuperUser creation. Renaming would orphan the existing identity hierarchy.
  • server_url is for the Control Plane’s own System Account connection. The address a browser dials is websocket_urls, which is a deployment default here and can be overridden per device in the console’s Settings page. See §2.1 — conflating the two is the configuration mistake with the least helpful symptom.
  • Set the encryption keys at install time. nats.encryption_key and nebula.encryption_key cannot be introduced retroactively for rows that already exist, and losing one loses the material it protected. They are also not what --encryptionEnv covers. See §2.2.
  • Audit retention runs on a schedule, not on every write. A misconfigured interval will just delay cleanup, not break ingestion.
  • The schema is embedded in the binary, not loaded from disk. Updating the embedded schema.json is only half the change: it reaches freshly-created databases only. To change the schema — or an API rule — on an existing deployment, the platform needs a new migrations/schema_update_*.go file, which runs at startup. This is the most common way a rule fix fails to ship. See Authorization §7 and Operations §5.1.
  • API rules are the platform's authorization layer. Nothing in config.yaml grants or restricts access; the rules in the embedded schema do all of it, with one deliberate hook beside them that refuses a relation pointing into another organization. See Authorization & Roles.

6. Where to Go Next