Self-hosting
Operations & troubleshooting
Run self-hosted Keyright in production: scheduled jobs, observability, security hardening, scaling and load balancing, backup and restore, upgrades, fully air-gapped operation, and a troubleshooting/FAQ for the things that come up on a first deploy.
Everything you need after the instance is up: keeping it healthy, scaling it, backing it up, upgrading it, running it air-gapped, and fixing what didn’t come up green. If you’re still standing it up, start with the deployment guide.
Operations
Scheduled jobs
The daily expiry/seat scan and webhook delivery run on in-process timers. On hosts that idle background timers (some serverless platforms), drive them from an external scheduler instead — set KEYRIGHT_CRON_SECRET and call:
curl -X POST -H "X-Cron-Secret: $SECRET" $BASE/internal/notifications/run
curl -X POST -H "X-Cron-Secret: $SECRET" $BASE/internal/webhooks/deliver
Observability
- Liveness:
GET /health(static 200 — never fails a node out over a transient DB blip). Readiness:GET /health/ready(also checks DB connectivity — 503 when unreachable). - Logs: structured JSON to stdout (Docker/K8s drivers, or the Windows Event Log via ANCM). Watch startup lines confirming migrations applied and the self-license status.
- Monitor: app
/health, PostgreSQL availability/connections/disk, and certificate expiry on the load balancer.
Diagnostics (2.2.0)
-
GET /admin/self-diagnostics(admin token) — a one-call health snapshot for triage: version, database reachability + pending migrations, self-license phase, enabled features, whether SMTP/alerts are configured, and uptime. Like/admin/self-license, it stays reachable even when the instance is gated, so you can diagnose a licensing stop.{ "version": "2.2.0", "uptimeSeconds": 3600, "database": { "reachable": true, "error": null, "pendingMigrations": 0 }, "selfLicense": { "status": "valid", "phase": "active", "isTrial": false, "daysUntilExpiry": 320, "runtimeBlocked": false }, "features": { "operatorMode": false, "sso": true, "whitelabel": true, "airgap": true }, "smtpConfigured": true, "alerts": { "rulesEnabled": 3, "errorsAlerted24h": 0 } } -
Error correlation ids — every unhandled server error (HTTP 500) is logged once with a short id and returns that id to the caller in the
X-Keyright-Error-Idresponse header (and anerrorIdfield in the body). When a customer reports an error, that id maps their report to exactly one log line.
Email alerts (2.2.0)
Get emailed when things happen — configured in the dashboard → Settings → Email alerts (or via the API):
- SMTP — set your own SMTP host/port/user/password/from in the dashboard (
POST /admin/alerts/smtp); the password is stored KEK-encrypted and is write-only (never returned). If you don’t set it, alerts fall back to theKEYRIGHT_SMTP_*env config. - Per-event rules — enable exactly the events you want and give each its own recipient list (
POST /admin/alerts/rules), so different events route to different addresses. A “send test email” button verifies your settings. - Events include customer-facing (license expiring/expired, seat-limit reached, revoked) and, importantly for self-hosted, your own instance’s license approaching expiry, entering grace, or gating — so a fail-closed instance never stops unexpectedly — plus system events (server error, database error, scheduled-job failure). Recurring conditions are de-duplicated so they email once per window, not on every scan.
Security hardening
- Keep
KEYRIGHT_ADMIN_TOKEN,KEYRIGHT_KEK, andKEYRIGHT_SIGNING_KEYin a secrets manager — never in an image or source control. - Expose only
:443(via the load balancer) publicly./admin/*and/internal/*should be reachable only by you/your network; the client/v1/*path is the only one your customers need. - Restrict database network access to the app nodes, and require TLS to the database.
Scaling & load balancing
The app is CPU-light and stateless; scale is about redundancy and database capacity, not app throughput.
| Scale | Active licenses | App nodes | PostgreSQL | Load balancer |
|---|---|---|---|---|
| Evaluation | any | 1 (0.5 vCPU / 512 MB) | bundled | none |
| Small | up to ~50k | 1 (1 vCPU / 1 GB) | 2 vCPU / 2 GB managed | optional TLS proxy |
| Standard HA | up to ~500k | 2–3 (1 vCPU / 1 GB) | HA (primary+standby) | yes |
| Large | 1M+ | 3–5+ autoscale | 4–8 vCPU / 16 GB+ HA + PgBouncer | yes |
Run 2+ app nodes behind any HTTP load balancer. Because the app is stateless and clients verify offline, no session affinity is needed — send any request to any node. Terminate TLS at the balancer and forward X-Forwarded-For (real client IP) and X-Forwarded-Proto (so the dashboard/portal generate https URLs); the app honours both. Concurrency is safe across nodes — credits and seats use serializable check-and-deduct, so no two nodes can double-spend a balance or a seat.
Backup & restore
Everything is in PostgreSQL — standard pg_dump / pg_restore. Also back up KEYRIGHT_KEK and KEYRIGHT_SIGNING_KEY separately and securely: without the KEK, the encrypted signing keys in the database cannot be recovered.
Upgrades
- Back up the database (and confirm you have the KEK/signing key).
- Pull the new version (
docker compose pull/ bump the Helmimage.tag/ drop in the new Windows bundle). - Restart. Migrations apply automatically on startup — one step, no manual SQL.
curl /health(checkversion) and confirm/admin/self-licensestill reports your license.
Upgrades are backward-compatible within a minor line; with 2+ nodes you can roll one at a time for zero downtime. Downgrades after a migration has run are not supported — which is why step 1 matters.
Air-gapped operation
Nothing here needs the internet at runtime. Your license and your customers’ licenses verify offline against embedded public keys; credits for disconnected machines use signed air-gapped credit blocks (an Enterprise feature): issue a signed block, redeem it offline, reconcile on reconnect. Two things for a clean air-gapped run:
- Disable GeoIP: set
KEYRIGHT_GEOIP_URL=(empty) so the dashboard’s machine-location lookup attempts nothing outbound. - Move the images in: on a connected host,
docker save ghcr.io/delta1-labs/keyright:2.2.2(andpostgres:16if you use the bundled DB) to a tarball, thendocker loadinside the air-gapped network.
Troubleshooting & FAQ
The container exits on startup complaining the KEK must be set. Set KEYRIGHT_KEK to a base64 32-byte value (openssl rand -base64 32). Note the app connects to the database and runs migrations first, so a DB-connection error surfaces before this one.
Startup error connecting to the database / “relation does not exist”. Check KEYRIGHT_DB (host, port, database, credentials, SSL Mode). The role must be able to create tables (Configuration → Database); confirm the database is empty on first boot and the role owns it (or has CREATE ON SCHEMA public on PG15+).
The pod is in CrashLoopBackOff / the container keeps restarting, logs show Npgsql Name or service not known / connection refused / a timeout. The app runs its migrations at startup, so it exits (and restarts) if it can’t reach the database when it boots — you won’t get a running-but-not-ready pod. On a fresh bundled-Postgres evaluation install a few restarts are expected while Postgres becomes Ready, and clear on their own. Otherwise it’s almost always a connectivity problem, not a bug: (1) the DB host doesn’t resolve or isn’t reachable from the app node (check DNS/egress, and any NSG/firewall or “allow Azure services” rule on a managed DB); (2) TLS mismatch — most managed Postgres needs SSL Mode=Require (add ;Trust Server Certificate=true if you don’t validate the CA); (3) wrong credentials/database name. Pre-flight before deploying: from a pod on the same network, confirm the app can resolve and reach the DB host and port. Once the DB is reachable the pod goes Ready on its own.
Every request returns 402. Your instance is gated — no valid license/trial key (Licensing). Check /admin/self-license; install a key via KEYRIGHT_SELF_LICENSE and restart.
/admin/* returns 401. Send X-Admin-Token: $KEYRIGHT_ADMIN_TOKEN. The client path /v1/* never needs it.
Does the image include a database? No — it’s app-only. Run PostgreSQL alongside it (see the architecture overview and Configuration → Database).
My customers’ apps can’t activate. Confirm the SDK’s service URL points at your instance (Wire your product) and that the product slug + embedded public key match the product you issue keys for.
My license expired — what happens? A paid license gets a 15-day read-only grace (runtime + reads keep working); a trial stops immediately. After grace, the instance stops entirely until you install a renewed key (Licensing).
Do I need a load balancer? Only for 2+ app nodes (Scaling). A single node doesn’t (you may still want a TLS-terminating proxy in front).
Ready to run Keyright in your own environment? Talk to sales for a trial or a Standard/Enterprise license, or see pricing.