Deployment¶
Pylon ships deploy artifacts for three targets. Choose the tab that matches your environment.
Systemd is the primary deployment target. The artifacts in deploy/systemd/ are
production-grade and include a hardened unit file, kernel tuning, and an annotated
environment-variable template.
Files¶
| File | Purpose |
|---|---|
deploy/systemd/pylon.service |
systemd unit — runs pylon as the pylon system user, sets LimitNOFILE=2000000, handles graceful shutdown via SIGTERM. |
deploy/systemd/99-pylon.sysctl.conf |
Kernel tuning drop-in — TCP buffer sizes, somaxconn, fs.file-max, fs.nr_open, and tcp_migrate_req for millions of idle WebSocket connections. |
deploy/systemd/pylon.env.example |
Template environment file — all variables documented inline; copy to /etc/pylon/pylon.env and edit. |
Install steps¶
The commands below give the deploy/systemd/ paths for a source checkout;
from a release tarball, the same files ship under systemd/ instead.
1. Apply kernel tuning (once per host, as root):
The drop-in sets fs.nr_open = 20000500, which must be ≥ LimitNOFILE in
the service unit (2000000). It also shrinks per-socket TCP buffer sizes to
save ~15 KiB RAM per idle WebSocket connection.
2. Build the binary:
3. Create the service account and install files:
# Create the system user.
useradd --system --no-create-home --shell /sbin/nologin pylon
# Install the binary.
install -m 0755 target/release/pylon /usr/local/bin/pylon
# Create the config directory (owned root, readable by pylon group).
install -d -m 0750 -o root -g pylon /etc/pylon
# Install and edit the environment file.
install -m 0640 -o root -g pylon \
deploy/systemd/pylon.env.example /etc/pylon/pylon.env
# Edit /etc/pylon/pylon.env — set adapter, Redis URL, etc.
# Install the apps config.
install -m 0640 -o root -g pylon \
apps.example.json /etc/pylon/apps.json
# IMPORTANT: change the "secret" field in apps.json.
# Install the systemd unit.
cp deploy/systemd/pylon.service /etc/systemd/system/
systemctl daemon-reload
4. Enable and start:
Day-2 operations¶
# Status
systemctl status pylon
# Graceful restart (SIGTERM → drain → restart)
systemctl restart pylon
Type=exec makes systemctl restart return as soon as the binary executes, so check systemctl is-active pylon after a restart.
# Tail logs
journalctl -u pylon -f
# A start that keeps failing (bad config value, unreadable apps.json) stops
# retrying after five attempts within a minute and the unit reports failed.
systemctl is-active pylon # failed
journalctl -u pylon -n 20 # the last attempt's error, e.g. invalid PYLON_SHUTDOWN_GRACE_MS="10s"
systemctl reset-failed pylon && systemctl start pylon # after fixing the config
# Health check
curl -s http://localhost:7000/health
curl -s http://localhost:7000/ready
Redis adapter (multi-node)¶
To run multiple nodes behind a load balancer, edit /etc/pylon/pylon.env on
every host:
If Redis runs on the same host, uncomment the Requires=redis.service lines in
pylon.service. See Clustering & Scaling for the full
multi-node setup guide.
Graceful shutdown¶
The unit sets KillSignal=SIGTERM and TimeoutStopSec=20. On systemctl stop
or systemctl restart, pylon:
- Flips
/readyto 503 immediately (LB stops sending new connections). - Waits
PYLON_SHUTDOWN_PREDRAIN_MS(default 2 s). - Sends a
pusher:error4200 frame followed by a WebSocket Close (4200) to all connections — Pusher's reconnect-immediately code, so clients reconnect to a surviving node without backoff. - Flushes up to
PYLON_SHUTDOWN_GRACE_MS(default 10 s). - Exits.
Worst-case drain is ~12 s. The 20 s TimeoutStopSec provides slack before
systemd force-kills the process.
Host file-descriptor limits
LimitNOFILE=2000000 covers 1 M WebSocket connections plus headroom for
epoll descriptors, timer fds, and sockets. Host fs.nr_open must be at
least this value — the 99-pylon.sysctl.conf drop-in ensures it.
Connection-count and memory-budget tuning are covered in Production Tuning.
Host prerequisites¶
Apply these before starting any container: the daemon-level nofile default cannot be
satisfied until fs.nr_open is raised.
Apply the kernel tuning drop-in on the Docker host before starting containers:
Ensure the Docker daemon allows high nofile limits by adding to
/etc/docker/daemon.json:
Restart the Docker daemon after editing this file.
Published image¶
A multi-arch image (linux/amd64 + linux/arm64) is published on each release:
ghcr.io/i-rocky/pylon:latest
ghcr.io/i-rocky/pylon:X.Y.Z # pinned release
ghcr.io/i-rocky/pylon:X.Y # floating minor
Single-node quick start¶
cp apps.example.json apps.json # from the release tarball or the repo root
# edit apps.json: set id, key and secret
chown 65534:65534 apps.json && chmod 0600 apps.json
The image runs as UID 65534 (nobody) with no shell, so a file it cannot read makes the
container exit with Permission denied; owning it to 65534 with mode 0600 keeps the secret
private and readable.
docker run -d --name pylon -p 7000:7000 \
-v "$PWD/apps.json:/etc/pylon/apps.json:ro" \
-e PYLON_APPS_PATH=/etc/pylon/apps.json \
--ulimit nofile=1048576:1048576 \
--stop-timeout 20 \
ghcr.io/i-rocky/pylon:latest
Volume-mount your apps.json at /etc/pylon/apps.json and pass the path
via PYLON_APPS_PATH. The --ulimit flag raises the file-descriptor limit
for the container. The default 10 s stop timeout is below the 12 s drain
worst case (2 s pre-drain plus 10 s grace); --stop-timeout 20 makes
docker stop use 20 s.
Two-node Compose cluster¶
deploy/docker/docker-compose.yml starts Redis 7, pylon-1 (host port 7000),
and pylon-2 (host port 7001), all sharing the same apps.json and using the
redis adapter.
# Copy and edit the apps config — change the secret!
cp apps.example.json deploy/docker/apps.json
chown 65534:65534 deploy/docker/apps.json && chmod 0600 deploy/docker/apps.json
# Build and start.
cd deploy/docker
docker compose up -d --build
# Verify both nodes.
docker compose ps
curl -s http://localhost:7000/health
curl -s http://localhost:7001/health
Rolling update¶
docker compose up -d --no-deps --build pylon-1
# Wait for pylon-1 to become healthy, then:
docker compose up -d --no-deps pylon-2
Each node's stop_grace_period: 20s in the compose file ensures the full
drain cycle completes before Docker kills the container.
No in-container health check
The pylon image is FROM scratch: a static binary, a CA bundle and a
one-line /etc/passwd, with no shell, curl or wget to probe with. Probe
GET /health and GET /ready over HTTP from outside the container —
which is what Kubernetes httpGet probes, load balancers and the CI smoke
test already do.
The Helm chart is at deploy/helm/pylon. It packages a Deployment,
Service, an apps Secret, liveness/readiness probes, a
HorizontalPodAutoscaler (opt-in), a PodDisruptionBudget, and security
contexts.
Node-level prerequisites¶
Kubernetes cannot apply most net.* sysctls per-pod. Before deploying pylon,
apply deploy/systemd/99-pylon.sysctl.conf to every node in the cluster
and ensure the container runtime allows nofile ≥ 2 000 000. See the comment
block at the bottom of deploy/helm/pylon/values.yaml for runtime-specific
instructions (containerd, docker-shim).
Install¶
The chart refuses to render until every app has a real secret, so the values file is written first with one generated in place:
# Single-node (local adapter, default):
cat > my-values.yaml <<EOF
apps:
- name: my-app
id: app
key: app-key
secret: $(openssl rand -hex 32)
capacity: 1000000
client_messages_enabled: false
enabled: true
webhooks: []
EOF
helm install pylon ./deploy/helm/pylon -f my-values.yaml
# Multi-node cluster (redis adapter):
helm install pylon ./deploy/helm/pylon -f my-values.yaml \
--set config.adapter=redis \
--set config.redisUrl=redis://my-redis:6379 \
--set replicaCount=3
Key values¶
| Value | Default | Purpose |
|---|---|---|
replicaCount |
1 |
Number of pylon pods. |
config.adapter |
local |
local or redis. Must be redis for replicaCount > 1. |
config.redisUrl |
"" |
Redis connection URL (required when adapter=redis). |
config.redisPrefix |
pylon |
Redis key prefix. |
config.workers |
0 |
Worker threads. 0 = one per CPU. |
config.memoryBudgetBytes |
0 |
Memory cap in bytes. 0 = auto. |
config.shutdownPredrainsMs |
2000 |
LB drain window after SIGTERM, before closing connections. |
config.shutdownGraceMs |
10000 |
Max time to flush in-flight connections. |
config.terminationGracePeriodSeconds |
0 |
Pod grace period. 0 = derive from the two settings above plus a safety margin, floored at 30s. |
autoscaling.enabled |
false |
Enable the HPA. |
autoscaling.minReplicas |
2 |
Minimum replicas when HPA is active. |
autoscaling.maxReplicas |
10 |
Maximum replicas when HPA is active. |
resources.requests.memory |
512Mi |
Pod memory request. |
resources.limits.memory |
8Gi |
Pod memory limit. |
existingSecret |
"" |
Name of a Secret you manage yourself (keys apps.json, redisUrl). When set, the chart creates no Secret. |
podDisruptionBudget.enabled |
true |
Render a PodDisruptionBudget; rendered only when replicaCount is above 1 or autoscaling is enabled. |
podDisruptionBudget.minAvailable |
1 |
Minimum pods that must stay up during a voluntary disruption. |
Byte and millisecond values are plain integers (memoryBudgetBytes: 2147483648, shutdownGraceMs: 10000); pylon has no unit suffixes, and the chart passes whatever is written through unchanged, so 2Gi or 10s fails at pod start with invalid PYLON_MEMORY_BUDGET_BYTES="2Gi" rather than being altered.
The chart refuses to render more than one replica, or autoscaling, on the local adapter, and renders the PodDisruptionBudget only when more than one pod can exist.
Autoscaling¶
Autoscaling needs the redis adapter, and helm upgrade without --reuse-values
re-reads the chart defaults, so pass the adapter again:
helm upgrade pylon ./deploy/helm/pylon \
--set config.adapter=redis \
--set config.redisUrl=redis://my-redis:6379 \
--set autoscaling.enabled=true \
--set autoscaling.minReplicas=2 \
--set autoscaling.maxReplicas=10
Graceful rollout¶
The Deployment template derives terminationGracePeriodSeconds from
config.shutdownPredrainsMs + config.shutdownGraceMs plus a safety
margin, floored at 30s. An explicit config.terminationGracePeriodSeconds
override is honoured as given, provided it's large enough to fit the
drain — the chart fails the render if it's too small or not a usable
non-negative integer. It also uses a rolling update strategy with
maxUnavailable: 0 to keep the full replica count serving traffic during
a rollout. The readiness probe (GET /ready) removes a pod from Service
endpoints as soon as it enters the drain phase.
TLS (Ingress)¶
Terminate TLS at the Ingress controller using cert-manager. The pylon pods and Service stay on plain HTTP. See TLS / SSL for the full Ingress manifest with the required WebSocket and timeout annotations.
Apps config and credentials¶
App secrets and the Redis URL are rendered into a Kubernetes Secret, never a
ConfigMap, and reach the pod as a mounted file (/etc/pylon/apps.json) and a
secretKeyRef (PYLON_REDIS_URL). The chart refuses to render while any
app's secret is empty or the placeholder CHANGE_ME — it is the HMAC
key behind every REST signature, channel-auth token and pusher:signin for
that app, so set a real one (openssl rand -hex 32) or point
existingSecret at a Secret you manage.
If you manage secrets outside Helm — an external secret manager,
sealed-secrets, or a CI-created Secret — set existingSecret to its name and
the chart creates none:
The Secret must carry two keys: apps.json (the app registry) and redisUrl.
Disruption budget¶
The chart renders a PodDisruptionBudget with minAvailable: 1 by default, so
a node drain or an autoscaler scale-down can never take every pylon pod at
once. The Deployment's maxUnavailable: 0 covers only rolling updates; a
voluntary eviction is a different path and needs its own budget.
For the full Helm values reference see deploy/helm/pylon/values.yaml and
deploy/README.md.
Health probes¶
All deployment targets expose the same two HTTP probes on the pylon port:
| Probe | Path | Healthy | Draining / starting |
|---|---|---|---|
| Liveness | GET /health |
200 "ok" | 200 "ok" (always) |
| Readiness | GET /ready |
200 "ready" | 503 "draining" or "starting" |
Configure your load balancer or Kubernetes readinessProbe to use /ready.
This ensures that a node in the pre-drain window stops receiving new connections
before its existing connections are closed.
Connection-count and memory tuning
File-descriptor limits (LimitNOFILE, --ulimit nofile, container runtime
settings) and kernel TCP buffer tuning are covered in Production Tuning. Set
PYLON_MEMORY_BUDGET_BYTES to cap memory consumption — see
Configuration for the full variable reference.
Readiness and shared dependencies¶
GET /ready answers on this node's own state alone: the per-core worker fleet
is up and the node is not draining. It deliberately does not probe Redis or
the app store.
Those are shared by every replica. If readiness included them, a single Redis or database outage would fail every pod's readiness at the same moment, the endpoint controller would remove every pod from the Service, and a degraded cluster would become an unreachable one — with no node left to serve the connections that still work. A node that has lost a shared dependency is degraded, not dead: existing WebSocket connections keep flowing, node-local delivery keeps working, and recovery needs no restart.
Alert on pylon_redis_connected == 0 and pylon_app_store_up == 0 instead —
see Observability for the rules.
Protecting /metrics¶
GET /metrics is open by default (back-compat). Because metrics expose app IDs,
connection counts, and infrastructure detail, bind pylon to a private interface
in any shared network. When network-level restriction is not enough — or as
defense in depth — arm the optional bearer gate:
With the token set:
- Every scrape must carry
Authorization: Bearer <token>. - A missing or wrong token returns 404 — deliberately not 401, so an unauthenticated prober cannot distinguish "metrics exist but are protected" from "no such route".
- The token is compared in constant time (no timing oracle on its value).
/healthand/ready(and/healthz//readyz) stay open — load balancers and kubelet probes never need the token.- An empty
PYLON_METRICS_TOKENis treated as unset (metrics stay open).
Prometheus scrape config with the token:
scrape_configs:
- job_name: pylon
static_configs:
- targets: ["pylon-host-1:7000"]
authorization:
type: Bearer
credentials: ${PYLON_METRICS_TOKEN}
In Kubernetes, put the token in a Secret and mount it into both pylon and the Prometheus configuration rather than embedding it in a ConfigMap. See Observability for the full metrics reference.