Observability¶
Pylon exposes Prometheus metrics, a liveness probe, and a readiness probe on the
same HTTP port as the REST API (PYLON_PORT, default 7000).
GET /metrics — Prometheus Exposition¶
Returns a Prometheus text-format (v0.0.4) snapshot.
By default the endpoint is unauthenticated (back-compat). Set
PYLON_METRICS_TOKEN to require Authorization: Bearer <token> on every
scrape; a request without the correct token then gets 404 (not 401, so an
unauthenticated prober cannot learn that the endpoint exists). /health and
/ready are never gated — load-balancer probes keep working without the
token. See Deployment: Protecting /metrics.
curl http://localhost:7000/metrics
# With the token gate armed:
curl -H "Authorization: Bearer $PYLON_METRICS_TOKEN" http://localhost:7000/metrics
Metrics Reference¶
All series use cardinality-safe labels only: worker (integer index), app
(app ID string), status (ok / failed). There are no per-channel
labels.
Process¶
| Series | Type | Labels | Description |
|---|---|---|---|
pylon_up |
gauge | — | Always 1; confirms the process is alive and the scrape succeeded |
pylon_rest_rate_limited_total |
counter | scope |
REST requests rejected with 429; scope="node", "app_events" or "app_reads". Always emitted, reading 0 while the matching limit is off |
pylon_app_store_up |
gauge | — | 1 = the app store answered its last probe; 0 = it failed or timed out. Always present. |
Per-App¶
| Series | Type | Labels | Description |
|---|---|---|---|
pylon_connections |
gauge | app |
Live WebSocket connections for the app |
pylon_channels_occupied |
gauge | app |
Channels with at least one subscriber |
pylon_subscriptions |
gauge | app |
Total channel subscriptions across all connections |
Which apps get a series, and whether it stays present at 0 versus
disappearing, depends on the app manager backend:
- Static-file app manager (
PYLON_APP_MANAGER=static/ the default): the app set is fixed at startup, so every configured app always emits all three series, reading0while idle.pylon_connections{app="x"} == 0alert rules work as written — the series never disappears. - Dynamic backends (SQL, Mongo): the app set is unbounded, so a series
only appears once an app has had at least one tracked connection, and drops
again once its connection count returns to zero — retaining every id ever
seen would be an unbounded-growth path. Write alerts against these apps
with
absent()(e.g.absent(pylon_connections{app="x"}) or pylon_connections{app="x"} == 0), not a bare== 0comparison, or the rule will silently stop evaluating the moment the app goes idle.
Clustered deployments: pylon_channels_occupied / pylon_subscriptions are cluster-wide, per node
pylon_connections is always this node's own local count. But
pylon_channels_occupied and pylon_subscriptions are sourced from the
configured adapter's channel view, and under PYLON_ADAPTER=redis
RedisAdapter::channels returns a cluster-wide view, not a per-node
one. A node reports these two series only for an app it has at least one
local connection for — a node with zero local connections for an app
reports 0 (from the static-file zero-seeding above), while a node with
even one local connection reports the FULL cluster-wide occupied-channel
and subscription count for that app, not just its own share. In a
multi-node deployment where an app's connections are spread across
nodes, avg by(app) or min by(app) aggregations across
pylon_channels_occupied{app="x"} / pylon_subscriptions{app="x"} can
therefore read low (the average is diluted by idle nodes' 0s; the min
is 0 whenever any node is idle for that app) — max by(app) is the
aggregation that reflects the true cluster-wide value in this mode.
Per-Worker Transport¶
| Series | Type | Labels | Description |
|---|---|---|---|
pylon_accepted_connections_total |
counter | worker |
Cumulative connections accepted by each worker since start |
pylon_broadcast_dropped_total |
counter | worker |
Broadcasts dropped because the worker hand-off channel was full |
pylon_codel_dropped_total |
counter | worker |
Frames discarded by the CoDel staleness check (stale frames removed from the queue before sending) |
pylon_drophead_dropped_total |
counter | worker |
Frames evicted by the per-connection drop-head queue: when a slow consumer's outbound queue is at its byte cap, the oldest queued frames are dropped to make room for newer ones (freshest-wins) |
pylon_mailbox_dropped_total |
counter | worker |
Frames dropped because a connection's inbound mailbox (the bounded direct-send channel for presence rosters, member events, user-targeted sends, watchlist notifications, cluster deliveries) was full when a producer tried to enqueue |
pylon_frame_limited_total |
counter | worker |
Connections closed with code 4100 for exceeding the per-connection inbound frame rate |
pylon_accept_limited_total |
counter | worker |
Sockets closed immediately after accept for exceeding the per-worker accept rate |
pylon_handshake_timeout_total |
counter | worker |
Pre-session connections reaped for exceeding PYLON_HANDSHAKE_TIMEOUT_MS, the absolute deadline from TCP accept within which a connection must finish its HTTP head, TLS, WebSocket upgrade and session establish. The reap sends no protocol close, so this counter is the only server-side signal that it happened |
pylon_inflight_bytes |
gauge | worker |
Bytes currently queued in each worker's outbound buffer |
pylon_inflight_bytes_sum |
gauge | — | Sum of pylon_inflight_bytes across all workers |
pylon_worker_budget_bytes |
gauge | — | Per-worker memory budget in bytes |
pylon_budget_factor |
gauge | — | PSI memory-pressure budget factor. A background loop polls kernel memory pressure (full avg10) once per second; above PYLON_PSI_THRESHOLD (default 15%) each worker's effective budget shrinks toward a 0.8 floor and recovers toward 1.0 when pressure clears. Steady-state range is 0.8–1.0 — a sustained value below 0.9 indicates real memory pressure, not queue backlog |
pylon_saturation_flag |
gauge | — | 1 while the node is shedding, 0 otherwise; omitted when the saturation monitor is not running. This is an actionable alert series — see the note below |
pylon_saturation_flag now has teeth — alert on it
The flag is raised by any worker whose queued outbound bytes reach 100 % of
its share of the memory budget (released below 80 %), or by a publisher that
found a worker's broadcast hand-off channel full. While it is 1 the node
is actively rejecting work: REST publishes get 503 + Retry-After: 1,
new subscriptions get a non-fatal 4004 LimitReached, inbound client-*
events are dropped silently, and new connections are closed with 4100.
Earlier builds cleared the underlying flag unconditionally every worker
loop, so it effectively never read 1 and none of those responses fired. It
now works, so a dashboard or alert that was never triggered before may start
firing after an upgrade. Alert on a sustained 1, and read it alongside
pylon_inflight_bytes (per worker) versus pylon_worker_budget_bytes.
Full diagnosis steps are in
Troubleshooting — Overload.
Webhook Pipeline¶
| Series | Type | Labels | Description |
|---|---|---|---|
pylon_webhook_enqueued_total |
counter | — | Webhook events successfully placed in the delivery mailbox |
pylon_webhook_dropped_total |
counter | — | Webhook events dropped because the mailbox was full or closed |
pylon_webhook_delivered_total |
counter | status |
Completed delivery attempts; status="ok" or status="failed" |
pylon_webhook_queue_depth |
gauge | — | Current number of events waiting in the webhook mailbox |
Cluster / Redis (only present on the Redis clustering path)¶
| Series | Type | Labels | Description |
|---|---|---|---|
pylon_cluster_cmd_dropped_total |
counter | — | ClusterCmd messages dropped on a full bridge channel |
pylon_cluster_publish_failed_total |
counter | — | Cross-node broadcast publishes that failed: the bridge channel was full or closed, or Redis rejected the PUBLISH. A REST trigger publishes cross-node before it delivers locally, so either failure answers 503 and no subscriber received the event. A WebSocket client event is handed to the cluster bridge first, so where the failure lands decides who received it: a hand-off refused by a full or closed bridge channel drops the event before any delivery — nobody receives it, not even this node's subscribers, and the sender is sent no error frame — while a Redis PUBLISH that fails inside the bridge after the hand-off succeeded leaves the event delivered to this node's subscribers and to no other node. Both cases are counted here and warned |
pylon_redis_connected |
gauge | — | 1 = Redis connection healthy; 0 = error/disconnected |
Prometheus Scrape Config¶
Add pylon as a scrape target in prometheus.yml:
scrape_configs:
- job_name: pylon
static_configs:
- targets:
- "pylon-host-1:7000"
- "pylon-host-2:7000"
# No auth needed unless PYLON_METRICS_TOKEN is set; restrict at the
# network layer regardless.
# authorization:
# type: Bearer
# credentials: ${PYLON_METRICS_TOKEN}
For Kubernetes, use a ServiceMonitor (Prometheus Operator) pointing at port
7000 with path /metrics.
Grafana¶
Import the series above into Grafana dashboards. Useful panel ideas:
- Connection density:
pylon_connections{app="…"}per app, plussum(pylon_inflight_bytes_sum)for buffer pressure. - Drop rates: rate of
pylon_broadcast_dropped_total,pylon_codel_dropped_total,pylon_drophead_dropped_total, andpylon_mailbox_dropped_total— non-zero values indicate backpressure (drop-head evictions mean slow consumers are losing their oldest queued frames; mailbox drops mean direct sends to a connection overran its inbound mailbox). - Webhook health:
pylon_webhook_dropped_totalrate andpylon_webhook_queue_depth— rising queue depth signals a slow upstream. - Redis health:
pylon_redis_connectedas a status panel; alert on< 1. - Memory pressure:
pylon_budget_factor— alert when sustained below 0.9 (the factor only ever ranges 0.8–1.0; it tracks kernel PSI memory pressure, not queue depth). - Overload / shedding:
pylon_saturation_flagas a status panel, alerting on a sustained1— the node is rejecting publishes with503and dropping client events while it reads high. Pair it withpylon_inflight_bytes / on() group_left pylon_worker_budget_bytesto see how close each worker is to its 100 % raise / 80 % release band.
Alerting¶
groups:
- name: pylon
rules:
- alert: PylonRedisDown
expr: pylon_redis_connected == 0
for: 2m
labels: { severity: critical }
annotations:
summary: "pylon {{ $labels.instance }} has lost Redis"
description: "Cross-node delivery is degraded on this node. /ready stays 200 by design; see Deployment."
- alert: PylonAppStoreDown
expr: pylon_app_store_up == 0
for: 2m
labels: { severity: critical }
annotations:
summary: "pylon {{ $labels.instance }} cannot reach its app store"
description: "REST auth answers 503 and new connections cannot resolve their app. /ready stays 200 by design; see Deployment."
The SQL probe shares the app manager's own connection pool with every other
lookup, so a burst of cache-missing queries can make a single probe time out
without the store actually being down — the reason PylonAppStoreDown carries
for: 2m rather than firing on the first missed probe.
GET /health — Liveness Probe¶
Always returns 200 ok as long as the process can handle HTTP requests. Use
this as a liveness probe: if it fails, restart the container.
Kubernetes liveness probe:
Also available at /healthz.
GET /ready — Readiness Probe¶
| Status | Body | Meaning |
|---|---|---|
200 OK |
ready |
Workers up and not draining — safe to route traffic here |
503 Service Unavailable |
draining |
Shutdown in progress; stop routing new connections |
503 Service Unavailable |
starting |
Workers not yet initialised |
Use this as a readiness probe: the load balancer or k8s controller stops sending new connections to a node when it returns non-200, which is exactly what you want during a graceful restart.
Kubernetes readiness probe:
readinessProbe:
httpGet:
path: /ready
port: 7000
initialDelaySeconds: 3
periodSeconds: 5
failureThreshold: 2
Also available at /readyz.
Load-balancer health checks
Point your load balancer's health check at /ready (not /health). This
ensures that draining nodes are removed from the rotation before their
connections are closed with Pusher code 4200.
Logs¶
Verbosity comes from RUST_LOG (default info); the format comes from
PYLON_LOG_FORMAT.
text (the default) is the human-readable formatter; it colours its output
only when stdout is a terminal, so journald, Docker and Kubernetes logs carry
no ANSI escape sequences. json emits one JSON object per line, which is what
a log pipeline wants:
{"timestamp":"2026-09-14T09:12:04.118273Z","level":"WARN","fields":{"message":"cluster publish failed; broadcast dropped","app":"app1","channel":"presence-room","error":"cluster publish failed: cluster bridge channel full or closed"},"target":"pylon::ws::handler"}
The message is at fields.message; every structured field pylon attaches —
app, channel, worker, error — is a sibling key under fields rather
than being flattened into prose. An unrecognised PYLON_LOG_FORMAT is a
startup error, not a silent fallback.