Skip to content

Observability

Pylon exposes Prometheus metrics, a liveness probe, and a readiness probe on the same HTTP port as the REST API (PYLON_PORT, default 7000).


GET /metrics — Prometheus Exposition

Returns a Prometheus text-format (v0.0.4) snapshot.

By default the endpoint is unauthenticated (back-compat). Set PYLON_METRICS_TOKEN to require Authorization: Bearer <token> on every scrape; a request without the correct token then gets 404 (not 401, so an unauthenticated prober cannot learn that the endpoint exists). /health and /ready are never gated — load-balancer probes keep working without the token. See Deployment: Protecting /metrics.

curl http://localhost:7000/metrics

# With the token gate armed:
curl -H "Authorization: Bearer $PYLON_METRICS_TOKEN" http://localhost:7000/metrics

Metrics Reference

All series use cardinality-safe labels only: worker (integer index), app (app ID string), status (ok / failed). There are no per-channel labels.

Process

Series Type Labels Description
pylon_up gauge — Always 1; confirms the process is alive and the scrape succeeded
pylon_rest_rate_limited_total counter scope REST requests rejected with 429; scope="node", "app_events" or "app_reads". Always emitted, reading 0 while the matching limit is off
pylon_app_store_up gauge — 1 = the app store answered its last probe; 0 = it failed or timed out. Always present.

Per-App

Series Type Labels Description
pylon_connections gauge app Live WebSocket connections for the app
pylon_channels_occupied gauge app Channels with at least one subscriber
pylon_subscriptions gauge app Total channel subscriptions across all connections

Which apps get a series, and whether it stays present at 0 versus disappearing, depends on the app manager backend:

  • Static-file app manager (PYLON_APP_MANAGER=static / the default): the app set is fixed at startup, so every configured app always emits all three series, reading 0 while idle. pylon_connections{app="x"} == 0 alert rules work as written — the series never disappears.
  • Dynamic backends (SQL, Mongo): the app set is unbounded, so a series only appears once an app has had at least one tracked connection, and drops again once its connection count returns to zero — retaining every id ever seen would be an unbounded-growth path. Write alerts against these apps with absent() (e.g. absent(pylon_connections{app="x"}) or pylon_connections{app="x"} == 0), not a bare == 0 comparison, or the rule will silently stop evaluating the moment the app goes idle.

Clustered deployments: pylon_channels_occupied / pylon_subscriptions are cluster-wide, per node

pylon_connections is always this node's own local count. But pylon_channels_occupied and pylon_subscriptions are sourced from the configured adapter's channel view, and under PYLON_ADAPTER=redis RedisAdapter::channels returns a cluster-wide view, not a per-node one. A node reports these two series only for an app it has at least one local connection for — a node with zero local connections for an app reports 0 (from the static-file zero-seeding above), while a node with even one local connection reports the FULL cluster-wide occupied-channel and subscription count for that app, not just its own share. In a multi-node deployment where an app's connections are spread across nodes, avg by(app) or min by(app) aggregations across pylon_channels_occupied{app="x"} / pylon_subscriptions{app="x"} can therefore read low (the average is diluted by idle nodes' 0s; the min is 0 whenever any node is idle for that app) — max by(app) is the aggregation that reflects the true cluster-wide value in this mode.

Per-Worker Transport

Series Type Labels Description
pylon_accepted_connections_total counter worker Cumulative connections accepted by each worker since start
pylon_broadcast_dropped_total counter worker Broadcasts dropped because the worker hand-off channel was full
pylon_codel_dropped_total counter worker Frames discarded by the CoDel staleness check (stale frames removed from the queue before sending)
pylon_drophead_dropped_total counter worker Frames evicted by the per-connection drop-head queue: when a slow consumer's outbound queue is at its byte cap, the oldest queued frames are dropped to make room for newer ones (freshest-wins)
pylon_mailbox_dropped_total counter worker Frames dropped because a connection's inbound mailbox (the bounded direct-send channel for presence rosters, member events, user-targeted sends, watchlist notifications, cluster deliveries) was full when a producer tried to enqueue
pylon_frame_limited_total counter worker Connections closed with code 4100 for exceeding the per-connection inbound frame rate
pylon_accept_limited_total counter worker Sockets closed immediately after accept for exceeding the per-worker accept rate
pylon_handshake_timeout_total counter worker Pre-session connections reaped for exceeding PYLON_HANDSHAKE_TIMEOUT_MS, the absolute deadline from TCP accept within which a connection must finish its HTTP head, TLS, WebSocket upgrade and session establish. The reap sends no protocol close, so this counter is the only server-side signal that it happened
pylon_inflight_bytes gauge worker Bytes currently queued in each worker's outbound buffer
pylon_inflight_bytes_sum gauge — Sum of pylon_inflight_bytes across all workers
pylon_worker_budget_bytes gauge — Per-worker memory budget in bytes
pylon_budget_factor gauge — PSI memory-pressure budget factor. A background loop polls kernel memory pressure (full avg10) once per second; above PYLON_PSI_THRESHOLD (default 15%) each worker's effective budget shrinks toward a 0.8 floor and recovers toward 1.0 when pressure clears. Steady-state range is 0.8–1.0 — a sustained value below 0.9 indicates real memory pressure, not queue backlog
pylon_saturation_flag gauge — 1 while the node is shedding, 0 otherwise; omitted when the saturation monitor is not running. This is an actionable alert series — see the note below

pylon_saturation_flag now has teeth — alert on it

The flag is raised by any worker whose queued outbound bytes reach 100 % of its share of the memory budget (released below 80 %), or by a publisher that found a worker's broadcast hand-off channel full. While it is 1 the node is actively rejecting work: REST publishes get 503 + Retry-After: 1, new subscriptions get a non-fatal 4004 LimitReached, inbound client-* events are dropped silently, and new connections are closed with 4100.

Earlier builds cleared the underlying flag unconditionally every worker loop, so it effectively never read 1 and none of those responses fired. It now works, so a dashboard or alert that was never triggered before may start firing after an upgrade. Alert on a sustained 1, and read it alongside pylon_inflight_bytes (per worker) versus pylon_worker_budget_bytes. Full diagnosis steps are in Troubleshooting — Overload.

Webhook Pipeline

Series Type Labels Description
pylon_webhook_enqueued_total counter — Webhook events successfully placed in the delivery mailbox
pylon_webhook_dropped_total counter — Webhook events dropped because the mailbox was full or closed
pylon_webhook_delivered_total counter status Completed delivery attempts; status="ok" or status="failed"
pylon_webhook_queue_depth gauge — Current number of events waiting in the webhook mailbox

Cluster / Redis (only present on the Redis clustering path)

Series Type Labels Description
pylon_cluster_cmd_dropped_total counter — ClusterCmd messages dropped on a full bridge channel
pylon_cluster_publish_failed_total counter — Cross-node broadcast publishes that failed: the bridge channel was full or closed, or Redis rejected the PUBLISH. A REST trigger publishes cross-node before it delivers locally, so either failure answers 503 and no subscriber received the event. A WebSocket client event is handed to the cluster bridge first, so where the failure lands decides who received it: a hand-off refused by a full or closed bridge channel drops the event before any delivery — nobody receives it, not even this node's subscribers, and the sender is sent no error frame — while a Redis PUBLISH that fails inside the bridge after the hand-off succeeded leaves the event delivered to this node's subscribers and to no other node. Both cases are counted here and warned
pylon_redis_connected gauge — 1 = Redis connection healthy; 0 = error/disconnected

Prometheus Scrape Config

Add pylon as a scrape target in prometheus.yml:

scrape_configs:
  - job_name: pylon
    static_configs:
      - targets:
          - "pylon-host-1:7000"
          - "pylon-host-2:7000"
    # No auth needed unless PYLON_METRICS_TOKEN is set; restrict at the
    # network layer regardless.
    # authorization:
    #   type: Bearer
    #   credentials: ${PYLON_METRICS_TOKEN}

For Kubernetes, use a ServiceMonitor (Prometheus Operator) pointing at port 7000 with path /metrics.


Grafana

Import the series above into Grafana dashboards. Useful panel ideas:

  • Connection density: pylon_connections{app="…"} per app, plus sum(pylon_inflight_bytes_sum) for buffer pressure.
  • Drop rates: rate of pylon_broadcast_dropped_total, pylon_codel_dropped_total, pylon_drophead_dropped_total, and pylon_mailbox_dropped_total — non-zero values indicate backpressure (drop-head evictions mean slow consumers are losing their oldest queued frames; mailbox drops mean direct sends to a connection overran its inbound mailbox).
  • Webhook health: pylon_webhook_dropped_total rate and pylon_webhook_queue_depth — rising queue depth signals a slow upstream.
  • Redis health: pylon_redis_connected as a status panel; alert on < 1.
  • Memory pressure: pylon_budget_factor — alert when sustained below 0.9 (the factor only ever ranges 0.8–1.0; it tracks kernel PSI memory pressure, not queue depth).
  • Overload / shedding: pylon_saturation_flag as a status panel, alerting on a sustained 1 — the node is rejecting publishes with 503 and dropping client events while it reads high. Pair it with pylon_inflight_bytes / on() group_left pylon_worker_budget_bytes to see how close each worker is to its 100 % raise / 80 % release band.

Alerting

groups:
  - name: pylon
    rules:
      - alert: PylonRedisDown
        expr: pylon_redis_connected == 0
        for: 2m
        labels: { severity: critical }
        annotations:
          summary: "pylon {{ $labels.instance }} has lost Redis"
          description: "Cross-node delivery is degraded on this node. /ready stays 200 by design; see Deployment."
      - alert: PylonAppStoreDown
        expr: pylon_app_store_up == 0
        for: 2m
        labels: { severity: critical }
        annotations:
          summary: "pylon {{ $labels.instance }} cannot reach its app store"
          description: "REST auth answers 503 and new connections cannot resolve their app. /ready stays 200 by design; see Deployment."

The SQL probe shares the app manager's own connection pool with every other lookup, so a burst of cache-missing queries can make a single probe time out without the store actually being down — the reason PylonAppStoreDown carries for: 2m rather than firing on the first missed probe.


GET /health — Liveness Probe

200 OK   body: ok

Always returns 200 ok as long as the process can handle HTTP requests. Use this as a liveness probe: if it fails, restart the container.

curl -f http://localhost:7000/health

Kubernetes liveness probe:

livenessProbe:
  httpGet:
    path: /health
    port: 7000
  initialDelaySeconds: 5
  periodSeconds: 10

Also available at /healthz.


GET /ready — Readiness Probe

Status Body Meaning
200 OK ready Workers up and not draining — safe to route traffic here
503 Service Unavailable draining Shutdown in progress; stop routing new connections
503 Service Unavailable starting Workers not yet initialised

Use this as a readiness probe: the load balancer or k8s controller stops sending new connections to a node when it returns non-200, which is exactly what you want during a graceful restart.

curl -f http://localhost:7000/ready

Kubernetes readiness probe:

readinessProbe:
  httpGet:
    path: /ready
    port: 7000
  initialDelaySeconds: 3
  periodSeconds: 5
  failureThreshold: 2

Also available at /readyz.

Load-balancer health checks

Point your load balancer's health check at /ready (not /health). This ensures that draining nodes are removed from the rotation before their connections are closed with Pusher code 4200.

Logs

Verbosity comes from RUST_LOG (default info); the format comes from PYLON_LOG_FORMAT.

text (the default) is the human-readable formatter; it colours its output only when stdout is a terminal, so journald, Docker and Kubernetes logs carry no ANSI escape sequences. json emits one JSON object per line, which is what a log pipeline wants:

{"timestamp":"2026-09-14T09:12:04.118273Z","level":"WARN","fields":{"message":"cluster publish failed; broadcast dropped","app":"app1","channel":"presence-room","error":"cluster publish failed: cluster bridge channel full or closed"},"target":"pylon::ws::handler"}

The message is at fields.message; every structured field pylon attaches — app, channel, worker, error — is a sibling key under fields rather than being flattened into prose. An unrecognised PYLON_LOG_FORMAT is a startup error, not a silent fallback.