feat: per-upstream circuit breaker to stop saturated-upstream pile-ups - #266
Conversation
When an upstream saturates, jussi currently keeps forwarding every request: each one waits out its timeout, gets cancelled, and the SQL the upstream already started keeps running and holding its connection. New requests keep arriving, so a slow backend never drains. This is the amplification loop behind the 2026-09-18 hivemind storm (jussi logged ~1.3M deadline cancellations in 2 hours; the hivemind ALB recorded them as ELB-side 400s while wallet SSR requests hit 504). Add a per-upstream-host circuit breaker (sliding 30s window, 50% failure rate with a 20-sample minimum, 10s open period with 25% jitter, and a single-probe half-open state): - closed: requests flow; failures (timeouts, transport errors, 5xx) feed the window - open: requests fail fast without dialing the upstream, so arrivals stop and in-flight upstream work can drain - half-open: exactly one probe request is admitted; success closes the breaker and clears the window, failure reopens for another cycle Broadcast transactions keep their no-retry path; the breaker wraps both the fail-fast and retry paths in callHTTPUpstream. Observability: jussi_upstream_circuit_state gauge per upstream host and jussi_upstream_circuit_rejects_total counter; rejections return a structured JSON-RPC error (CodeUpstreamResponseErr, 'Upstream temporarily unavailable') so clients see a normal upstream error class. The breaker is deliberately not retried through (rejections are not RetriableUpstreamError), so the existing bounded retry logic cannot re-enter an open breaker.
…ns, env config Review follow-ups for the circuit breaker PR: 1. get_state workaround bypassed the breaker (github review n.1). callSteemd dialed the upstream directly — beta-hivemind in production, i.e. the exact backend this breaker protects — with a 15s sub-request timeout, invisible to the breaker. Route it through Allow()/Record() like every other upstream call; the emulated paths are idempotent reads, so failing fast while open is acceptable. 2. Structured JSON-RPC errors were flattened before reaching clients (review n.2). HandleError used a type switch, so any *JSONRPCError wrapped by fmt.Errorf fell through to the generic Internal error: code 1100 and message lost, plus one ERROR log line per rejection (~hundreds/sec while open). Switch to errors.As so wrapped typed errors survive; the generic path for plain errors is unchanged (regression-tested). 3. Half-open probe could be answered by a stale request (review n.3). Allow() now returns an opaque ProbeToken and Record() requires the matching token to decide the half-open outcome. Tokens are generation-numbered so a token from a previous cycle is rejected; a liveness guard re-admits a probe if the previous one never records, so half-open cannot deadlock. 4. Breaker parameters are now configurable via environment variables (review n.4): JUSSI_UPSTREAM_CIRCUIT_ENABLED / _WINDOW_SECONDS / _FAILURE_RATE / _MIN_SAMPLES / _OPEN_DURATION_SECONDS / _JITTER_FRACTION, wired through the standard viper binding table. Defaults match the previous hardcoded tuning; disabling the breaker (ENABLED=false) makes it a pass-through. Also fix malformed object-form ttls/timeouts entries in TEST_UPSTREAM_CONFIG.json so it parses under loadUpstreamConfig (found while testing the env override path).
Review follow-ups pushed (commit e78fe40)All four review points are addressed in the second commit on steemit/jussi#266. 1.
|
…ker state in /health Closes out the remaining review items: Integration tests (review item 3): a new handler-level suite exercises the breaker against a real httptest upstream — - failure window trips the breaker and requests stop reaching the upstream while it is open (hit counter proves zero leakage) - after the upstream heals and the open period lapses, a single probe gets through, succeeds, and closes the breaker so traffic resumes - rejections surface as the structured 'Upstream temporarily unavailable' JSON-RPC error (code 1100) with breaker context in data.details, not the generic Internal error - upstream.circuit.enabled=false is a true pass-through Per-host metric labels (review item 4): Registry.For now returns the breaker's registry key (scheme://host) alongside the breaker, and the jussi_upstream_circuit_state / jussi_upstream_circuit_rejects_total series use that key instead of the full URL, so all URL variants of one upstream produce exactly one series. /health observability: the health payload now includes circuit_states (per-upstream breaker state), plus circuit_degraded and circuit_worst_state flags when any breaker is open or half-open. This gives the ELB health path an explicit, scrape-independent view of breaker state alongside the Prometheus gauge.
Dashboard panels, Prometheus alert rules, and the cross-service view for the circuit breaker, tuned against the 2026-09-18 incident profile. Written for the watchtower stack (Prometheus + Grafana + OpenObserve); no new jussi-side collection is required.
Background
During the 2026-09-18 production incident, hivemind's DB connection pool saturated and every request jussi forwarded waited out its 3s deadline, got cancelled, and was immediately replaced by the next arrival. Jussi logged ~1.3M
context deadline exceedederrors in 2 hours; the hivemind internal ALB recorded those cancelled requests as ~1.3M ELB-side 400s (they never reached a target); wallet SSR requests behind it hit 60s nginx timeouts and users saw steemitwallet.com 504s. Recovery only happened after the environment was scaled out — i.e. when the pool got enough slack for in-flight SQL to finish faster than new work arrived.That is the gap this PR closes: when an upstream saturates, jussi has no mechanism to stop feeding it. A circuit breaker converts saturation into fast failures, which stops the arrival burst, lets the in-flight upstream work drain, and lets the backend self-heal — the same mechanism that scaling provided, without needing a human at 7am.
What this adds
A per-upstream-host circuit breaker (
internal/upstream/breaker.go), wired intoRequestProcessor.callHTTPUpstream:Default tuning: 30s sliding window, 50% failure rate with a 20-sample minimum (so a handful of timeouts never trips it), 10s open period with 25% jitter (so multiple jussi instances don't probe a recovering backend in lockstep). Values chosen against the incident's traffic profile; all configurable via
BreakerConfig.Design notes:
RetriableUpstreamError, so the existing bounded retry cannot re-enter an open breaker.Observability
jussi_upstream_circuit_stategauge per upstream host (0=closed, 1=open, 2=half-open)jussi_upstream_circuit_rejects_totalcounter per upstream hostCodeUpstreamResponseErr, message "Upstream temporarily unavailable"), keeping them in the same client-visible error class as other upstream failuresTesting
internal/upstream/breaker_test.gocovering: stays-closed under low failure rate, trips at threshold, min-samples guard, rejects-while-open, single-probe half-open, close-on-successful-probe, reopen-on-failed-probe, window sliding, and registry keying. All pass.go build ./...,go vet, and the fullgo test ./...suite pass.Operational follow-ups (not in this PR)
configfiles.json/08-upstream.config) can still be tuned independently — e.g. raisingget_statespecifically — and composes with the breaker: the breaker handles saturation, the timeout handles per-request budgets.