Skip to content

Rank the quick-switch list on ids and pins, then load only the members it returns - #319

Merged
SiteRelEnby merged 1 commit into
mainfrom
fix/top-fronters-member-projection
Sep 24, 2026
Merged

SiteRelEnby merged 1 commit into
mainfrom
fix/top-fronters-member-projection

Conversation

@SiteRelEnby

Copy link
Copy Markdown
Contributor

GET /v1/members/top-fronters (the quick-switch list) has had a slow tail for a while: p50 around 40 ms, p99 parked at the 2.5 s bucket edge for hours at a stretch, at about fifty requests an hour. That shape is one client's every request being slow, not a slow query for everyone.

What it wasn't. The scoring query. Seeded a scratch database with 51k fronts over two years and ran EXPLAIN ANALYZE on the exact statement: 19 ms, using the (system_id, ended_at) index through a BitmapOr, 12.6k rows in the window.

What it was. The line after it. The handler loaded every member of the system as a full ORM row, encrypted bio and note included, to sort by pin and score and return at most eight. With 1,500 members carrying 20 KB bios that load is 151 ms on a fast local box; production's largest system has 15,682 fronts and 1,641 groups, and a roster to match, on a slower box, polled every time the quick-switch list opens.

The change. The ordering runs on a projection of id and quick_switch_pin for the whole roster (1.7 ms for the same 1,500 members), then the full rows are loaded for the ids that survive the limit, in ranking order. Same result, same ordering rules, same tiebreakers; the existing endpoint tests and the SQL/Python parity test pin that, and a new test checks the returned rows are fully hydrated (bio present) rather than the projection leaking through.

Plus one gauge. sheaf_system_member_count_max and its sheaf_systems_by_member_count distribution, next to the other per-system maxima. The roster size is what every load-everyone endpoint scales with, and until now there was no way to see the largest one without a scratch database.

Not confirmed on production (that would need log_min_duration_statement for a day); the fix is correct and cheap regardless, so ship it and watch the route's p99.

…s it returns

GET /v1/members/top-fronters scored the roster in the database already
(one aggregate query, tens of milliseconds against fifty thousand
fronts), then loaded every member of the system as a full ORM row,
encrypted bio and note included, to sort by pin and score and return at
most eight. On a large system with long bios that hydration was the
entire cost of the request, and it ran on every quick-switch poll: the
route's p50 sat at 40 ms while a slow tail of about one request in six
took two seconds, which is one big system's client polling.

The ordering now runs on a projection of id and quick_switch_pin for the
whole roster, and the full rows are loaded afterwards for the ids that
survive the limit, in ranking order. Same result, same ordering rules,
same tiebreakers; the existing endpoint and parity tests pin that.
Measured on a scratch database with 1,500 members carrying 20 KB bios:
151 ms for the old load, 1.7 ms for the projection.

Also adds sheaf_system_member_count_max and sheaf_systems_by_member_count
next to the other per-system maxima and distributions. The roster is
what every load-everyone endpoint scales with, and there was no way to
see the largest one without a scratch database.
@SiteRelEnby
SiteRelEnby merged commit 037cf45 into main Sep 24, 2026
24 checks passed
@SiteRelEnby
SiteRelEnby deleted the fix/top-fronters-member-projection branch September 24, 2026 02:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant