Building Server Actors on Effect Clusters
Tanner Scadden
Sep 21, 2026
The most expensive part of serving a hospital schedule is building it the first time. A schedule is the result of combining a large amount of operational state: who is working, who is unavailable, which shifts need coverage, where staffing requirements have changed, and how those constraints interact across a unit and date range. Most changes affect only a small part of that picture. Someone calls out, a shift is reassigned, time off is approved, or a staffing requirement changes. Recomputing the entire schedule for each of these changes would repeat a lot of work whose result has not changed.
To accommodate for the above, we built storage for our actors (which are powered by effect cluster), while keeping postgres as the source of truth with fast bootups as shard ownership is reassigned. Drain a snapshot to S3 periodically, read the snapshot on boot, then catch up to the current state as needed using a cursor.
Avoiding state divergence in the projections (reread, don’t apply a delta).
Projections can derive from full events or incremental diffs. We originally chose diffs for their compact size, but the cost was high: emitters must understand projection structure to compute removals, and ordering/idempotency guarantees remain hard to enforce. We found engineering was often sinking time into how did it get into this state and dealing with races, and decided to pivot to full events instead, with a twist.
The principle here is: events should be oriented around the domain that needs to be revalidated, and they are used to know what to re-fetch from Postgres on each read, keeping current state authoritative. This prevents inconsistencies that can arise from dropped or out-of-order delta events.
To support this, we built around an event cursor: a position marker in the event stream that lets any actor catch up from any historical point. All state changes are recorded in a shared domain_events table, written atomically in the same transaction as the in-memory projection update.
Since we keep track of the unit that is affected by this mutation, and what dates are not invalid from this change. This design empowers us to do a few things pretty simply:
Actor polls a table for specific sources and unitId(s) as needed
The actors can invalidate only specific sources and times within their projections
Be replica lag resistant. As you poll the db for a txid that is greater than the cursor you have, and it is not there yet, that is ok. You’ll keep polling and progressing as it catches up.
Can keep actors as real time as possible by sending a ping to bypass the polling interval and check immediately
We can catch up to the current state in bulk by aggregating the sources and their unions to refetch only what is needed to be caught up.
Pathway to S3 as the storage for actors
We keep projections in memory to avoid write amplification to Postgres and to prevent connection storms during deployments or shard rebalancing, when thousands of actors might move the assigned server simultaneously.
We evaluated multiple storage options, but hit a critical constraint: schedule projections frequently exceed 10MB of JSON—far beyond Redis's capacity with its single-threaded architecture when thousands of writes come in at the same time. Inspired by Cursor's Git at Any Scale approach, we adopted a similar pattern: drain the snapshots to S3 with compression after writes, modeling the state as:
This enables us to:
Enable efficient recovery during shard transfers: actors load the precomputed snapshot and replay only missed events from the cursor forward, minimizing startup time and avoiding full recompilation.
Support breaking changes gracefully: old projection versions can be safely discarded, since event recovery depends only on the cursor position, not the snapshot schema.
The best part, the snapshots own nothing, and are fully disposable. If no snapshot is found or corrupted, build from Postgres. Similarly, a snapshot failing to save has limited consequences. Actors continue to operate normally, and the next process will either start from an older snapshot and catch up, or rebuild from postgres with a slight performance hit.
Making it the default
We put these rules into DurableEntity, a wrapper around our Effect Cluster entities. A projection provides the domain-specific parts: how to construct itself from Postgres and how to refresh its state when something changes. The wrapper handles loading snapshots, catching them up, saving new ones, and falling back to reconstruction.
This design is boring, and that is part of the beauty. There is enough innovation happening around changing how hospitals do scheduling & staffing while adopting Effect Clusters early. Postgres remains the source of truth, we ate the cost of some performance in favor or eliminating race conditions, and everything can reuse the same building blocks.