Don't build business logic on an expiry notification
Some designs look elegant on a whiteboard and fall apart the first time you ask how they fail. The one I want to describe was like that.
We needed to do some important background work at the end of a customer’s session: take what they’d entered and send it to a downstream system. The session data already lived in a distributed store with an expiry time. So the first proposal was neat: when the store tells us a session key has expired, that’s the signal to do the work.
Why an expiry notification is a bad trigger
It sounds efficient. You don’t need a scheduler, the store already knows when things expire, and you just listen. But once we walked through the failure cases, three problems showed up.
- Everyone hears it. Notifications from shared infrastructure can reach every application instance listening, and on shared infrastructure, sometimes more than one environment. Without extra coordination, the same piece of work gets done twice.
- Nobody might be listening. If the application is restarting or mid-deployment at the moment the key expires, the notification can arrive with nobody there to handle it. And by definition, the data it refers to has just gone.
- You can’t easily reason about it. When something goes wrong, there’s no record of work that was due, only an event that may or may not have been received. You can’t replay it, and it’s hard to prove what happened.
The underlying issue is that an expiry notification is a side effect of how the cache manages memory. It was never designed to be a reliable business event, and building on it means inheriting all its delivery quirks.
Separate recording the work from doing it
The design we moved to splits the job in two.
Recording the work. When a session has enough information to create a work item, the latest payload is stored with its own expiry, and its identifier and due time go into a queue ordered by time. Now the work exists somewhere durable, independent of any running process.
Doing the work. A scheduled job runs at a configurable interval, reads the items that have reached their due time, fetches their payloads, sends them downstream and removes them from the queue when they succeed. If the application restarts, nothing is lost. The next run picks up where the last one left off.
A few details mattered more than they looked:
- The processing window. The payload has to outlive the customer’s visible session by long enough for the scheduler to reach it. Getting that gap right, between when the browser session ends and when the stored data expires, is what makes the design work.
- Claiming work. With several instances and parallel deployment slots, two processors could pick up the same item. The design uses a lock or lease per item, so only the one that claims it does the work. Where infrastructure is shared across environments, each environment gets its own queue.
- Knowing it works. We defined the monitoring before the code: logging the whole path from session to downstream submission, alerts for scheduler failures, checks for items that expired before being processed, detection of duplicates, and a reconciliation check that compares qualifying journeys against submitted work. A silent drop in volume is as important a signal as an error.
Scope the first version on purpose
The legacy system did this job with a more elaborate mechanism, and there was a pull to rebuild all of it. I pushed for a narrower first phase: process one item at the end of the session, and nothing more. That kept the first release small enough to test properly and get right.
Why we didn’t go fully event-driven, yet
Others on the team rightly raised the cleaner end state: a durable message broker with a serverless consumer. That would give us proper retries, dead-lettering and better observability out of the box.
We agreed to defer it. It meant new infrastructure and wider architectural change before customers saw any benefit. The queue-and-scheduler pattern reused capabilities the platform already had, and it removed the riskiest failure modes now. It also leaves a clean seam: recording the work is already shaped like publishing an event, so moving to a broker later replaces the scheduler, not the whole design.
That’s the trade-off I’d make again: remove the failure mode first with what the team can run today, and keep the door open for the bigger design.
The questions I now ask
Whenever background work hangs off a user action, I ask:
- What triggers this, and was that trigger designed to be reliable?
- Where does the work exist between “due” and “done”?
- What happens if two processes pick it up at once?
- What happens if nobody’s running when it’s due?
- How would we know if it silently stopped?
Most designs survive those questions with small changes. The ones that don’t are far cheaper to fix in a design review than after an incident.