For most of the last decade, the default answer to when should this run? was now. Streaming pipelines, event buses, live dashboards, sub-second APIs. Latency was the scoreboard, and batch processing carried the faint smell of mainframes and overnight tape.
Then teams started putting model inference in the critical path, and the arithmetic changed. A model call is not a database read. It costs real money per invocation, it varies wildly in duration, and it fails in ways retries do not always fix. Wrap that in a synchronous request and you have handed your user’s loading spinner to a dependency you do not control.
So the batch job came back — quietly, wearing new clothes. We call it a queue, a job runner, an agent inbox, a nightly enrichment pass. The pattern is the same one the mainframe people had: collect the work, do it when conditions are favourable, notify when done.
The interesting part is what it buys you beyond cost. Batched work can be reordered, deduplicated, cached against near-identical inputs, and — crucially — reviewed before it lands. A human can look at two hundred draft outputs in a queue far more cheaply than at two hundred interruptions.
The design question worth asking on every AI feature is not how fast can this be? It is does anyone actually notice if this takes four minutes? More often than product instinct suggests, the answer is no. And where the answer is no, asynchronous design gives you cheaper, calmer, more inspectable systems.
Real-time is a requirement. It was never supposed to be a personality.
