Batch vs Stream and Data Modeling

Choose when work runs and shape data for correctness first, then deliberately denormalize proven read bottlenecks.

Batch runs a bounded set later; streaming handles events continuously while data modeling controls read and write costBatch runs a bounded set later; streaming handles events continuously while data modeling controls read and write cost

ChoiceUse whenCost
BatchReports, backfills, daily billingResults arrive later
StreamFraud alerts, live dashboardsOrdering/replay/deduplication
Normalized tablesIntegrity and flexible writesReads may join
Denormalized read modelMeasured read path is costlyDuplicates must stay in sync

Streams need an ordering key, idempotency, retention and DLQ. Batch jobs need isolation so reporting cannot starve checkout. A denormalized projection needs an owner and a measurable freshness delay.