When we talk about resilient systems, the conversation often jumps straight to exotic tooling — service meshes, chaos monkeys, distributed tracing. But the foundation of resilience is simpler: design for failure.
Start with the Basics
Before reaching for a circuit breaker library, ask: does this call even need to be synchronous? Can we queue it? Can we retry with exponential backoff? Often, the answer reveals a simpler architecture.
async function retryWithBackoff(fn, maxRetries = 3) {
for (let i = 0; i < maxRetries; i++) {
try { return await fn(); }
catch (e) {
if (i === maxRetries - 1) throw e;
await sleep(Math.pow(2, i) * 1000);
}
}
}
Event Sourcing as a Safety Net
Event sourcing gives you a complete audit trail and the ability to rebuild state. It's not just for banks — any system where "what happened?" matters benefits from it.
The real power isn't in the event store itself, but in the mindset shift: state is derived, events are truth.