When we talk about resilient systems, the conversation often jumps straight to exotic tooling — service meshes, chaos monkeys, distributed tracing. But the foundation of resilience is simpler: design for failure.

Start with the Basics

Before reaching for a circuit breaker library, ask: does this call even need to be synchronous? Can we queue it? Can we retry with exponential backoff? Often, the answer reveals a simpler architecture.

async function retryWithBackoff(fn, maxRetries = 3) {
  for (let i = 0; i < maxRetries; i++) {
    try { return await fn(); }
    catch (e) {
      if (i === maxRetries - 1) throw e;
      await sleep(Math.pow(2, i) * 1000);
    }
  }
}

Event Sourcing as a Safety Net

Event sourcing gives you a complete audit trail and the ability to rebuild state. It's not just for banks — any system where "what happened?" matters benefits from it.

The real power isn't in the event store itself, but in the mindset shift: state is derived, events are truth.