Treat connectivity as a budget

An AI feature that depends on every upstream service being reachable is only as available as the weakest link in that chain. Before choosing a model or interface, decide which operations must continue locally, which can wait, and which must stop safely.

This is especially important when a system sits inside communication, accessibility, hospitality, or organizational workflows. A network failure should not leave people guessing whether an action completed or whether a human is now responsible.

Four practical patterns

  • Keep the critical path narrowLimit the services required for the action the user needs right now. Optional enrichment should not block essential work.
  • Make retries recoverablePersist work before transmission, use idempotent operations, and make delayed synchronization observable.
  • Expose degraded stateShow when data is stale, automation is paused, or a dependency is unavailable instead of presenting false confidence.
  • Preserve human controlDefine who owns the next action when automation cannot complete it, and retain enough context for a safe handoff.

What changes for AI systems

AI adds uncertainty to ordinary distributed-systems failure. A delayed request can become a duplicate action; missing context can produce a confident but inappropriate answer; and a model fallback can silently change quality.

The response is not to describe the system as resilient. It is to design observable boundaries: version the context used for a decision, record the origin of important data, make automated actions distinguishable from human ones, and escalate when confidence or connectivity falls below the workflow’s safe threshold.

The design test

Ask what the user sees after the network disappears midway through an action. If the system cannot explain what completed, what remains queued, and who is now responsible, its failure behavior is not finished.