GeneralBackendDevOps/CloudHigh signal

Fault Tolerance

A system's ability to keep working correctly even when part of it fails — a server crashes, a network link drops — instead of the whole system going down.

Think of it like

Like a building staying structurally sound even if one support beam fails, because the load was designed to redistribute across the others.

Junior or senior?

Junior sounds like

Claims their system is 'fault tolerant' as a buzzword.

Senior sounds like

Can describe a specific real failure it tolerated gracefully, and how.

Ask them

“Tell me about a real failure your system survived because of how it was designed. What would've happened without that design?”