Two production workflows that looked green and quietly weren't. Turned into boring, provable reliability.
A green health check sitting on top of a dead pipeline is the worst failure mode there is. A shallow ping was reporting success while the workflow wrote zero rows for hours.
Retry policies on all 9 HTTP nodes so transient timeouts stop becoming outages. Idempotent dedup so retries never double-process. Service names instead of hardcoded IPs so a reboot is a non-event. A health check rewritten to alert on any errored execution in 24 hours. Secrets moved out of workflow JSON and rotated.
This is what an audit actually finds: retries, dedup, false-green monitoring, and hardcoded assumptions. The fixes are unglamorous and they are the whole job.