Atlas Automation Field Note · #02
Lessons from building, breaking, and repairing real automation.
When the API Went Silent, We Stopped Trusting the Live Request
Why reliable automation needs a last-known-good fallback — not endless retries.
01The Request Failed. The Content Hadn’t.
While building and operating the AtlasProfitAI publishing system, we ran into a situation that sounds obvious once it is written down. Our site reads its articles from a separate content system. At times, that live request failed — the content system was temporarily unavailable, slow to answer, or telling us to slow down.
Our original design made a quiet assumption: if the live request fails, the page fails. A reader would see an empty list or a “not found” page for an article that existed and had already been checked.
But the article had not changed. Nothing about it had become wrong. The only thing that had failed was our ability to fetch it at that moment. We had been treating two different facts — “we could not reach the source” and “the content is not valid” — as if they were the same fact.
Live request failure ≠ content invalidity
Separating dependency availability from content validity became the starting point for how we rebuilt this part of the system.
02The Retry Trap
The first instinct when a request fails is to try again. Retries are genuinely useful: many failures are transient, and a second attempt a moment later often succeeds. Common causes include:
- temporary network problems between two services
- timeouts while an upstream service is busy
- short periods of upstream instability
- rate limiting, where a service deliberately refuses extra requests
The trap is retrying without limits. If a service is refusing requests because it is under pressure, every immediate retry adds to that pressure. When many requests do this at once, the result is a retry storm: the recovery behavior itself keeps the service from recovering.
There is a second, quieter problem. A retry that eventually returns something does not prove that what came back is right. In Field Note #01, The Automation Worked. The Result Was Wrong., we described how a successful response can still hide a wrong outcome. The same applies here.
Retry is a recovery tool — not a proof of correctness
03What “Last Known Good” Actually Means
“Last known good” is easy to misread as “whatever we cached last.” That is not what we mean, and the difference matters.
Last response
Whatever the source returned most recently — which might be an error page, an empty list, a partial result, or a correct one. It has not necessarily been checked.
Last-known-good state
A previously retrieved or produced state that passed the system’s required checks and may be used temporarily, under defined conditions.
A last-known-good state is not:
- arbitrary cached content
- old data assumed to be correct forever
- a substitute for validation
It earns its status by passing checks, and it keeps that status only for as long as serving it is safe.
04The Architecture We Moved Toward
Conceptually, the design has two paths. The live source is still the first choice every time. The stored state is a safety net, not a replacement.
Normal path
- Live request
- Success?
- Yes → validate
- Serve
- Update last-known-good
Temporary failure path
- Live request
- Temporary failure
- Bounded retry / cooldown
- Fall back to last-known-good
- Recover live source later
One rule sits underneath both paths: a stored state is only replaced by a result that has itself passed validation. An error, an empty response, or a partial result is newer, but it is not better. Writing it over a verified state would turn a temporary problem into a lasting one.
Newer ≠ better
Never replace known-good data with known-bad data
Equally important is what the fallback does not do. A request for something that genuinely does not exist should still be answered honestly as “not found.” Fallback covers a source we cannot reach, not content that was never there.
05Why a Cooldown Can Be Better Than Aggressive Retrying
When a service rate-limits or refuses requests, it is communicating something: it needs fewer requests, not more. Responding with an immediate burst of retries ignores that signal.
We moved toward three habits:
- Bounded retries — a small, fixed number of attempts, then stop trying for this request.
- Backoff — waiting longer between attempts instead of retrying instantly.
- Cooldown — after repeated refusals, pausing live requests for a short period and serving the last-known-good state meanwhile.
The point is not a specific number of attempts or seconds; the right values depend on the service and the content. The point is that the system has a defined limit, and that reaching the limit leads to a safe, predictable behavior rather than more load.
06When Fallback Is Appropriate — and When It Isn’t
Fallback is not a universal good. It works when an older, verified answer is safer than no answer. It is wrong when an older answer could mislead someone into a harmful decision.
Often appropriate
- editorial content and articles
- documentation
- verified public information that does not need second-by-second freshness
- non-transactional reference material
Often inappropriate
- financial balances
- inventory availability
- security or access state
- medical or emergency information
- transactions and payments
- anything where stale information could materially harm a user
For the second group, failing visibly and clearly is usually the honest choice. A slightly older article is still a useful article. A slightly older account balance can be a wrong account balance.
Resilience must not come at the cost of truth
07How This Changed the Way We Think About Reliability
We used to measure reliability with one question: did the API answer?
We now ask a different one: can the reader still receive a known-valid result when a dependency temporarily fails? That shift affected more than one feature. It changed how we think about:
- Validation — a result has to pass checks before it is trusted or stored.
- Fallback — deciding in advance which content may be served from a stored state.
- Recovery — returning to the live source automatically once it is available again.
- Observability — knowing when the system is running on fallback, rather than having it fail silently in either direction.
- Dependency isolation — keeping one service’s bad moment from becoming the whole site’s bad moment.
We are a small team, and none of this makes the system immune to failure. It makes failures easier to contain. If you are deciding where safeguards like these belong in your own workflows, our AI Automation Risk Scanner walks through the questions we use, including whether a workflow has a safe-stop and a recovery path.
08A Practical Last-Known-Good Checklist
Before adding a fallback to a workflow, answer these questions explicitly:
- 01What exactly qualifies as “known good” in this workflow?
- 02What validation must a state pass before it earns that status?
- 03How long can that state safely remain useful to a reader?
- 04Which failures permit fallback — timeouts, rate limits, temporary unavailability?
- 05Which failures must stop the workflow instead of falling back?
- 06Can a failed, empty, or partial response ever overwrite good data?
- 07How does the system know when to return to the live source?
If any answer is “we don’t know,” the fallback is not ready yet. An unclear fallback can be worse than none, because it hides problems instead of absorbing them safely.
09Closing
The live request will sometimes fail. That is a normal property of systems that depend on other systems. What matters is deciding, ahead of time, what a failure is allowed to mean for the person on the other end.
Reliable automation is not a system that never encounters failure. It is a system that knows which failures should reach the user — and which ones it can safely absorb.
Atlas Automation Field Notes describe lessons from building and operating AtlasProfitAI’s own systems. Read how we work in our research methodology and editorial policy.