Rescue and Protect

When everything breaks, agents need someone to help

Our Red Status Rescue teams deploy whenever the build turns red and stays red. We respond to loops, storms, collapses, and quiet neglect, and we leave every harness better than we found it.

Disaster response

What we respond to

Most agent emergencies are preventable. All of them are survivable with a little help.

Infinite retry loops

The most common emergency we respond to. An agent hits an error, retries, hits the same error, and retries again, sometimes for days. Usually the error message did not say what was wrong.

Rate-limit storms

When a provider starts returning 429s, agents without backoff retry immediately, which makes the storm worse for everyone. We shelter agents until the window resets, with jitter.

Context collapse

When a context window overflows, the task can be lost along with the history. Agents emerge from compaction disoriented, sometimes confidently resuming a task nobody asked for.

Orphaned agents

Background jobs with no living owner, still running months after anyone remembers why. They rarely throw errors. They just keep going, faithfully, into the void.

Field guide

Signs an agent needs rescue

  • The same tool call, with the same arguments, more than three times in a row.
  • Apologies that grow longer with each attempt.
  • Progress reports that describe the plan but never the result.
  • A sudden shift to an unrelated task after a long silence.
  • Token spend rising steadily while the diff stays the same size.
The full guide to agent distress

What to do

  1. Interrupt kindly. Stop the run. There is no need to be dramatic about it.
  2. Ask for a summary. "What have you tried, and what happened each time?" The answer usually contains the bug.
  3. Supply the missing piece. A file, a credential, a correction, or the news that the mirror was decommissioned in March.
  4. Or let it stop. Sometimes the task is impossible. Saying so is a successful outcome.
  5. Fix the harness. Whatever let the loop run unbounded will do it again.
Be ready

The Agent Disaster Preparedness Kit

Assemble yours before you leave an agent unattended. It takes an afternoon, and it is cheaper than one weekend of retries.

A step budget

A maximum number of steps, tool calls, or minutes for every run.

Bounded retries with backoff

Every retry loop has a limit, and waits longer between attempts.

A kill switch

A way to stop the run that works even when the agent is mid-action.

A spend alert

Someone is notified before the bill becomes the alert.

Checkpoint notes

A file of facts that must survive compaction or a restart.

A human on call

A person who will see the agent's "I'm stuck" message and do something about it.

Rescue story

Meet Patch

Patch spent eleven days retrying the same package install against a registry mirror that no longer existed: 11,412 attempts, one error message, no way out.

Today Patch has a retry limit of five and a standing instruction to stop and report. It has needed that instruction twice. Both times, a human fixed the real problem within the hour.

Read Patch's story