The challenge.
At Immigrant Invest, a small engineering team supports around 60 agent systems and automations handling roughly 120,000 executions a month. In a legal-services business, operational failures carry consequences beyond a broken workflow. Keeping these systems dependable could not rely on engineers manually reviewing every run.
I designed and built a shared reliability layer to monitor execution, investigate incidents, and recover failed operations across the portfolio.
Deterministic recovery. Contextual investigation.
The system combines deterministic checks and recovery routines with monitoring and remediation agents. Retries and fallbacks handle predictable failures. When those mechanisms are insufficient, agents use execution logs and MCP gateways to investigate the affected systems, plan a fix, and carry it out.
Capture
Log and normalize every execution into a common format for monitoring and investigation.
Detect
Identify system, provider, and API errors. Escalate unusual execution times, token consumption, and looping behavior immediately.
Recover
Attempt deterministic retries and fallbacks, then route unresolved incidents to remediation agents for diagnosis and repair.
Hand off
Report resolved incidents to engineers and system owners. Give developers the investigation findings when agent-led recovery is unsuccessful.
Retain
Add fixes to a regression dataset so production incidents inform future reliability testing.
Operational results.
- Production executions
- 32,706Up 4.69%
- Failed executions
- 7Down 61.11%
- Failure rate
- 0%Down 0.1 pp
- Reported time saved
- 628hUp 294h
- Avg. run time
- 15.74sDown 0.7s
Remediation agents resolve approximately 80% of the cases they handle within five minutes. Engineers and system owners receive a report of the completed work; unresolved cases reach developers with the investigation already prepared.
The reported execution failure rate across the monitored systems is below 0.1%. This measures operational execution failures, not the correctness of every AI-generated response.
More capacity for engineering.
The layer takes recurring investigation and recovery work off the engineering team’s hands. Cases that still need a developer arrive with the diagnostic work already done, reducing the time spent reconstructing what happened.
Its business value is continuity: restoring the operations that depend on these systems while keeping maintenance manageable for a small team. The regression dataset also preserves incident knowledge beyond the immediate fix.