All work

Immigrant Invest · Production reliability

AI Reliability Layer

A shared layer for monitoring, diagnosis, and automated recovery across production agents and automations.

Active systems and automations
~60
Executions per month
~120k
Cases resolved by remediation agents
~80%
Reported execution failure rate
<0.1%

The challenge.

At Immigrant Invest, a small engineering team supports around 60 agent systems and automations handling roughly 120,000 executions a month. In a legal-services business, operational failures carry consequences beyond a broken workflow. Keeping these systems dependable could not rely on engineers manually reviewing every run.

I designed and built a shared reliability layer to monitor execution, investigate incidents, and recover failed operations across the portfolio.

Deterministic recovery. Contextual investigation.

The system combines deterministic checks and recovery routines with monitoring and remediation agents. Retries and fallbacks handle predictable failures. When those mechanisms are insufficient, agents use execution logs and MCP gateways to investigate the affected systems, plan a fix, and carry it out.

  1. Capture

    Log and normalize every execution into a common format for monitoring and investigation.

  2. Detect

    Identify system, provider, and API errors. Escalate unusual execution times, token consumption, and looping behavior immediately.

  3. Recover

    Attempt deterministic retries and fallbacks, then route unresolved incidents to remediation agents for diagnosis and repair.

  4. Hand off

    Report resolved incidents to engineers and system owners. Give developers the investigation findings when agent-led recovery is unsuccessful.

  5. Retain

    Add fixes to a regression dataset so production incidents inform future reliability testing.

Operational results.

Production snapshot
Production executions
32,706Up 4.69%
Failed executions
7Down 61.11%
Failure rate
0%Down 0.1 pp
Reported time saved
628hUp 294h
Avg. run time
15.74sDown 0.7s
Workflow dashboard snapshot. Failure rate rounded as displayed.

Remediation agents resolve approximately 80% of the cases they handle within five minutes. Engineers and system owners receive a report of the completed work; unresolved cases reach developers with the investigation already prepared.

The reported execution failure rate across the monitored systems is below 0.1%. This measures operational execution failures, not the correctness of every AI-generated response.

More capacity for engineering.

The layer takes recurring investigation and recovery work off the engineering team’s hands. Cases that still need a developer arrive with the diagnostic work already done, reducing the time spent reconstructing what happened.

Its business value is continuity: restoring the operations that depend on these systems while keeping maintenance manageable for a small team. The regression dataset also preserves incident knowledge beyond the immediate fix.