Projects & experiments

# Things I’m building.

## [Case studies](https://boiko.ai/work/#case-studies-heading)

- ### [AI Reliability Layer](https://boiko.ai/work/ai-reliability-layer/)

  A shared layer for monitoring, diagnosis, and automated recovery across production agents and automations.

  Immigrant Invest · Production reliability

- ### [Engineering Coordination Layer](https://boiko.ai/work/engineering-coordination-layer/)

  A multi-agent system that helps me maintain a high standard of project coordination by connecting discussions, decisions, and project updates to clear tasks, owners, and next steps.

  Immigrant Invest · Engineering operations

- ### [Company Brain](https://boiko.ai/work/company-brain/)

  Shared company context from 17 internal systems, available to employee agents and internal products through permission-aware access.

  Immigrant Invest · Context infrastructure

- ### [PII Sanitization Pipeline](https://boiko.ai/work/pii-sanitization/)

  A shared pipeline for preparing confidential client and partner data for AI workflows, with field-level policies and local processing before downstream agents receive it.

  Immigrant Invest · Data privacy

- ### [Presales Automation](https://boiko.ai/work/multi-agent-presales/)

  A multi-agent presales system that qualifies leads, answers questions using a knowledge base, and schedules sales calls, with shared context across voice, WhatsApp, email, and SMS.

  Immigrant Invest · Presales automation

- ### [Document Intelligence Platform](https://boiko.ai/work/document-intelligence-platform/)

  An agent-routed platform that prepares 100K+ confidential documents for downstream systems through local OCR and LLM processing, selective Azure fallback, and legal review.

  Immigrant Invest · Document intelligence

- ### [Quote Agent](https://boiko.ai/work/quote-agent/)

  A multi-agent quoting system that interprets site assessments, compares purchasing options, and prepares costed proposals with code-based checks and human approval.

  Paul Kick · Renovation operations · 2024

## [Personal projects](https://boiko.ai/work/#personal-heading)

- ### [pgwarden](https://boiko.ai/work/pgwarden/)

  Give AI assistants Postgres access with per-person permissions, PII masking, human-approved writes, and an auditable query history.

  MCP infrastructure · Database access control

  [pgwarden on GitHub](https://github.com/B0yko/pgwarden)

- ### [proof-of-done](https://boiko.ai/work/proof-of-done/)

  Check a coding agent’s verification claims against command results after its latest edits, with a Claude Code stop hook and a transcript audit CLI.

  Developer tools · Verification evidence

  [proof-of-done on GitHub](https://github.com/B0yko/proof-of-done)

- ### [agent-claimcheck](https://boiko.ai/work/agent-claimcheck/)

  Check AI agent success claims against tool receipts and state evidence, with calibrated decisions and a review queue for uncertain cases.

  Agent evaluation · Claim verification

  [agent-claimcheck on GitHub](https://github.com/B0yko/agent-claimcheck)

- ### [Tern](https://boiko.ai/work/tern/)

  Search a video archive by what was said, shown, or written on screen. Built to run locally on Apple Silicon.

  Multimodal search · Desktop app

  [Tern on GitHub](https://github.com/B0yko/tern)

- ### [taskdistill](https://boiko.ai/work/taskdistill/)

  Turn a recurring LLM task into a small local model, with measured quality and automatic escalation to the original API.

  Model distillation · Local inference

  [taskdistill on GitHub](https://github.com/B0yko/taskdistill)

- ### [booking-truth](https://boiko.ai/work/booking-truth/)

  Check whether an AI booking agent did what it promised, using calendar-state evaluation, injected failures, and a guarded reference agent.

  Agent evaluation · Booking reliability

  [booking-truth on GitHub](https://github.com/B0yko/booking-truth)

- ### [wellbrief](https://boiko.ai/work/wellbrief/)

  Ask questions across drilling reports and build pre-drill risk briefs, with traceable figures and source quotes. Runs offline by default.

  Offline retrieval · Decision support

  [wellbrief on GitHub](https://github.com/B0yko/wellbrief)

## [Research & experiments](https://boiko.ai/work/#research-heading)

- ### [8 DGX Sparks: LLM Serving Capacity](https://boiko.ai/work/dgx-spark-serving-capacity/)

  An empirical study of workload-dependent concurrency, prefill interference, and interconnect performance across an eight-node inference fleet.

  Inference systems · September 2026

Source: [https://boiko.ai/work/](https://boiko.ai/work/)
