All work

Immigrant Invest · Data privacy

PII Sanitization Pipeline

A shared pipeline for preparing confidential client and partner data for AI workflows, with field-level policies and local processing before downstream agents receive it.


Preparing confidential data for AI.

At Immigrant Invest, client and partner information appears in CRM records, applications, documents, call transcripts, messages, and internal notes. In a legal-services business, these sources contain personal identifiers and confidential information that downstream agents should not receive by default.

I put a shared sanitization pipeline in place as part of data ingestion and normalization. It prepares a separate representation for agent use, applying sensitivity rules before the data becomes prompt context or enters an agent-facing retrieval index.

Start with the data structure.

Known sensitive fields in CRM records, lead forms, and other structured inputs are explicitly classified. Deterministic policies exclude those fields and their values from agent inputs by default. Their treatment follows the schema instead of depending on a model to recognize each value.

Free-text fields need content analysis even when they sit inside a structured record. A note or conversation can contain sensitive information anywhere in the text.

Process unstructured content locally.

Documents, transcripts, and communications require analysis of the content itself. This combines pattern-based detection with local language-model processing to assess information in context. Documents are handled with on-premise OCR and local models, including DeepSeek.

The local processing stage has access to the source material so it can identify sensitive content. Downstream business agents receive the prepared representation. Keeping document processing local and controlling what reaches those agents are separate parts of the design.

  • Records

    CRM, databases, applications

  • Documents

    Locally extracted text and OCR

  • Conversations

    Messages and call transcripts

  • Notes

    Internal written context

Ingestion and normalization

Apply field-level policies and analyze unstructured content locally.

Prepared data

A sanitized representation for downstream agent workflows and retrieval.

Sanitization sits upstream of agent context and retrieval indexing.

Prepare the data before execution.

Moving sanitization into ingestion makes the prepared data reusable across agent runs while the source content and sanitization policy remain unchanged. The expensive analysis does not need to be repeated every time an agent executes.

Changes to the data, policy, or detection configuration require reprocessing. New messages, changed CRM values, and fresh tool responses still need to pass the same boundary before becoming agent context; a previously processed source does not make future updates safe automatically.

A shared standard for data preparation.

The pipeline gives the team a common place to manage sensitive-data handling across workflows. Its value is reducing unnecessary exposure and repeated processing while keeping data preparation consistent as new agents and sources are added.

Sanitization does not replace access controls. A processed record can still contain confidential context, so access to it still needs to be restricted to the workflows and people that need it.