A smaller model for a specific job.
Classifying a support message or extracting fields from an invoice rarely needs a general-purpose model on every request. taskdistill captures existing API traffic, prepares training data, and fine-tunes a small Qwen2.5 model with LoRA and MLX on Apple Silicon.
The pipeline handles duplicate removal, train/validation/test splits, and checks for data leakage. An OpenAI-compatible server then answers locally, escalating requests to the original model when confidence falls below the selected threshold.
Choosing when to escalate.
Model selection, checkpoints, calibration, and the escalation threshold use validation data. Selection functions reject test splits at runtime. Reports compare the student, teacher, cascade, and simpler baselines on held-out data, with confidence intervals and cost and latency estimates.
What the evaluation showed.
On 3,075 Banking77 test examples, the cascade agreed with recorded teacher answers 97.6% of the time while escalating 23.7% of requests. Agreement measures how closely it follows the teacher; accuracy against the dataset labels is reported separately.
The invoice experiment exposed a different outcome: field-level F1 against the teacher fell to 94.1% on unseen test layouts, below the 97% validation target. Both results are included in the published reports.
Current scope.
Single-turn classification and JSON extraction, with replayable teacher responses, dataset cards, and per-run reports. The local training path targets Apple Silicon; the project also includes a PyTorch backend tested on CPU.