Quality before cost.
Running a smaller model locally is useful only if it can do the job well enough. local-enough evaluates local models, cloud models, and non-LLM baselines on five back-office tasks: CRM extraction, intent classification, PII redaction, meeting summarization, and vendor matching.
The benchmark records task quality, output validity, latency, throughput, per-task cost, and an estimate of hardware and energy costs. It selects routes on a calibration split and reports held-out test results separately.
Turn measurements into a route.
A generated plan connects the evaluation to an OpenAI-compatible router. Deterministic checks reject malformed results, and requests can fall back when a selected route fails. A local-only setting refuses cloud escalation; if no local candidate meets the quality bar, the router returns an error by default instead of quietly lowering the standard.
A useful “no” is a result.
In the published reference run, a local option cleared the quality bar for two of five tasks: intent classification through a TF-IDF baseline and vendor matching through a small local model. The best local PII-redaction result reached 92.2% against a 95% bar, so the local-only route declined that task.
The reference workload suggested a lower-cost mixed route, but no single cloud model met every task bar. The cost comparison is a scenario estimate, not a measured production saving.
Where the evidence stops.
Four of the five datasets use templates, and the reference local runs used one Apple Silicon machine, English inputs, a dated price snapshot, and configured rather than measured power consumption. The router is a reference implementation and has no built-in authentication. Each new workload needs its own quality bar and evaluation.