All work

Inference systems · September 2026

8 DGX Sparks: LLM Serving Capacity

An empirical study of workload-dependent concurrency, prefill interference, and interconnect performance across an eight-node inference fleet.

DGX Spark systems
8
Configured request slots
40
Text streams · median ≥20 tok/s
~11
Code streams · median ≥20 tok/s
~28

Research question.

How much concurrent work can a local inference fleet sustain at an interactive generation rate? I evaluated eight DGX Spark systems with English prose and Python workloads, using a median per-stream rate of 20 generated tokens per second as a capacity reference. The aim was to establish practical concurrency limits and identify the bottlenecks behind them.

Experimental setup.

Eight NVIDIA DGX Spark systems, each with a GB10 processor and 128 GB of unified memory, host four deployments. The four-node DeepSeek deployment uses a direct-cabled ring and a modified NCCL build. All three model configurations use speculative decoding.

Inference fleet configuration
ModelNodesServing configurationSlots
DeepSeek-V4.1-Flash4SGLang · tensor and expert parallelism16
GLM-5.3-Flash2vLLM · EXL3 4-bit8
Qwen3.8-Flash-Next1 + 1vLLM · NVFP4 · independent replicas8 + 8

Measurement protocol.

Synchronized waves of 1–16 streams use distinct prompts from sets of 16 Python tasks and 16 English writing topics. Requested output lengths are 1,000 tokens for code at temperature 0, and 300–600 tokens for prose at temperature 0.7. Reasoning is disabled. Each concurrency level has a warm-up wave and one to five repeated trials.

Each stream’s decode rate is (output tokens − 1) divided by the interval between its first and last token. The figures take the median across streams within each trial, then the median across trials. Queueing and prefill are assessed separately through time to first token. Aggregate throughput divides the sum of (output tokens − 1) across streams by the shared interval from the earliest first token to the latest last token. It therefore differs from concurrency multiplied by the median stream rate.

Concurrency depends on the workload.

Code generation sustains higher concurrency at the selected median-rate threshold in every tested deployment. Summing the deployment capacities gives approximately 11 English-text streams or 28 code streams under their respective test conditions. These are alternative workloads, not a simultaneous mixed-load result. Higher speculative draft acceptance on code is consistent with the advantage, although temperature and output length also differ.

English text: median stream rate falls as concurrency rises. The 20-token-per-second reference admits 2 DeepSeek streams, 1 GLM stream, and 4 Qwen streams per replica. The DeepSeek two-stream observation uses shorter outputs and one trial.Python code: median stream rates stay at or above 20 tokens per second at 8 DeepSeek streams, 4 GLM streams, and 8 Qwen streams per replica. Qwen has no reported two-stream code measurement.
Median decode rates; dashed lines mark 20 tok/s. Markers are observations and connecting lines guide the eye. Qwen represents one node; its two-stream code point was not reported. DeepSeek’s two-stream text point uses 300 output tokens and one trial; the other DeepSeek text points use 600 tokens and two trials.
Estimated concurrency at the median 20-token-per-second decode reference
DeploymentEnglish textPython code
DeepSeek · 4 nodes~2 streams¹8 streams
GLM · 2 nodes1 stream4 streams
Qwen · each of 2 replicas4 streams8 streams

¹ DeepSeek’s two-stream text estimate depends on the shorter-output observation. The 600-token measurements alone place the threshold between 1 and 4 streams. At four code streams, GLM’s median is 21.1 tok/s and its minimum is 19.4 tok/s.

Prefill can dominate responsiveness.

A 32k-token uncached prompt on Qwen reduced concurrent decode rates to roughly 4 tok/s for about 16 seconds. New short requests waited 13–15 seconds for their first token. On DeepSeek, reusing a cached 40k-token document reduced time to first token from about 17 seconds to 0.4 seconds.

These observations make prompt reuse and prefill isolation relevant to capacity planning. A decode-only concurrency estimate does not capture the delays introduced by long incoming documents.

Diagnosing an interconnect bottleneck.

The four-node ring initially sustained only 12.5–13 Gb/s per link. RDMA loopback reproduced the slowdown without a cable, directing the investigation toward the host-side NIC path. System logs associated the affected hosts with NIC hot-plug reinitialization. Rebooting with cables connected restored link performance; the initialization path remained the working root-cause explanation.

After link recovery, DeepSeek’s aggregate code throughput at 16 streams improved by 27%, and prefill throughput roughly doubled.

Reported measurements before and after network recovery
MeasurementBeforeAfter
Link throughput~13 Gb/s111.6 Gb/s
DeepSeek aggregate decode · 16 code streams180 tok/s228 tok/s
DeepSeek prefill1.15–1.2k tok/s2.2–2.4k tok/s

Implications and limitations.

The results favor workload-aware routing and separate budgets for decode concurrency and long-context prefill. Configured request slots alone are insufficient for interactive capacity planning. Model selection still requires an independent evaluation of output quality.

The 20 tok/s reference applies to median decode rate and does not guarantee a minimum speed for every request. Prompts are synthetic, repetition counts vary, and the figures report point estimates without confidence intervals. English-text results from one Qwen replica are extrapolated to the second; full agent workflows and output quality were not evaluated.

GLM code measurements were collected on 25 September and its text measurements on 26 September with the same configuration but a different server session. Output length also matters: in a separate GLM comparison at eight streams, extending responses from 400 to 1,000 tokens reduced the rate from 19.8 to 17.3 tok/s.

DeepSeek’s four- and sixteen-stream code records flag output-length validation failures; individual completion lengths were not retained in those exports. Long-context tests include prompts up to approximately 86k tokens and 40–56k-token cache checks, but do not establish performance at the full configured context limits.

DGX Spark systems and networking equipment on the lab desk.