- Published on
The Cached-Read Bottleneck: Speeding Up DeepSeek V4.1 Engram on Four DGX Sparks
- Authors

- Name
- Yusheng Zheng (云微)
- @yunwei37
The Cached-Read Bottleneck: Speeding Up DeepSeek V4.1 Engram on Four DGX Sparks
When a model server stalls on storage, the obvious story is that the SSD is too slow. In this case, that story was wrong.
We were serving DeepSeek V4.1 Flash across four NVIDIA DGX Spark systems. Its two Engram n-gram tables are roughly 200 GiB, too large to keep in host memory without taking away the unified memory needed by the GPUs. Our serving stack therefore kept the tables on local NVMe and read only the rows needed by each step.
The implementation was sensible: split the work into chunks, submit them to a 32-thread Python pool, and wait for the reads to finish. Yet many of those rows were already in Linux's page cache. The actual storage work was often tiny; creating futures, waking threads, contending on the GIL, and collecting completions cost much more than the cached reads themselves.
The final change was small. Try cached rows synchronously with preadv(..., RWF_NOWAIT), retry only the rows that report EAGAIN, and send the residual cold rows to the existing blocking pool. The old path remains the fallback on kernels or filesystems that do not support NOWAIT.
The isolated hot-read component became as much as 8.30x faster. That number is real, but it is not a model-throughput result. We deployed the patch through the production Flux path, repeated endpoint tests, and then watched real traffic. Short fixed workloads improved substantially; long-prefill TTFT did not; unmatched production windows were too different to support a general throughput claim.
That gap between an exciting microbenchmark and a defensible production conclusion is the main subject of this post.
The system we were trying to fit
The serving topology was:
OpenAI-compatible gateway
|
DeepSeek V4.1 Flash
vLLM TP=4
/ / \ \
DGX Spark DGX Spark DGX Spark DGX Spark
rank 0 rank 1 rank 2 rank 3
\_________ ConnectX ring _________/
each rank reads its Engram rows from NVMe
The checkpoint is a very large MoE model with two Engram lookup tables. On DGX Spark, CPU and GPU share the same 128 GiB memory pool. Materializing approximately 200 GiB of Engram tables in host memory is therefore not merely expensive: it makes the model impossible to serve in the intended configuration.
The disk-backed reader solved the capacity problem. Each tensor-parallel rank read its own table rows from a local file and dequantized them before copying the staging result to the GPU. A process-wide ThreadPoolExecutor with 32 workers provided queue depth for cold reads.
Conceptually, the old path looked like this:
for each Engram job:
divide requested rows into chunks
submit every chunk to a 32-thread pool
wait for every future
continue the GPU forward pass
This is a good shape when most reads reach physical storage. It becomes expensive when most requested pages are already resident.
Start at the API, then move down the stack
We did not begin with a favorite profiler or a proposed code change. We started with the user-visible service and moved downward until the evidence localized one blocking layer.
The investigation used five views:
| Layer | Tool or evidence | Question |
|---|---|---|
| service | fixed OpenAI-compatible requests and Prometheus | Is the server healthy, and what is the end-to-end baseline? |
| process | pidstat and process counters | Is host CPU scheduling material? |
| syscall | strace timing plus file-descriptor mapping | Which calls stall and which files do they touch? |
| VM/page cache | mincore, refaults, reclaim, and PSI | Are pages resident or under memory pressure? |
| storage/component | iostat and real-offset byte comparisons | Is this bandwidth, latency, or software orchestration? |
The active ranks already had perf, strace, bpftrace, iostat, pidstat, and NVIDIA profiling tools. We did not enable a full GPU trace first. Once the CPU-side call chain showed the forward pass waiting on a synchronous host reader, tracing downstream GPU idleness would have added detail without changing the immediate diagnosis.
That choice matters in practice. Profiling is not a contest to collect every possible trace. The best next tool is the one that separates the remaining hypotheses.
The first clue: slow tails, even without reclaim
Syscall timing showed repeated positional reads in the request path. In one 30-second sample, the median read was only 221 microseconds, but P99 was 8.9 milliseconds. Forty-three percent of observed decode batches contained at least one call slower than 4 milliseconds.
The batch waited for every future, so a handful of cold pages set the latency of the whole Engram step.
Memory evidence ruled out a simple “the machine is swapping” explanation. After cleaning stale shared-memory artifacts from an older process, a rank showed no kswapd or direct reclaim activity and zero short-window memory PSI. File refaults remained high, and the Engram tail remained.
Then a smaller measurement exposed a second cost. An isolated batch took 9.4 milliseconds even though the sum-derived syscall service time was about 0.94 milliseconds. The difference was not NVMe bandwidth. It was Python future construction, submission, wakeups, GIL handoffs, and completion processing.
Cached replays made the gap obvious:
| 642-720 row reads | Time |
|---|---|
| existing 32-thread path | about 5.2 ms |
serial blocking preadv | about 0.9 ms |
serial RWF_NOWAIT | about 0.65-0.75 ms |
More threads could change cold-read queue depth, but no thread count removes the overhead of scheduling cached reads through an executor.
The measurement mistake that nearly fooled us
The first hot/cold comparison classified a page by attempting a NOWAIT read and then measured the candidate on the same page. That is invalid: the classification operation can itself initiate I/O or change residency before the timed arm begins.
This kind of error is easy to miss because the benchmark still looks careful. It has two arms, real files, real offsets, and precise timers. The problem is causal: measurement changes the state being measured.
We replaced that classifier with mincore, which observes page residency without reading the page. The corrected harness selected separate samples for the old and new arms with the same requested resident/cold mix and used separate destination buffers.
mincore is still page-granular, and residency can change between classification and execution. The corrected method reduces the race; it does not make page-cache state immutable. That limitation belongs beside the result, not in a footnote added later.
The candidate: use the kernel page cache as the cache
The production candidate deliberately did not add a user-space row cache. On a unified-memory machine, duplicating page-cache contents would consume the same memory the model and KV cache need.
The algorithm was:
for every requested row:
try preadv(..., RWF_NOWAIT)
if complete: keep it
if EAGAIN: put it on a pending list
retry only the pending rows once with RWF_NOWAIT
send rows still pending to the existing blocking thread pool
if NOWAIT is unsupported:
run the unchanged blocking path for the whole batch
The second NOWAIT pass is an optimization, not a correctness requirement. On this kernel and filesystem, a first EAGAIN was often followed by an immediate hit, consistent with I/O starting while the first call refused to wait. If the second call still reports EAGAIN, the blocking path guarantees progress.
The implementation also preserves the old full-read loop. Partial reads continue until the row is complete; a zero or terminal short read raises instead of returning corrupt bytes.
We rejected several larger alternatives:
- A user-space cache duplicated the kernel page cache and increased unified-memory pressure.
- More Python read threads did not remove cached-read scheduling overhead.
- Parallelizing the four logical Engram jobs was slower in the component harness.
io_uringwas technically attractive, but the image had noliburingor Python binding; adding a C/FFI dependency was disproportionate to this fix.- GPUDirect Storage was not available in the running environment.
The useful property of the selected change is that every unsupported or unresolved case flows into the reader that was already serving production.
Component correctness before performance
We compared 3,000 candidate reads byte-for-byte against the existing reader using real model files, real row offsets, and real row sizes. There were zero mismatches.
Forced-path tests covered:
| Case | Expected behavior | Result |
|---|---|---|
| cached row | finish in the NOWAIT path | passed |
one EAGAIN, then hit | finish on the second pass | passed |
persistent EAGAIN | use the blocking pool | passed |
| partial read | continue until complete | passed |
| NOWAIT unsupported | use the old whole-batch path | passed |
| terminal short read | raise an error | passed |
Only after these checks did we compare timing.
The component result
The corrected mixed-residency comparison improved on all four ranks:
| Rank | 50% classified-cold | 100% classified-cold |
|---|---|---|
| 0 | 5.78x | 2.35x |
| 1 | 3.62x | 2.22x |
| 2 | 2.00x | 2.57x |
| 3 | 6.69x | 1.94x |
The all-cold improvement does not mean NOWAIT made the SSD faster. The candidate quickly completes rows that become available, avoids futures for rows already present, and reserves the pool for the residual set.
The largest cached prefill-shaped component case represented 8,192 tokens and 196,608 total row-read calls across four logical jobs:
| Reader | Runs | Median |
|---|---|---|
| existing 32-thread path | 1041.573, 1052.327, 1056.587 ms | 1052.327 ms |
| serial NOWAIT path | 126.787, 129.530, 106.160 ms | 126.787 ms |
That is the 8.30x result. Resource counters explain why:
| Metric | Existing | Candidate |
|---|---|---|
| wall time | 1060.682 ms | 132.480 ms |
| user CPU | 713.406 ms | 70.017 ms |
| system CPU | 1638.170 ms | 62.461 ms |
| voluntary context switches | 296,162 | 0 |
This proves a hot-page orchestration problem. It does not prove an 8.30x model speedup.
Deploying one change through production
The prototype changed only the tracked Engram reader. We kept the image, four-rank topology, 32-thread fallback, model arguments, storage, cache configuration, credentials, and service routing unchanged.
The change was committed, pushed, and reconciled by Flux. The four ranks mounted the new generated patch ConfigMap. The head became Ready 7 minutes 44 seconds after creation, and all four containers remained at zero restarts.
Before rollout, the exact source compiled and passed the forced failure paths. After rollout, health checks and a real request through the normal production gateway succeeded.
Keeping the rollout narrow mattered. If we had changed thread counts, cache sizes, model batching, and the reader together, even a large improvement would have had no defensible cause.
Fixed endpoint A/B
The fixed requests used the real server, temperature zero, thinking disabled, and exactly 64 output tokens.
| Workload | Before | After | Observed change |
|---|---|---|---|
| C1 median output rate | 12.07 tok/s | 35.49 tok/s | 2.94x |
| C1 median TTFT | 0.712 s | 0.245 s | -65.6% |
| C1 median total time | 6.015 s | 2.082 s | -65.4% |
| C4 median per-request output rate | 8.48 tok/s | 20.28 tok/s | 2.39x |
| C4 wall aggregate | 28.49 tok/s | 62.96 tok/s | 2.21x |
| 13,316-token prompt TTFT | 5.889 s | 5.989 s | +1.7% |
| 13,316-token prompt total time | 7.687 s | 6.616 s | -13.9% |
The short paths improved materially end to end. The long-prefill TTFT did not improve. Its 100 ms increase is within the uncertainty of one before/after sample, but it is still evidence against a universal latency claim.
There was another confounder: the old short baseline began with two ambient requests, while the new short run began idle. The endpoint numbers prove the deployed path can be faster; they do not isolate the exact share caused by NOWAIT.
Real traffic gives a less convenient answer
The production service was not an idle benchmark machine. During one four-request probe, eight requests were active, including large-context work. Each small request completed successfully, but took about 112 seconds.
Another 64-token request started behind five long requests. It reached first output after 597.722 seconds, then decoded at 9.24 tok/s and completed after 604.649 seconds. It did not fail. The optimized reader removed one cost, but it did not solve head-of-line delay from large active requests.
Ten-minute Prometheus windows immediately before and after rollout were also heterogeneous:
| Production metric | Before | After |
|---|---|---|
| completed requests/s | 0.10 | 0.09 |
| prompt tokens/s | 27,748.54 | 6,922.03 |
| generation tokens/s | 45.85 | 51.80 |
| average / max running | 3.27 / 5 | 3.75 / 8 |
| max waiting | 2 | 6 |
| TTFT p50 / p95 | 2.21 / 14.67 s | 2.06 / 73.20 s |
| inter-token latency p50 / p95 | 0.18 / 0.36 s | 0.17 / 0.31 s |
| end-to-end latency p50 / p95 | 9.91 / 88.50 s | 19.29 / 189.00 s |
| request errors | 0 | 0 |
Generation tokens per second rose about 13%, while prompt traffic fell 75% and queue pressure increased. These windows prove real traffic and initial correctness. They cannot establish a fleet-wide throughput improvement because the requests were different.
What 24 hours of production did establish
At the 24-hour follow-up, all four current ranks were Ready with zero restarts. Prometheus reported, rounded because range-query boundaries are interpolated:
| 24-hour evidence | Value |
|---|---|
| normally completed requests | 3,921 |
| error / abort outcomes | 0 / 0 |
| prompt tokens | 1.064 billion |
| generation tokens | 4.378 million |
| scheduler preemptions | 0 |
| recent one-hour ITL p50 | 168 ms |
| recent one-hour TTFT p95 | 4.86 s |
This is strong stability and real-workload evidence. It still lacks an equivalent 24-hour old-reader control with the same prompt lengths, output lengths, cache-hit rates, and concurrency. So the production conclusion remains:
- the optimized reader is deployed and stable;
- fixed short requests improved substantially;
- long-prefill TTFT was essentially unchanged in the measured pair;
- unmatched production traffic showed no errors and slightly higher generation rate, but worse tail latency under heavier load;
- a matched, stratified canary is still needed for a general throughput claim.
Upstream moved while we were measuring
DeepSeek V4.1 support in vLLM was moving quickly during this work. The official tracking issue #56400 lists several Engram paths.
The closest official proposal is vLLM PR #56757, which adds file-backed Engram tables using memory-mapped shards. As of September 19, it remains open and needs a rebase. It does not contain a RWF_NOWAIT cached-page fast path.
Adjacent work includes:
- PR #56926, serialized offloaded lookups and huge-page packing;
- PR #57023, Mooncake-backed tables; and
- PR #57651, sharing host tables across colocated DP replicas.
These solve important placement and sharing problems. None duplicates the explicit Python row-read scheduling bottleneck measured here.
The most useful upstream contribution is therefore not another broad disk-offload RFC. It is a focused result on #56757 after its rebase, with the corrected residency method, byte-equality tests, CPU/context-switch evidence, and production A/B. The exact Python preadv patch belongs first in the four-DGX-Spark recipe that owns this reader.
The mistakes worth remembering
- Assuming storage was the bottleneck. Cached reads were cheap; Python scheduling was expensive.
- Using a classifier that changed page-cache state. A benchmark can be precise and still be causally invalid.
- Counting only visible content in a reasoning stream. DeepSeek may stream reasoning in a separate field; the client had to count both.
- Treating an 8.30x component result as a model result. The full server has queueing, prefill, decode, collectives, and other host work.
- Comparing unmatched production windows. Higher output rate alongside lower prompt traffic and deeper queues is not a clean A/B.
- Ignoring saturation. A faster component cannot eliminate head-of-line blocking behind long requests.
- Adding a second cache too early. Linux already had the relevant pages; duplicating them would consume scarce unified memory.
- Reaching for the largest profiler first. A short syscall-to-page-cache evidence chain localized the problem without restarting the service for a full GPU trace.
A reusable profiling pattern
The method generalizes beyond Engram:
start with a real request
-> identify the blocking process phase
-> map slow syscalls to files and source
-> separate residency, reclaim, and storage service
-> reproduce the exact component with real offsets
-> verify bytes and failure paths
-> deploy one change
-> repeat fixed endpoint tests
-> observe production traffic without overclaiming
The key is to keep the claim at the same level as the experiment. Component benchmarks explain mechanism. Fixed endpoint tests show practical effect. Production counters show stability and workload behavior. They become a causal production-throughput claim only when the workload distributions are matched.
Final result
The optimization stayed in production because it preserved bytes, retained the proven blocking fallback, improved fixed short-request performance, and served more than a billion prompt tokens in the first 24-hour window without errors, aborts, preemptions, or rank restarts.
The best number was not 8.30x. It was the causal chain:
disk-backed Engram made the model fit
-> cached rows still went through a 32-thread Python pool
-> syscall and page-cache evidence separated I/O from scheduling
-> RWF_NOWAIT completed hot rows without futures
-> unresolved rows retained the old blocking path
-> fixed endpoint workloads improved after deployment
-> 24-hour production soak stayed correct and stable
-> the general throughput claim remains open pending a matched canary
That last line is part of the result. A useful systems optimization report should explain not only why a patch helped, but also what the measurements still cannot prove.