DeepSeek
DeepSeek Long-Context Flakiness, Measured: Phase Sensitivity in Chunked KV-Cache Compression
Short answer
- Researchers from ByteDance Seed, with collaborators at Princeton, Stanford, and UC Berkeley, have isolated a structural cause behind DeepSeek's intermittent long-context failures. DeepSeek-V4's chunked KV-cache compression compresses tokens at a fixed stride of 4, and retrieval accuracy depends on where a token sits inside that cycle — its "phase." At 128K context, the same fact can differ by up to 40 percentage points in retrieval accuracy depending on phase alone.
- Base checkpoints are hit hardest: DeepSeek-V4-Flash-Base averaged 52.6% on needle-in-a-haystack retrieval with a best-to-worst phase gap of 40.23 points. Post-training narrows the gap but does not close it. DeepSeek-V4.1-Flash was the steadiest model tested, averaging 92.38% with a gap of 6.09 points.
- Average benchmark scores conceal the problem entirely. A retrieval task that works today can fail after a two-token edit to the prompt, with no error raised. There is no official fix as of October 10, 2026 — treat any single-position retrieval check as unreliable, and test across multiple token offsets instead.
If you have run a long-context pipeline on DeepSeek, you have probably met the version of the model that Chinese developers call "chōufēng," roughly "acting up." The same document, the same question, and retrieval that lands cleanly on one run and whiffs on the next. A paper quietly posted to arXiv on September 28, 2026 now gives that flakiness a mechanism, a name, and a number. The name is phase sensitivity, and the number is a 40-point swing.
The paper, Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression (arXiv:2609.36322), comes from a ByteDance Seed team working with researchers from Princeton, Stanford, and UC Berkeley. Chinese tech media picked it up on October 9 under headlines like "ByteDance Seed finds why DeepSeek 'glitches.'" The substance is more interesting than the headline, and it matters to anyone evaluating or building on long-context models.
What chunked KV-cache compression does, and what it breaks
Long-context inference in transformers gets bottlenecked by the KV cache: memory and attention costs grow with context length. DeepSeek-V4's answer, like a growing family of efficiency designs, is chunked KV-cache compression: the model groups consecutive tokens into fixed-size windows and compresses each window into fewer cache entries at a fixed stride. For DeepSeek-V4's compressed sparse attention (CSA), that stride is S = 4 tokens. Fewer entries, less memory, cheaper attention.
The side effect is a new positional coordinate the pre-LLM era never had: a token's phase, its position relative to the compression-window boundaries. The team measured retrieval performance as a function of phase and found a systematic asymmetry. The same information is easy to retrieve at one phase and hard at another, and the pattern repeats with the compression stride. That periodic variation is what they call phase sensitivity.
Two details make this more than a curiosity. First, it is not a DeepSeek-specific bug: the team pretrained a family of transformers from scratch on a Qwen3-0.6B backbone across multiple KV-compression designs and reproduced phase sensitivity in every compressed variant. Second, the control behaves: DeepSeek-V3.1-Base, which has no chunked compression, showed the positional reversal in only 4 of 64 tested inputs, with tiny margins. The mechanism tracks the compression, not the brand.
The numbers: 40 points between neighbors
The core evaluation is a controlled needle-in-a-haystack task at 128K context, scored by answer-prefix accuracy across eight residue groups (the target key's position modulo 8, spanning two stride-4 cycles). The headline finding: in the DeepSeek-V4 family, retrieval accuracy differs by up to 40 percentage points across phases.
| Model | Avg. accuracy | Best-to-worst phase gap |
|---|---|---|
| DeepSeek-V4-Flash-Base | 52.6% | 40.23 pp |
| DeepSeek-V4-Flash-0731 (post-trained) | 91.4% | 19.14 pp |
| DeepSeek-V4-Pro-Base | 72.1% | 34.77 pp |
| DeepSeek-V4-Pro-0813 (post-trained) | 84.4% | 14.84 pp |
| DeepSeek-V4.1-Flash | 92.38% | 6.09 pp |
Read the first row the way an engineer experiences it. A base checkpoint that averages 52.6% is not "a mediocre model." It is a model that is genuinely good at retrieving some facts and genuinely bad at retrieving the same facts two tokens away. The average is a fiction that both groups of prompts are being flattened into. Post-training clearly helps, with accuracy climbing and gaps shrinking in every paired comparison, but a 14-to-19-point gap surviving post-training in V4-family models is not noise. DeepSeek-V4.1-Flash, evaluated on 2,560 prompts per residue group, was the steadiest of the five at 92.38% mean and 6.09 points of spread.
The uncomfortable part is the averaging. Every summary metric you normally look at (a LongBench-style score, a RAG eval average, a vendor dashboard) is a mean over positions. Phase sensitivity lives in the variance across positions, which those summaries throw away. The paper's blunt framing: high average accuracy can coexist with systematic positional failures.
The demo that makes it concrete: completions flipping every 4 tokens
Before the haystack runs, the paper opens with a smaller, nastier demonstration on code completion. The authors hand the models a function with fillers of varying length, and the preferred completion alternates between two answers, flipping with a period of exactly four tokens, matching the CSA stride. Move the cursor two tokens and the model's answer changes; move it back and it changes again. The reversals held across four filler families.
The obvious guess — that the damage comes from key-value pairs straddling a window boundary — does not survive the paper's own analysis. In a reference model with window 8 and stride 8, phases zero through six all share a window with their values, and accuracy still varies by more than 50 percentage points among them. The mechanism the authors trace instead is phase specialization: different attention components learn to carry different phases, some concentrating on a subset of positions, so the load is uneven by construction. Gradient-flow analysis in idealized retrieval models suggests sharp phase specialization is what training dynamics naturally favor.
What this means if you build on the DeepSeek API
Don't confuse the two "caches." DeepSeek's API-side context caching is a billing feature — repeated prefixes hit disk cache and cost less. Chunked KV-cache compression is inside the model weights. One changes your invoice, the other changes what the model can find. Both can be in play at 128K, and they are unrelated.
Intermittent retrieval failure is not necessarily your prompt, and not necessarily the service. If a long-context call silently drops a fact it retrieved an hour ago, phase sensitivity is now a documented candidate. This kind of failure never raises an API error — it returns a confident wrong answer, which makes it easy to misdiagnose. Our error-code reference covers the failures that announce themselves; this is the failure that doesn't.
Test across offsets. The paper's methodological point is the practical one: evaluating a compressed model requires measuring across phases. Concretely, when you validate a prompt template, shift the payload by a few tokens and run it again — at 128K, a template that works at one padding length can fail at the next. (This offset-sweep practice is our engineering recommendation, not the paper's; the paper stops at the measurement methodology.)
Duplicate what matters. Until there's a fix, the boring mitigation for critical facts is redundancy: state the key constraint in more than one place in the prompt rather than trusting a single insertion point to land on a good phase. Again, engineering folk wisdom applied to a newly measured problem.
If you're picking a variant for long-context work: V4.1-Flash is the one this paper measures as steadiest, and it's also the variant our launch coverage and deployment guide already track. A 6.09-point spread is not immunity, but it's a different world from 40.
Caveats, honestly stated
The paper evaluates open-weight checkpoints: base and post-trained releases. What a live API deployment does (routing, configurations, any server-side mitigation) is not visible from here, so treat these numbers as the model's behavior, not a contract about any specific endpoint. The finding generalizes to any chunked KV-cache compression design, but the 40-point figure is specific to the DeepSeek-V4 family at 128K on this needle-in-a-haystack task. As of October 10, 2026, we are not aware of an official DeepSeek response or a user-facing fix; if one lands, the calculus above changes. This page is a snapshot of one paper's measurements on one day, and we checked every number above against the arXiv source.
Sources: Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression (arXiv:2609.36322, submitted September 28, 2026) · IT之家, October 9, 2026 (Chinese-language coverage). Numbers verified against the paper's abstract, Section 2, Table 7 discussion, and Appendix C on October 10, 2026.