RESEARCH · #005

Trying “DFlash,” a Diffusion-Model Approach to Parallel Draft-Token Generation, on Gemma

In the concept edition and the implementation/benchmark edition, we covered a speed-up technique for LLM generation called MTP. This time, we verify DFlash, an unusual draft model that uses diffusion to predict multiple tokens at once, comparing its real-world performance against Gemma's native Assistant model. Part 3 of 3.

2026.09.11 11 min read Part 3 of 3

In the concept edition and the implementation/benchmark edition, we covered a speed-up technique for LLM generation called MTP (Multi-Token Prediction). To recap briefly: a lightweight "draft model" predicts a handful of tokens ahead of time, and the main model checks them all at once. When the guesses are right, you leap ahead several tokens in a single step, which is what makes the whole thing feel faster.

There’s more than one way to build that draft model, and the one we’re looking at this time, DFlash, takes an unusual approach: it uses a diffusion model — the kind of technique you’d normally associate with image generation — to predict multiple tokens all at once instead of one at a time. It claims to support a wide range of models and to significantly outperform EAGLE-3, an existing approach. Those are the claims worth testing directly.

The catch is that benchmarks like this are usually measured in an environment the vendor sets up. It’s harder to find a case where someone ran it on their own GPU, against an opponent that already has a dedicated, well-optimized MTP model of its own — Gemma-4’s native Assistant model.

So this time, using the same setup as the implementation/benchmark edition (an RTX 3060 with 12GB VRAM, llama.cpp, the same JavaScript coding task), we directly compared DFlash against Gemma-4-12B-it’s Assistant model. The short version: DFlash did not outperform the Assistant model. The reasons are technically clear, though, and they draw a fairly clear picture of where DFlash is strong and where it isn’t — that’s what we’ll dig into below.

Overview of New Techniques

Two new token-prediction techniques appeared in quick succession:

DSpark was released to speed up inference specifically for DeepSeek’s own DeepSeek-V4, and it only supports DeepSeek-V4 / DeepSeek-V4-Flash.

DFlash, on the other hand, supports a much wider range of models, each with its own dedicated DFlash model. Like Google’s Assistant model, it’s designed to be bolted on as an add-on to achieve token prediction.

Z-Lab is a research group led by Zhijian Liu [2], an assistant professor at UC San Diego who is also a research scientist at NVIDIA. The lab works across the algorithm, systems, and application layers to make AI smaller, faster, and more efficient.

This time, we wanted to understand how much of a speed-up DFlash actually delivers, verified through hands-on testing.

How DFlash Predicts Tokens

DFlash is introduced in the paper "DFlash: Block Diffusion for Flash Speculative Decoding" [3].

Speculative token prediction itself isn’t new — the earliest paper on the idea was published by Google DeepMind in 2023, "Accelerating Large Language Model Decoding with Speculative Sampling" [4]. Improvements continued quietly from there, culminating in 2025 in an MTP draft model called EAGLE-3. Even so, it apparently never escaped the autoregressive paradigm, and in the end didn’t deliver a dramatic speed improvement.

Meanwhile, diffusion models — the noise-removal mechanism commonly used in image generation — have been making their way into the LLM space. It started with Meta’s LLaDa, and in Japan, KDDI’s ELYZA Lab team released a model called "ELYZA-Diffusion-Instruct-1.0-Dream-7B" [5].

Z-Lab’s DFlash brings that diffusion-model property into the MTP draft model. The basic mechanism is shown below.

[Input Token (N)] --> [Main Model] -- h_on ----------------------------------------> [Predicted: N+1]
                                       |                                             ^
                                       v                                             |
                                  [KVCache]                                          |
                                       |                                             |
+--------------------------------------+--------------------------------+            |
| DFlash Drafter                       v (KV data injection)            |            |
|                                 [KVCache]                             |            |
|                                      |                                |            |
|                              [Diffusion Model]                        |            |
|                                      |                                |            |
|                               [Token Decoder]                         |            |
|                                      |                                |            |
|                            +---------v--------+                       |            |
|                            | Predicted N+2    |                       |            |
|                            |       +          |                       |            |
|                            | Predicted N+3    |                       |            |
|                            |       +          |                       |            |
|                            | Predicted N+4    |                       |            |
|                            +---------+--------+                       |            |
+--------------------------------------+--------------------------------+            |
                                       |                                             |
+--------------------------------------+---------------------------------------------+-------+
| Processing inside Main Model         v                                             |       |
|                             * Uses causal attention to mask;                       |       |
|                               computes probs in parallel                           |       |
|                                      +------------------------------+              |       |
|                                      v                              v              |       |
|                         (Match found)                  (No match)                  |       |
|                         Include matching portion       Nothing included in output  |       |
|                         -> Predicted: N+2                                          |       |
|                         -> Predicted: N+3                                          |       |
+------------------------------------------------------------------------------------+-------+

Figure 1. DFlash mechanism.

Looking at the mechanism, it resembles the Assistant model implemented in Gemma, but the big difference is in how the draft model itself predicts tokens. The latter half — the token-evaluation stage — closely mirrors Gemma’s approach.

With DFlash, the KV cache is built independently. As before, when the main model predicts token N+1, it uses that state to update its own KV cache. Data extracted from that update is then injected into DFlash’s own KV cache, which is where DFlash’s processing begins.

Gemma’s Assistant model runs this prediction step using a very small neural network, sequentially and at high speed, producing as many predicted tokens as needed before handing them off to the evaluation logic.

DFlash, by contrast, doesn’t use an autoregressive model inside its small neural network — it uses a diffusion model. Here, much like generating an image, it produces all of the needed predicted tokens, in the correct order, in one shot. The longer the maximum token length, the longer this takes, but it’s still dramatically faster than doing it sequentially. What follows is the same as before: causal-attention-based masking runs in parallel, and the result determines which tokens are allowed to be output together.

In short: DFlash trades sequential autoregressive generation for a single-shot diffusion process, predicting all candidate tokens simultaneously rather than one by one.
[For Gemma Assistant]
[Input Token (N)] --> [Main Model] --> [Predicted: N+1]
                            |
                            +-> [Predicted: N+2] -(seq)-> [Predicted: N+3] -(seq)-> [Predicted: N+4]
                                                                    |
                                                                    v
                                                     [Verify] --> Up to n OK --> [Confirmed: N+2, N+3...]

[For DFlash]
[Input Token (N)] --> [Main Model] --> [Predicted: N+1]
                            |
                            |   +-- [Predicted: N+2]
                            +---+-- [Predicted: N+3]   <-- (Outputs all at once, order included)
                            |   +-- [Predicted: N+4]
                            |               |
                            |               v
                            +--------->  [Verify] --> Up to n OK --> [Confirmed: N+2, N+3...]

Figure 2. Difference between Gemma’s Assistant model and DFlash.

Running It in llama.cpp

First, get the model files.

hf download google/gemma-4-12B-it
hf download z-lab/gemma4-12B-it-DFlash

This time we’re using llama.cpp build 9850. It’s a good idea to grab the latest build. Use the Python conversion tool included with it to convert each model to GGUF format. Start with the main model.

$ python convert_hf_to_gguf.py \
~/.cache/huggingface/hub/models--google--gemma-4-12B-it/snapshots/cxxxxxx...a/ \
--outtype bf16 --outfile gemma-4-12B-it-b16.gguf

Next, convert the DFlash model to GGUF. The important thing here is --target-model-dir. When converting a DFlash model, it needs to be given a reference to the main model — that’s the parameter it uses to build the converted DFlash output.

$ python convert_hf_to_gguf.py \
~/.cache/huggingface/hub/models--google--gemma-4-12B-it/snapshots/cxxxxxx...a/ \
--outtype bf16 --outfile models/gemma-4-12B-it-DFlash-b16.gguf \
--target-model-dir \
~/.cache/huggingface/hub/gemma-4-12B-it-qat-q4_0-unquantized/snapshots/c202...a/

Quantizing these produces 4-bit models.

$ /opt/llama/bin/llama-quantize \
models/gemma-4-12B-it-b16.gguf models/gemma-4-12B-it-q4_0.gguf q4_0
 
$ /opt/llama/bin/llama-quantize \
models/gemma-4-12B-it-DFlash-b16.gguf models/gemma-4-12B-it-DFlash-q4_0.gguf q4_0

At runtime, add the following arguments. Since the DFlash model acts as an add-on draft/MTP model, use --model-draft to point to it and enable MTP.

/opt/llama/bin/llama-server --model /opt/llama/models/gemma-4-12B-it-q4_0.gguf \
--model-draft /opt/llama/models/gemma-4-12B-it-DFlash-q4_0.gguf \
-t 4 --prio 2 --temp 1.0 --top-p 0.95 --top-k 64 --host 0.0.0.0 --port 8001 \
--device CUDA0 -mg 0 -sm layer --fit on -fa on -c 163840 -ctv q8_0 -ctk q8_0 \
--no-warmup --no-cache-prompt --cache-ram 0 \
--spec-type draft-dflash --spec-draft-n-max 4 --spec-draft-device CUDA0 \
--chat-template-kwargs '{"enable_thinking":true}'

A quick rundown of the arguments:

Verification: Using DFlash on Gemma-4-12B-it

We ran the following simplified verification setup.

The maximum number of predicted tokens is set to 4 (matching the previous verification’s settings, for comparability). We use the llama.cpp web frontend to ask a JavaScript coding task over four turns, measuring token throughput speed.

Resource Usage

Memory usage came out as follows. This run used text-only mode, so multimodal-model requirements are excluded.

Data category Assistant CUDA0 Assistant CPU DFlash CUDA0 DFlash CPU
Weight data 6,390.19 540.00 6,637.69 787.50
KV cache data 1,360.00 — 1,360.00 —
KV cache for Sliding Window 765.00 — 765.00 —
Gated DeltaNet compute buffer 533.80 180.80 533.80 180.80
MTP model weight data 226.90 144.00 390.42 —
MTP token-to-piece cache size — — 1.94 —
MTP model KV cache size — — 640.00 —
MTP model Sliding Window KV cache size — — 136.00 —
MTP model Gated DeltaNet compute buffer 532.78 180.79 519.50 176.02
Total 9,808.67 1,045.59 10,984.35 1,226.38

Table 1. Memory usage breakdown by MTP method (units: MiB).

Of that, the memory used specifically by the MTP model itself was roughly double for DFlash.

Assistant CUDA0 Assistant CPU DFlash CUDA0 DFlash CPU
MTP model memory usage 759.68 324.79 1,687.86 176.02

Table 2. Total memory used by the MTP model (units: MiB).

The main driver here is how the two approaches use the KV cache. The Assistant model shares its KV cache with the main Gemma model. DFlash, by contrast, keeps its own independent KV cache — meaning it has to hold the same cache structure as the main Gemma model a second time. That’s the biggest factor behind the difference.

Token Speed Comparison (Assistant vs DFlash)

As in the previous verification, we compared token ingestion speed turn by turn. There was no major throughput difference between the Assistant model and DFlash for reading tokens in. For output speed per turn, the Assistant model was clearly ahead — DFlash consistently trailed by about 10–20 tokens per second.

Turn Assistant Read (tok/s) DFlash Read (tok/s) Assistant Gen (tok/s) DFlash Gen (tok/s)
1 138.34 53.72 62.80 39.37
2 710.71 652.47 62.04 48.42
3 733.53 739.52 66.68 52.40
4 735.43 708.22 62.77 48.73

Table 3. Token read-in and generation speed by turn.

Overall, inference with DFlash trailed the Assistant model. The Assistant model was already a fairly well-optimized setup going in, which may be part of why the diffusion model’s theoretical advantage didn’t stand out here.

Let’s also look at the token acceptance rate.

Assistant Acceptance Assistant Accepted Assistant Generated DFlash Acceptance DFlash Accepted DFlash Generated
Max 88.27% 2,196 2,648 62.02% 1,892 3,304
Min 75.34% 1,564 2,064 41.03% 1,331 2,812
Avg 81.06% 1,868 2,305 53.42% 1,663 3,129

Table 4. Acceptance rate comparison — Assistant (left) vs. DFlash (right).

This suggests that DFlash’s underwhelming result comes down to its lower acceptance rate. If model tuning progresses further and a faster-processing diffusion model becomes possible, DFlash might eventually surpass the Assistant model.

Trying to Improve Things

Through this verification, we confirmed that DFlash falls a bit short of the Assistant model. But maybe tweaking the parameters would speed things up? With that in mind, we tried raising the number of predicted tokens. We bumped --spec-draft-n-max from 4 to 15.

prompt eval time = 122.83 ms / 31 tokens ( 3.96 ms per token, 252.39 tokens per second)
eval time = 48833.63 ms / 1893 tokens ( 25.80 ms per token, 38.76 tokens per second)
total time = 48956.46 ms / 1924 tokens
graphs reused = 1444
draft acceptance = 0.10901 ( 1174 accepted / 10770 generated), mean len = 2.64

The complete opposite of what we hoped for: acceptance dropped, and speed dropped along with it. It generated a lot more candidate tokens, but almost all of them were rejected — simply lengthening the prediction window clearly wasn’t the answer.

We also tried rebuilding DFlash on Gemma-4-12B-it-QAT. The results were not what we expected. Output was so slow that we stopped measuring after the second turn.

prompt eval time = 475.69 ms / 31 tokens ( 15.34 ms per token, 65.17 tokens per second)
eval time = 102682.33 ms / 1963 tokens (52.31 ms per token, 19.12 tokens per second)
total time = 103158.02 ms / 1994 tokens
graphs reused = 1953
draft acceptance = 0.00025 ( 2 accepted / 7844 generated), mean len = 1.00

This configuration is not viable: throughput dropped further, and the acceptance rate fell to 0.25% — the lowest we observed in this test.

Why Couldn’t It Beat Gemma-4-Assistant?

The likely cause here is that Gemma-4-Assistant’s draft model is simply fast enough that it processed tokens more quickly than DFlash’s draft model.

As covered earlier, the speed advantage a diffusion-based drafter like DFlash offers over a conventional MTP draft model comes from being able to output candidate tokens “all at once.” Against that, Gemma-Assistant has the following working strongly in its favor, which appears to have canceled out DFlash’s advantage:

First, Gemma-Assistant comes with KV-cache sharing built in from the start, while DFlash keeps an independent KV cache that has to be injected fresh every time. That’s an advantage for Gemma-Assistant.

On top of that, Gemma-Assistant’s hidden dimension is only 1,024, with a 4-layer structure. DFlash’s hidden dimension matches the main model at 3,840, with 5 layers — meaning its compute cost is far heavier than Gemma-Assistant’s. However good DFlash’s diffusion model is at generating tokens in a single batch, more compute per step still means more time spent per step.

In this case, we should conclude that Gemma-Assistant simply already had a more optimized setup, and DFlash wasn’t able to get ahead of it.

Conclusion

This time we introduced DFlash, a technique from the UC San Diego (UCSD) research team Z-Lab that takes the MTP technology covered previously and pushes it in a new direction. Applying a diffusion model to a draft model was a genuinely ambitious idea, but in this test it wasn’t able to beat Gemma’s native draft model, Assistant.

Tracing the cause, the token acceptance rate was lower than the native draft model’s, and that penalty appears to have been the dominant factor. Without a way to understand how to raise that acceptance rate, it’s hard to say at this point whether further speed gains are realistic.

We also think the matchup itself didn’t help. Gemma-4’s Assistant model is built with practicality in mind and makes full use of Gemma-4’s shared KV-cache mechanism, whereas DFlash has to write to an independent KV cache — and that overhead seems to have tipped the balance toward lower throughput.

On the broader question of putting diffusion models to work in LLMs, LLaDa and Dream are the well-known names, but more recently Google itself has released a model called Diffusion Gemma. We’re currently in the middle of testing it ourselves, and the throughput numbers so far are notably high. We plan to cover those results in a dedicated article.

Supplementary Information

DFlash Model Structure and Connection to the Main Model
The DFlash model, which adds MTP capability to Gemma-4-12B-it, is architecturally distinct from Gemma-4-Assistant: rather than sharing a KV cache the way Assistant does, it’s a fully separate model. In terms of its activation function and related choices, it’s actually closer to a Qwen-style model. Its LA/GA notation follows Gemma-4’s own convention — LA meaning Sliding Window Attention and GA meaning Full Attention.

Unlike the Assistant model, this approach uses non-causal attention, and internally relies on Masked Diffusion — that’s the key structural difference. This mechanism lets DFlash derive all of its predicted tokens at once; from there, the rest of the pipeline follows the same logic as the Assistant model, outputting whichever tokens get accepted.

What Is EAGLE3?
EAGLE3 is version 3 of the EAGLE (Extrapolation Algorithm for Greater Language Model Efficiency) series. Both EAGLE and successor models like P-EAGLE treat the sequential bottleneck of causal attention as the core problem, and both are focused on how to parallelize that part of the pipeline further.

References

  1. DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence — arXiv
  2. Zhijian Liu / Z-Lab — z-lab.ai
  3. DFlash: Block Diffusion for Flash Speculative Decoding — arXiv
  4. Accelerating Large Language Model Decoding with Speculative Sampling — arXiv
  5. ELYZA-Diffusion-Instruct-1.0-Dream-7B — Hugging Face
C
Compute Cluster Research
RESEARCH #005 — September 11, 2026