Qwen3.6 / Qwopus3.6 27B Q6_K
Baseline decode9.48 tok/s
DFlash decode8.34 tok/s
DFlash delta-12.1%
Draft acceptance30.8%
Mean accepted length3.45
Do not promote DFlash on MiniV yet.
DFlash activated and accepted draft tokens, but the isolated Vulkan DFlash build was slower than baseline across all tested compatible models. Baseline and DFlash used the same build; only speculative mode changed.
Prefill and decode are separated. Decode uses llama-server predicted timing; wall decode is shown separately. Status SPEED_VALID_EMPTY_CONTENT means 256 completion tokens were generated, but content was empty due thinking-template behavior in this isolated PR build; this report is runtime-only, not quality scoring.
| Model | Mode | Load sec | Prefill tok/s | Decode tok/s | Wall tok/s | Draft accept % | Status |
|---|---|---|---|---|---|---|---|
| Qwen3.6 / Qwopus3.6 27B Q6_K | baseline | 64.465 | 45.255 | 9.481 | 9.021 | — | SPEED_VALID_EMPTY_CONTENT |
| Qwen3.6 / Qwopus3.6 27B Q6_K | dflash | 70.16 | 41.807 | 8.337 | 7.949 | 30.8 | SPEED_VALID_EMPTY_CONTENT |
| Qwen3.6 35B-A3B Q6_K_XL | baseline | 91.402 | 77.259 | 51.072 | 43.802 | — | SPEED_VALID_EMPTY_CONTENT |
| Qwen3.6 35B-A3B Q6_K_XL | dflash | 94.191 | 83.797 | 33.717 | 30.603 | 37.4 | SPEED_VALID_EMPTY_CONTENT |
| Gemma4 26B-A4B-it Q6_K_XL | baseline | 64.135 | 176.638 | 44.831 | 41.136 | — | SPEED_VALID_EMPTY_CONTENT |
| Gemma4 26B-A4B-it Q6_K_XL | dflash | 65.16 | 159.646 | 31.1 | 29.207 | 34.3 | SPEED_VALID_EMPTY_CONTENT |