Corsair AI Workstation 300 · MiniV Vulkan UMA

DFlash Pilot Benchmark

Run dflash-matrix-20260708_153700 · 2026-07-08 · public sanitized metrics
Decision

Do not promote DFlash on MiniV yet.

DFlash activated and accepted draft tokens, but the isolated Vulkan DFlash build was slower than baseline across all tested compatible models. Baseline and DFlash used the same build; only speculative mode changed.

Context8192
Output cap256 tokens
Drafter quantQ8_0
BackendRADV Vulkan
Buildllama.cpp e8f1015

Qwen3.6 / Qwopus3.6 27B Q6_K

Baseline decode9.48 tok/s
DFlash decode8.34 tok/s
DFlash delta-12.1%
Draft acceptance30.8%
Mean accepted length3.45

Qwen3.6 35B-A3B Q6_K_XL

Baseline decode51.07 tok/s
DFlash decode33.72 tok/s
DFlash delta-34.0%
Draft acceptance37.4%
Mean accepted length3.92

Gemma4 26B-A4B-it Q6_K_XL

Baseline decode44.83 tok/s
DFlash decode31.10 tok/s
DFlash delta-30.6%
Draft acceptance34.3%
Mean accepted length3.64

Per-case metrics

Prefill and decode are separated. Decode uses llama-server predicted timing; wall decode is shown separately. Status SPEED_VALID_EMPTY_CONTENT means 256 completion tokens were generated, but content was empty due thinking-template behavior in this isolated PR build; this report is runtime-only, not quality scoring.

ModelModeLoad secPrefill tok/sDecode tok/sWall tok/sDraft accept %Status
Qwen3.6 / Qwopus3.6 27B Q6_Kbaseline64.46545.2559.4819.021SPEED_VALID_EMPTY_CONTENT
Qwen3.6 / Qwopus3.6 27B Q6_Kdflash70.1641.8078.3377.94930.8SPEED_VALID_EMPTY_CONTENT
Qwen3.6 35B-A3B Q6_K_XLbaseline91.40277.25951.07243.802SPEED_VALID_EMPTY_CONTENT
Qwen3.6 35B-A3B Q6_K_XLdflash94.19183.79733.71730.60337.4SPEED_VALID_EMPTY_CONTENT
Gemma4 26B-A4B-it Q6_K_XLbaseline64.135176.63844.83141.136SPEED_VALID_EMPTY_CONTENT
Gemma4 26B-A4B-it Q6_K_XLdflash65.16159.64631.129.20734.3SPEED_VALID_EMPTY_CONTENT

Methodology guardrails