MiniV Best-Real TurboQuant + MTP Validation

2026-07-08 · Bench v2.2 · separated prefill / decode / MTP metrics
Proven optimizations only
Best production path remains TurboQuant + MTP.

This validation uses the already-proven MiniV path: TurboQuant turbo4/turbo3, MTP speculative decoding, ubatch=1024, flash attention, performance governor, swappiness=10, THP=always, adaptive timeouts, and deterministic task checkers. No DFlash data is mixed into this benchmark.

JSON datasetResults CSVBack to gallery

Winner signal

architect-35b-q6 led short/medium agent tasks at 76.3 decode tok/s average and passed deep-context retrieval at 49.4 decode tok/s with 91.9% MTP acceptance. The 4-task full average is 69.6 decode tok/s.

8/8Tasks passed
80.5Best per-task decode tok/s
91.3%Architect avg MTP acceptance
360.2Architect deep prefill tok/s

Aggregate results

ScopeModelPassScorePrefillDecodeMTPElapsed
3-task short/mediumarchitect-35b-q63/31.00286.376.391.0%12.6
3-task short/mediumqwen36-35b-q63/31.00304.268.190.6%29.0
4-task fullarchitect-35b-q64/41.00304.869.591.2%32.6
4-task fullqwen36-35b-q64/41.00321.062.590.2%45.8

Per-task results

ModelTaskStatusScorePrefillDecodeMTPTotal s
architect-35b-q6tool_callPASS1.00210.080.590.0%3.0
architect-35b-q6coding_fixPASS1.00235.575.392.7%8.0
architect-35b-q6context_heavyPASS1.00413.473.090.4%26.8
qwen36-35b-q6tool_callPASS1.00214.472.892.8%3.8
qwen36-35b-q6coding_fixPASS1.00250.866.990.4%20.6
qwen36-35b-q6context_heavyPASS1.00447.564.588.5%62.5
architect-35b-q6deep_context_retrievalPASS1.00360.249.491.9%92.8
qwen36-35b-q6deep_context_retrievalPASS1.00371.145.789.3%96.1