CORSAIR_AI_WS_300 :: MINIV_VULKAN

MiniV Vulkan UMA · Corsair AI Workstation 300 64GB · all-local trust gates · Updated July 16, 2026 · current operation + Bench v2.2
📍 ABOUT THIS GALLERY Start here
Can an entire AI lab run on one compact workstation — no cloud, no subscriptions, no external GPUs?

This gallery proves it can. Every benchmark, every model, and every 3D scene was generated entirely on-premise by a single 64GB machine the size of a game console. The question isn't whether local AI is viable — it's how fast and how reliable it can be when you tune it properly.

For the tech-curious: We tested 21 AI models for speed, quality, and deep-context retrieval. We generated 49 working 3D/WebGL apps from scratch. We pushed context windows to 120,000 tokens. All on integrated graphics.

🖥️ The Machine Corsair AI Workstation 300 — a mini PC with AMD Ryzen AI MAX 385 and 64GB unified memory. Think 48GB VRAM without a dedicated GPU.
🔒 100% Private No data leaves the machine. No API keys, no per-token billing. Your prompts, your code, your IP — all stay local.
⚡ Fast Enough 71 tok/s on the champion model — ~3,500 words/min. MTP acceptance hits 94%. 120K token retrieval in 92 seconds.
📊 What's Inside 21 models benchmarked · 84 test runs · 49 interactive 3D apps · 120K token context ladder · DFlash speculative decoding pilot.
📑 REPORT INDEX Quick navigation
🛡️ CURRENT LOCAL OPERATION public-safe runtime policy · 2026-07-16
Open WebUI via model gateThe local general-chat UI now reaches inference through an allowlist gate, not through a direct multi-model router connection.
Single-resident guardrailOnly architect-35b-q6 is exposed to the UI while the router keeps one resident model. This prevents destructive model-load thrashing.
Local-first memoryThe local librarian is operating read-only with governed proposal/apply/rollback. Cloud review is optional, never the memory dependency.
Public boundaryThis gallery documents evidence and routing decisions only. Access URLs, credentials, private traces and internal filesystem details are intentionally excluded.
🚀 BEST-REAL TURBOQUANT + MTP VALIDATION July 08, 2026 · 8/8 PASS
Proven optimizations only · no DFlash mixed in
Architect-35B Q6 remains the strongest current MiniV production lane.

Top 35B Q6 MoE/MTP lanes were re-tested with the proven stack: TurboQuant turbo4/turbo3, MTP speculative decoding, ubatch=1024, flash attention, performance governor, swappiness=10, THP=always, adaptive timeouts, and separated llama-server timings.

PASS rate8/8 deterministic tasks passed across architect-35b-q6 and qwen36-35b-q6.
Architect short/medium76.3 decode tok/s average · 91.0% MTP acceptance · 3/3 PASS.
Architect deep context360.2 prefill tok/s · 49.4 decode tok/s · 91.9% MTP · PASS.
🧭 CORRECTED 120K CONTEXT RETRIEVAL LADDER July 10, 2026 · 86 rows · 21 aliases
Token-accounted retrieval envelope · public-safe evidence
Architect-35B Q6 retrieves exactly through the 120K request class.

The corrected ladder records authoritative server-reported input tokens instead of relying on client estimates. architect-35b-q6 passed every class through 120K, reaching 117,814 observed prompt tokens. The result reinforces the current TurboQuant + MTP production profile; it does not justify a global Q8 promotion.

Visible retrieval PASS82 exact visible-answer passes. PARTIAL and non-pass outcomes remain explicit rather than being hidden inside a single score.
Stable default envelope131K configured context · 120K request class · 117.8K observed prompt tokens on architect’s largest visible PASS.
Method disciplinePrefill and decode stay separate. This ladder is a retrieval probe—not a generic quality or speed leaderboard.
DFLASH PILOT UPDATE July 08, 2026 July 08, 2026
Runtime pilot · measured negative
DFlash activated, but was slower than baseline on MiniV Vulkan.

Q8 DFlash drafters were tested against Qwen3.6-27B, Qwen3.6-35B-A3B, and Gemma4-26B-A4B using a separate llama.cpp DFlash Vulkan build. Prefill, decode, wall decode, and draft acceptance remain separated. Decision: do not promote DFlash to production MiniV yet.

Best baseline testedQwen3.6 35B-A3B Q6_K_XL: 51.07 decode tok/s
Best DFlash testedQwen3.6 35B-A3B DFlash: 33.72 decode tok/s
Acceptance observed30.8%–37.4% draft-token acceptance, still not enough to beat baseline.
🎮 VISUAL 3D BENCHMARK GALLERY 11 models · 49 HTML deliverables · 0 runtime issues
Direct local-model generation · 5 visual prompts · pass4 runtime verified
11 local models generated 49 working 3D/WebGL/Canvas HTML deliverables — zero strict runtime issues.

Each model received 5 visual prompts: Three.js Particle Galaxy, FPS Raycasting Engine, 3D Flight Simulation, Wave Ocean Shader, and Breakout Canvas Game. All deliverables were runtime-gated with headless browser verification. SHA256 integrity manifest included.

Models testedarchitect-35b-q6, coding-coder-27b-q6/q8, nex-n2-mini-q6, ornith-35b-q6, qwen36-35b-q6/q8, reasoning-27b-q6, q8-architect-35b, q8-reasoning-27b
Runtime gate49/49 complete · 0 markdown fences · 0 strict browser/runtime issues · 6 benign warnings only
Pi-dev harness15 additional pi-agent-generated files verified across architect, ornith, and qwen lanes
🧠 ORCHESTRATION CREDIT N30
Benchmark methodology owner
N30
All benchmark orchestration, methodology design, and AI model evaluations directed by N30. Bench v2.2 methodology: separated prefill/decode metrics, MTP acceptance tracking, and deep context retrieval scoring. Strict trust remains separated from draft/recovery lanes.
📊 FINAL TRUST-GATE METRICS MiniV Vulkan UMA · Bench v2.2
21
Models Evaluated
84
Total Benchmark Runs
11
Perfect Score (1.00)
71
Best Decode tok/s
94%
Best MTP Accept
92s
Best Deep Context
EXECUTIVE DECISION LAYER MiniV default lane
What changed
Bench v2.2 delivers separated prefill/decode metrics, MTP acceptance tracking, and deep context retrieval scoring across 21 models.

21 models benchmarked with Bench v2.2 — separated prefill/decode/MTP metrics. 11 models achieved perfect score. Deep context retrieval solved: 310s→92s via ubatch optimization. OS-level optimizations: CPU governor performance, ubatch-size tuning, and Transparent Huge Pages (THP) enabled.

Championminiv-qwopus36-35b-a3b-q6-mtp is the local champion for MiniV Vulkan UMA — 71 tok/s decode, 90% MTP accept, 92s deep context.
DisciplineFramework-ready is useful, but only LOCAL_CHAMPION with perfect score and deep context pass moves into production routing.
EvidenceBench v2.2 separates prefill/decode from llama-server API timings object, tracks MTP acceptance, and scores deep context retrieval with adaptive timeouts.
🧪 N30 TRUST PIPELINE Bench v2.2 methodology
Smoke gate

Agentic/structured tasks with adaptive timeouts by context size. Bench v2.2 separates prefill/decode metrics from llama-server API timings object. Thinking tokens isolated via reasoning_content field.

Framework gate

Framework-spec tasks: executable contracts, BMAD/story breakdown, spec review, and execution planning. MTP acceptance rate tracked for multi-token prediction models.

Deep context gate

Long-context retrieval scoring (deep_context_score) with timing measurement (deep_context_time_s). OS-level optimizations: CPU governor performance, ubatch-size tuning, THP enabled.

🏆 QUALITY LEADERBOARD 21 models · Bench v2.2
📋 MINIV VULKAN UMA — FINAL ALL-LOCAL COMPARISON Bench v2.2 final run
Swipe horizontally to inspect all gate columns. Prefill/decode separated, MTP and deep context shown.
#ModelClassPassDecodePrefillMTPDeep ctxStatus
1 miniv-qwopus36-35b-a3b-q6-mtp LOCAL_CHAMPION 4/4 71.2 320.3 90% 92s · 1.00 MODEL_DONE
2 miniv-qwen36-35b-a3b-q6 FRAMEWORK_READY 4/4 63.8 332.7 91% 99s · 1.00 MODEL_DONE
3 miniv-qwopus36-35b-a3b-q8-mtp FRAMEWORK_READY 4/4 62.5 361.8 92% 93s · 1.00 MODEL_DONE
4 miniv-qwen36-35b-a3b-q8 FRAMEWORK_READY 4/4 54.4 338.6 89% 105s · 1.00 MODEL_DONE
5 miniv-ornith1-35b-q6 FRAMEWORK_READY 4/4 52.8 328.2 97s · 1.00 MODEL_DONE
6 miniv-ornith1-35b-q6-100k FRAMEWORK_READY 4/4 52.6 328.1 97s · 1.00 MODEL_DONE
7 miniv-nex-n2-mini-q6 FRAMEWORK_READY 4/4 49.5 336.1 88s · 1.00 MODEL_DONE
8 miniv-qwen3-coder-next-q4 SMOKE_PROMOTED 3/4 50.0 228.7 106s · 1.00 MODEL_DONE
9 miniv-qwopus35-9b-coder-q6-mtp SMOKE_PROMOTED 3/4 48.1 439.7 89% 110s · 1.00 MODEL_DONE
10 miniv-agentworld-35b-a3b-q6 FRAMEWORK_READY 4/4 46.5 343.6 120s · 0.67 MODEL_DONE
11 miniv-ornith1-35b-q8 SMOKE_PROMOTED 3/4 45.3 372.5 103s · 1.00 MODEL_DONE
12 miniv-gemma4-26b-a4b-q6 FRAMEWORK_READY 4/4 39.7 404.8 84s · 0.67 MODEL_DONE
13 miniv-gemma4-26b-a4b-q8 FRAMEWORK_READY 4/4 36.7 421.9 83s · 0.67 MODEL_DONE
14 miniv-agentworld-35b-a3b-q8 FRAMEWORK_READY 4/4 32.6 329.3 114s · 0.67 MODEL_DONE
15 miniv-ornith1-9b-q8 SMOKE_PROMOTED 3/4 22.6 489.7 94s · 0.67 MODEL_DONE
16 miniv-qwopus36-27b-v2-q6-mtp FRAMEWORK_READY 4/4 19.1 127.2 94% 301s · 1.00 MODEL_DONE
17 miniv-qwopus36-27b-coder-q6-mtp FRAMEWORK_READY 4/4 19.0 128.6 93% 310s · 1.00 MODEL_DONE
18 miniv-gemma4-12b-q6 FRAMEWORK_READY 4/4 16.6 279.5 166s · 0.67 MODEL_DONE
19 miniv-qwopus36-27b-coder-q8-mtp FRAMEWORK_READY 4/4 14.7 135.7 93% 293s · 1.00 MODEL_DONE
20 miniv-gemma4-12b-q8 FRAMEWORK_READY 4/4 13.2 286.2 159s · 0.67 MODEL_DONE
21 miniv-qwopus36-27b-v2-q8 FRAMEWORK_READY 4/4 7.4 139.0 300s · 1.00 MODEL_DONE
🧬 JACKRONG QWOPUS36-EVAL STYLE COMPARISON framing, not apples-to-apples
Public reference frame

Jackrong/Kyle Hessling Space: Qwopus3.6-27B v1-preview Q4_K_M on RTX 5090, 16-prompt suite: 5 agentic, 5 web-design, 6 canvas/WebGL. Published copy reports 62.3 tok/s average, 87.4k generated tokens, 23.4 min runtime.

We reuse the comparison vocabulary — agentic reasoning, production UI, and canvas/WebGL — while showing N30's stricter local multi-pass trust gates with Bench v2.2 metrics.

Local gate mapping
  • Smoke → agentic trust probes with adaptive timeouts by context size.
  • Framework → spec/framework tasks with MTP acceptance rate tracking.
  • Deep context → long-context retrieval scoring and timing.
  • Policy → prefill/decode separated from llama-server API timings object; thinking tokens isolated via reasoning_content.
🔗 PROVENANCE & OFFICIAL LINKS public evidence
Source run: Bench v2.2 · 21 models · 84 runs · 2026-07-07 · Methodology: separated prefill/decode metrics, MTP acceptance tracking, deep context retrieval scoring, thinking tokens via reasoning_content, adaptive timeouts by context size.