Engineering 21 June 2026
Our Local AI Coding Stack: Qwen3.6-35B on vLLM — the Real Numbers
How we run Qwen3.6-35B as a coding assistant on our own hardware — vLLM, FP8, 256K context, speculative decoding. With real measurements from live operation.
Read more →