Qwen3.8-Flash-Next locally on a 12 GB RTX 4070 A community write-up... - Vibeus
Qwen3.8-Flash-Next locally on a 12 GB RTX 4070 A community write-up reports running Qwen3.8-Flash-Next (125B-A6B MoE with a 51B n-gram table) on an RTX 4070 with 12 GB of VRAM, 64 GB of DDR5-5600 memory and a Gen4 NVMe drive under Linux. The author says generation improved from about 6 to nearly 20 tokens per second after patches and optimizations, with 20.65 tokens per second reported for an MTP variant. The setup uses a 4.27 bpw AtomicChat quant, lazy n-gram SSD offloading and --fit on --fit-target 512. The post also reports a 64k context window and 300–350 tokens per second for prompt processing. The author says the 27B dense model may be more suitable for systems with 24 GB of VRAM, while this setup targets a low-VRAM, high-RAM configuration. Источник