Mimo 2.6 Flash may become my default local LLM on a cluster of 2 DGX Sparks

159 views
Nick Antonaccio
Nick AntonaccioAdmin
Sep 26, 2026 at 21:19 (edited, 1 revision)
#1

I'm currently running Mimo 2.6 Flash on 2 clustered DGX Spark machines, at native shipped precision, ~20 tokens per second.

This is exciting because nothing was re-quantized, so this model should be as precise as it is on the API. That hasn't been the case for either Deepseek V4 Flash or GLM 5.3 Flash (both those models need to run at significant compression to enable speed and large context), so this model could potentially beat the quality of both those previous leaders in my locally hosted LLM stable - at the same speed.

The current setup is 4‑bit MXFP4 experts + FP8 E4M3 attention/dense + BF16 o_proj, FP8 KV cache. vLLM detects it as quantization=fp8 and runs the experts through its DEEPGEMM_MXFP4 MoE backend. No dequantization/GGUF step was involved.

I'll provide more feedback and examples soon!

Nick Antonaccio
Nick AntonaccioAdmin
Sep 26, 2026 at 21:21
#2

One down side is that this model takes more than 10 minutes to re-load on a reboot, but I don't expect to have to reboot often.

Nick Antonaccio
Nick AntonaccioAdmin
Sep 27, 2026 at 00:38 (edited, 1 revision)
#3

I was surprised at how much trouble Mimo 2.6 Flash had with a basic Rubik's Cube solver. It's working, but the task took a while and required a few debug iterations:

https://com-pute.com/nick/rubiks--mimo26f.html

Even Qwen 3.6 35a3 knocked out that challenge first shot, without any issues or iterations required whatsoever.

We'll see how the rest of the testing goes...

Nick Antonaccio
Nick AntonaccioAdmin
Oct 01, 2026 at 01:05 (edited, 1 revision)
#4

I've been comparing Mimo 2.6 Flash against GLM 5.3 Flash all week. On the 2 DGX Spark clusters, Mimo is definitely a bit faster, and on some tasks it does a better job. For example, here's a little 1-off vibe coded 3D racing game:

https://com-pute.com/nick/3D_racing--mimo2.6f--turbo-circuit-3d.html

That's better than most of the frontier models could do at the beginning of the year.

For all the obscure knowledge questions I've tried with those 2 models, GLM3 Flash just seems to know more details, even at IQ3 quant - but for most situations, that simply doesn't matter much, because all these models are great at compiling research online. Even models from 6 months ago do great at looking up info, if there's an Internet connection.

So for the moment, I'm keeping 1 DGX cluster running GLM Flash and another running Mimo Flash, to see if there's any clear consensus. As it stands now, they both feel great.

BTW, I'm still eagerly awaiting version 4 of the current Qwen 3.8 Flash Next architecture to be released. It runs faster than both MiMo and GLM Flash, at 4 bit quant, on a single 128Gb Strix Halo, DGX Spark, Mac Studio, etc. Its file size and memory use is much smaller than Mimo or GLM Flash, but even the current 3.8 Flash Next preview model seems to have genuinely comparable world knowledge, coding skills, and general agentic capability. In a few cases, Qwen 3.8 Flash Next has downright beaten Mimo and GLM at challenging coding tasks. It's now the default model that I keep loaded on Strix Halo machines.

All these best of breed free local models are already world changing - it's so exciting to imagine what we'll be running on local hardware by this time next year.

Nick Antonaccio
Nick AntonaccioAdmin
Oct 08, 2026 at 23:58 (edited, 7 revisions)
#5

I've been keeping a copy of Mimo 2.6 Flash running on a cluster of 2 DGX Sparks. It's quicker than GLM 5.3 Flash, and around the same class of capability. I keep GLM 5.3 Flash running on a separate cluster of 2 DGX Sparks, and use them both interchangeably.

Please login to post a reply.

© 2026 AI By Nick.