Post History

Current version by Nick Antonaccio

Current VersionSep 13, 2026 at 12:07

Donato Capitella put together everything needed to run Qwen3.8-Flash-Next on AMD Strix Halo, with the best performance yet. He's using the Halogen engine to get 752 tokens/s prefill and 41 tokens/s generation at 32K context, and that configuration completed all 19 tasks on his benchmark tests (18 on the first attempt). That's the best combination of speed and performance yet for any model on his leaderboard, making that configuration of Qwen3.8-Flash-Next the current leading model all-around for Strix Halo:

https://www.youtube.com/watch?v=Nm_zN6RQ_eE

That's pretty freakin amazing results: a seriously capable model (beating Deepseek v4 Flash in Donato's benchmarks), running on a single very portable laptop machine (no clustered hardware), which you can currently purchase for $2850. That's hard to beat right now.

His new toolbox is available at:

https://github.com/kyuz0/ai-toolbox-cockpit

Please keep in mind that these successful results were all achieved with the preview model which shipped with Qwen3.8-Flash-Next. Qwen 4 models running on that same next architecture are expected to be released before the end of the year. That should be something genuinely special for the self-hosting crowd.

Previous Versions
Version 1Sep 13, 2026 at 12:07

Donato Capitella put together everything needed to run Qwen3.8-Flash-Next on AMD Strix Halo, with the best performance yet. He's using the Halogen engine to get 752 tokens/s prefill and 41 tokens/s generation at 32K context, and that configuration completed all 19 tasks on his benchmark tests (18 on the first attempt). That's the best combination of speed and performance yet for any model on his leaderboard, making that configuration of Qwen3.8-Flash-Next the current leading model all-around for Strix Halo:

https://www.youtube.com/watch?v=Nm_zN6RQ_eE

His new toolbox is available at:

https://github.com/kyuz0/ai-toolbox-cockpit

Please keep in mind these successful results are all with the preview model that shipped with Qwen3.8-Flash-Next - Qwen 4 models running in that same next architecture are expected before the end of the year. That should really be something special.