Donato Capitella put together everything needed to run Qwen3.8-Flash-Next on AMD Strix Halo, with the best performance yet. He's using the Halogen engine to get 752 tokens/s prefill and 41 tokens/s generation at 32K context, and that configuration completed all 19 tasks on his benchmark tests (18 on the first attempt). That's the best combination of speed and performance yet for any model on his leaderboard, making that configuration of Qwen3.8-Flash-Next the current leading model all-around for Strix Halo:
https://www.youtube.com/watch?v=Nm_zN6RQ_eE
That's pretty freakin amazing results: a seriously capable model (beating Deepseek v4 Flash in Donato's benchmarks), running on a single very portable laptop machine (no clustered hardware), which you can currently purchase for $2850. That's hard to beat right now.
His new toolbox is available at:
https://github.com/kyuz0/ai-toolbox-cockpit
Please keep in mind that these successful results were all achieved with the preview model which shipped with Qwen3.8-Flash-Next. Qwen 4 models running on that same next architecture are expected to be released before the end of the year. That should be something genuinely special for the self-hosting crowd.