Note from the future: see the recommended sampling settings in the post below - they make a difference. Initial issues appear to be resolving now. If you're using LM Studio, use the Unsloth versions, before the LM Studio Community version. Also be sure to try the Ridge quant by Empero. It runs on machines with as little as 12Gb VRAM.
Be sure to see all the info about the previous 3.6 versions of Qwen, plus Gemma 4: https://aibynick.com/thread/22
When you watch all the first review videos, everyone's response to Qwen 3.8 27b is as expected - it's being called the best model that can run on small consumer GPUs. And the benchmark hype all seems to show that it produces better output than 3.6, in most tests.
The only problem is that Qwen 3.8 27b thinks a lot. In fact, I'm experiencing significant issues with it terminating before completing tasks, because tasks are running so long. Perhaps another harness may help with that issue (the problematic terminations all occurred in Pi), but I've seen many reviews expressing a similar problem - even small software generations end up using nearly all the 256k context. Results are great, when the job is complete, but I'm regularly seeing jobs not completing.
To be clear, I'm not experiencing long context termination issues with only machines that have small GPUs. The task termination issue has occurred on both DGX Spark and Strix Halo machines, using both the q4 and q6 quants of 3.8 27b. So, this is not an isolated issue with just one version, one machine, or one configuration.
The other issue is that the 27b model is slow for a smallish model. The q4 quant on Strix Halo and the q6 quant on DGX Spark both run at just over 12 tokens per second. That's usable, but no where near the 50-60 tps I typically get with the Qwen 3.6 35a3 MOE model, on those same hardware platforms (and on faster hardware such as RTX 5090 and RTX 6000, the 3.6 MOE model runs over 200 tps).
I understand that historically, the dense 3.6 27b (dense) model is generally expected to produce higher quality output than the 3.6 35a3 (MOE) model, but for the sorts of tasks I've used Qwen 3.6 to accomplish, the MOE version has always done a great job, and at many times the speed. What that means in practice is that I can perform more iterations in less time, and craft better output overall, using the MOE model. Honestly, I haven't needed any better quality than I get from the MOE model. It's fantastic. When I need more world knowledge and capability, I use Deepseek V4 Flash, and for vision, Mimo 2.5. The Gemma 4 models, Hy3, and Stepfun 3.7 Flash also get thrown into the local AI mix. Orchestrating tasks with that stable of brains has been effective for such a wide variety of workflows.
One thing I think important to note is that Qwen 3.5 122b (MOE) performs more than twice as fast as 3.8 27b (dense) on both Strix Halo and DGX Spark machines. That bigger class of model has much more world knowledge than 27 billion parameters can ever contain. So, I'm very eager to see if Alibaba will release a 120ish billion parameter version of 3.8. I have a sense that a new 100b+ sized MOE model from Qwen could likely be a Deepseek V4 Flash killer.
So for now, I'm disappointed with Qwen 3.8 27b. I'll be extremely eager to try any mixture of experts Qwen 3.8 model. I expect a 35b or 122b size MOE model will likely be a serious improvement over other comparable options. I'll also try other harnesses to see if fewer task termination issues are experienced - and/or perhaps an update to Pi will help reduce the problems I'm currently seeing with 3.8 27b. Either way, I think we'll find that in practice, 3.8 27b is not the magic pill so many people hoped it would be. I'll keep my ears open, and continue testing...