Post History

Current version by Nick Antonaccio

Current VersionAug 27, 2026 at 23:00

Note from the future: see the recommended sampling settings in the post below - they make a difference. Initial issues appear to be resolving now. If you're using LM Studio, use the Unsloth versions, before the LM Studio Community version. Also be sure to try the Ridge quant by Empero. It runs on machines with as little as 12Gb VRAM.

Be sure to see all the info about the previous 3.6 versions of Qwen, plus Gemma 4: https://aibynick.com/thread/22

When you watch all the first review videos, everyone's response to Qwen 3.8 27b is as expected - it's being called the best model that can run on small consumer GPUs. And the benchmark hype all seems to show that it produces better output than 3.6, in most tests.

The only problem is that Qwen 3.8 27b thinks a lot. In fact, I'm experiencing significant issues with it terminating before completing tasks, because tasks are running so long. Perhaps another harness may help with that issue (the problematic terminations all occurred in Pi), but I've seen many reviews expressing a similar problem - even small software generations end up using nearly all the 256k context. Results are great, when the job is complete, but I'm regularly seeing jobs not completing.

To be clear, I'm not experiencing long context termination issues with only machines that have small GPUs. The task termination issue has occurred on both DGX Spark and Strix Halo machines, using both the q4 and q6 quants of 3.8 27b. So, this is not an isolated issue with just one version, one machine, or one configuration.

The other issue is that the 27b model is slow for a smallish model. The q4 quant on Strix Halo and the q6 quant on DGX Spark both run at just over 12 tokens per second. That's usable, but no where near the 50-60 tps I typically get with the Qwen 3.6 35a3 MOE model, on those same hardware platforms (and on faster hardware such as RTX 5090 and RTX 6000, the 3.6 MOE model runs over 200 tps).

I understand that historically, the dense 3.6 27b (dense) model is generally expected to produce higher quality output than the 3.6 35a3 (MOE) model, but for the sorts of tasks I've used Qwen 3.6 to accomplish, the MOE version has always done a great job, and at many times the speed. What that means in practice is that I can perform more iterations in less time, and craft better output overall, using the MOE model. Honestly, I haven't needed any better quality than I get from the MOE model. It's fantastic. When I need more world knowledge and capability, I use Deepseek V4 Flash, and for vision, Mimo 2.5. The Gemma 4 models, Hy3, and Stepfun 3.7 Flash also get thrown into the local AI mix. Orchestrating tasks with that stable of brains has been effective for such a wide variety of workflows.

One thing I think important to note is that Qwen 3.5 122b (MOE) performs more than twice as fast as 3.8 27b (dense) on both Strix Halo and DGX Spark machines. That bigger class of model has much more world knowledge than 27 billion parameters can ever contain. So, I'm very eager to see if Alibaba will release a 120ish billion parameter version of 3.8. I have a sense that a new 100b+ sized MOE model from Qwen could likely be a Deepseek V4 Flash killer.

So for now, I'm disappointed with Qwen 3.8 27b. I'll be extremely eager to try any mixture of experts Qwen 3.8 model. I expect a 35b or 122b size MOE model will likely be a serious improvement over other comparable options. I'll also try other harnesses to see if fewer task termination issues are experienced - and/or perhaps an update to Pi will help reduce the problems I'm currently seeing with 3.8 27b. Either way, I think we'll find that in practice, 3.8 27b is not the magic pill so many people hoped it would be. I'll keep my ears open, and continue testing...

Previous Versions
Version 9Aug 27, 2026 at 23:00

Note from the future: see the recommended sampling settings in the post below - they make a difference. Initial issues appear to be resolving now. If you're using LM Studio, use the Unsloth versions, before the LM Studio Community version. Also be sure to try the Ridge quant by Empero. It runs on machines with as little as 12Gb VRAM.

By sure to see all the info about the previous 3.6 versions of Qwen, plus Gemma 4: https://aibynick.com/thread/22

When you watch all the first review videos, everyone's response to Qwen 3.8 27b is as expected - it's being called the best model that can run on small consumer GPUs. And the benchmark hype all seems to show that it produces better output than 3.6, in most tests.

The only problem is that Qwen 3.8 27b thinks a lot. In fact, I'm experiencing significant issues with it terminating before completing tasks, because tasks are running so long. Perhaps another harness may help with that issue (the problematic terminations all occurred in Pi), but I've seen many reviews expressing a similar problem - even small software generations end up using nearly all the 256k context. Results are great, when the job is complete, but I'm regularly seeing jobs not completing.

To be clear, I'm not experiencing long context termination issues with only machines that have small GPUs. The task termination issue has occurred on both DGX Spark and Strix Halo machines, using both the q4 and q6 quants of 3.8 27b. So, this is not an isolated issue with just one version, one machine, or one configuration.

The other issue is that the 27b model is slow for a smallish model. The q4 quant on Strix Halo and the q6 quant on DGX Spark both run at just over 12 tokens per second. That's usable, but no where near the 50-60 tps I typically get with the Qwen 3.6 35a3 MOE model, on those same hardware platforms (and on faster hardware such as RTX 5090 and RTX 6000, the 3.6 MOE model runs over 200 tps).

I understand that historically, the dense 3.6 27b (dense) model is generally expected to produce higher quality output than the 3.6 35a3 (MOE) model, but for the sorts of tasks I've used Qwen 3.6 to accomplish, the MOE version has always done a great job, and at many times the speed. What that means in practice is that I can perform more iterations in less time, and craft better output overall, using the MOE model. Honestly, I haven't needed any better quality than I get from the MOE model. It's fantastic. When I need more world knowledge and capability, I use Deepseek V4 Flash, and for vision, Mimo 2.5. The Gemma 4 models, Hy3, and Stepfun 3.7 Flash also get thrown into the local AI mix. Orchestrating tasks with that stable of brains has been effective for such a wide variety of workflows.

One thing I think important to note is that Qwen 3.5 122b (MOE) performs more than twice as fast as 3.8 27b (dense) on both Strix Halo and DGX Spark machines. That bigger class of model has much more world knowledge than 27 billion parameters can ever contain. So, I'm very eager to see if Alibaba will release a 120ish billion parameter version of 3.8. I have a sense that a new 100b+ sized MOE model from Qwen could likely be a Deepseek V4 Flash killer.

So for now, I'm disappointed with Qwen 3.8 27b. I'll be extremely eager to try any mixture of experts Qwen 3.8 model. I expect a 35b or 122b size MOE model will likely be a serious improvement over other comparable options. I'll also try other harnesses to see if fewer task termination issues are experienced - and/or perhaps an update to Pi will help reduce the problems I'm currently seeing with 3.8 27b. Either way, I think we'll find that in practice, 3.8 27b is not the magic pill so many people hoped it would be. I'll keep my ears open, and continue testing...

Version 8Aug 27, 2026 at 23:00

Note from the future: see the recommended sampling settings in the post below - they make a difference. Initial issues appear to be resolving now. If you're using LM Studio, use the Unsloth versions, before the LM Studio Community version. Also be sure to try the Ridge quant by Empero. It runs on machines with as little as 12Gb VRAM.

When you watch all the first review videos, everyone's response to Qwen 3.8 27b is as expected - it's being called the best model that can run on small consumer GPUs. And the benchmark hype all seems to show that it produces better output than 3.6, in most tests.

The only problem is that Qwen 3.8 27b thinks a lot. In fact, I'm experiencing significant issues with it terminating before completing tasks, because tasks are running so long. Perhaps another harness may help with that issue (the problematic terminations all occurred in Pi), but I've seen many reviews expressing a similar problem - even small software generations end up using nearly all the 256k context. Results are great, when the job is complete, but I'm regularly seeing jobs not completing.

To be clear, I'm not experiencing long context termination issues with only machines that have small GPUs. The task termination issue has occurred on both DGX Spark and Strix Halo machines, using both the q4 and q6 quants of 3.8 27b. So, this is not an isolated issue with just one version, one machine, or one configuration.

The other issue is that the 27b model is slow for a smallish model. The q4 quant on Strix Halo and the q6 quant on DGX Spark both run at just over 12 tokens per second. That's usable, but no where near the 50-60 tps I typically get with the Qwen 3.6 35a3 MOE model, on those same hardware platforms (and on faster hardware such as RTX 5090 and RTX 6000, the 3.6 MOE model runs over 200 tps).

I understand that historically, the dense 3.6 27b (dense) model is generally expected to produce higher quality output than the 3.6 35a3 (MOE) model, but for the sorts of tasks I've used Qwen 3.6 to accomplish, the MOE version has always done a great job, and at many times the speed. What that means in practice is that I can perform more iterations in less time, and craft better output overall, using the MOE model. Honestly, I haven't needed any better quality than I get from the MOE model. It's fantastic. When I need more world knowledge and capability, I use Deepseek V4 Flash, and for vision, Mimo 2.5. The Gemma 4 models, Hy3, and Stepfun 3.7 Flash also get thrown into the local AI mix. Orchestrating tasks with that stable of brains has been effective for such a wide variety of workflows.

One thing I think important to note is that Qwen 3.5 122b (MOE) performs more than twice as fast as 3.8 27b (dense) on both Strix Halo and DGX Spark machines. That bigger class of model has much more world knowledge than 27 billion parameters can ever contain. So, I'm very eager to see if Alibaba will release a 120ish billion parameter version of 3.8. I have a sense that a new 100b+ sized MOE model from Qwen could likely be a Deepseek V4 Flash killer.

So for now, I'm disappointed with Qwen 3.8 27b. I'll be extremely eager to try any mixture of experts Qwen 3.8 model. I expect a 35b or 122b size MOE model will likely be a serious improvement over other comparable options. I'll also try other harnesses to see if fewer task termination issues are experienced - and/or perhaps an update to Pi will help reduce the problems I'm currently seeing with 3.8 27b. Either way, I think we'll find that in practice, 3.8 27b is not the magic pill so many people hoped it would be. I'll keep my ears open, and continue testing...

Version 7Aug 19, 2026 at 12:05

Note from the future: see the recommended sampling settings in the post below - they make a difference. Initial issues appear to be resolving now. Also be sure to try the Ridge quant by Empero. It runs on machines with as little as 12Gb VRAM.

When you watch all the first review videos, everyone's response to Qwen 3.8 27b is as expected - it's being called the best model that can run on small consumer GPUs. And the benchmark hype all seems to show that it produces better output than 3.6, in most tests.

The only problem is that Qwen 3.8 27b thinks a lot. In fact, I'm experiencing significant issues with it terminating before completing tasks, because tasks are running so long. Perhaps another harness may help with that issue (the problematic terminations all occurred in Pi), but I've seen many reviews expressing a similar problem - even small software generations end up using nearly all the 256k context. Results are great, when the job is complete, but I'm regularly seeing jobs not completing.

To be clear, I'm not experiencing long context termination issues with only machines that have small GPUs. The task termination issue has occurred on both DGX Spark and Strix Halo machines, using both the q4 and q6 quants of 3.8 27b. So, this is not an isolated issue with just one version, one machine, or one configuration.

The other issue is that the 27b model is slow for a smallish model. The q4 quant on Strix Halo and the q6 quant on DGX Spark both run at just over 12 tokens per second. That's usable, but no where near the 50-60 tps I typically get with the Qwen 3.6 35a3 MOE model, on those same hardware platforms (and on faster hardware such as RTX 5090 and RTX 6000, the 3.6 MOE model runs over 200 tps).

I understand that historically, the dense 3.6 27b (dense) model is generally expected to produce higher quality output than the 3.6 35a3 (MOE) model, but for the sorts of tasks I've used Qwen 3.6 to accomplish, the MOE version has always done a great job, and at many times the speed. What that means in practice is that I can perform more iterations in less time, and craft better output overall, using the MOE model. Honestly, I haven't needed any better quality than I get from the MOE model. It's fantastic. When I need more world knowledge and capability, I use Deepseek V4 Flash, and for vision, Mimo 2.5. The Gemma 4 models, Hy3, and Stepfun 3.7 Flash also get thrown into the local AI mix. Orchestrating tasks with that stable of brains has been effective for such a wide variety of workflows.

One thing I think important to note is that Qwen 3.5 122b (MOE) performs more than twice as fast as 3.8 27b (dense) on both Strix Halo and DGX Spark machines. That bigger class of model has much more world knowledge than 27 billion parameters can ever contain. So, I'm very eager to see if Alibaba will release a 120ish billion parameter version of 3.8. I have a sense that a new 100b+ sized MOE model from Qwen could likely be a Deepseek V4 Flash killer.

So for now, I'm disappointed with Qwen 3.8 27b. I'll be extremely eager to try any mixture of experts Qwen 3.8 model. I expect a 35b or 122b size MOE model will likely be a serious improvement over other comparable options. I'll also try other harnesses to see if fewer task termination issues are experienced - and/or perhaps an update to Pi will help reduce the problems I'm currently seeing with 3.8 27b. Either way, I think we'll find that in practice, 3.8 27b is not the magic pill so many people hoped it would be. I'll keep my ears open, and continue testing...

Version 6Aug 18, 2026 at 13:53

Note from the future: see the recommended sampling settings in the post below - they make a difference. Initial issues appear to be resolving now. Also be sure to try the Ridge quant by Empero. It runs faster on every architecture.

When you watch all the first review videos, everyone's response to Qwen 3.8 27b is as expected - it's being called the best model that can run on small consumer GPUs. And the benchmark hype all seems to show that it produces better output than 3.6, in most tests.

The only problem is that Qwen 3.8 27b thinks a lot. In fact, I'm experiencing significant issues with it terminating before completing tasks, because tasks are running so long. Perhaps another harness may help with that issue (the problematic terminations all occurred in Pi), but I've seen many reviews expressing a similar problem - even small software generations end up using nearly all the 256k context. Results are great, when the job is complete, but I'm regularly seeing jobs not completing.

To be clear, I'm not experiencing long context termination issues with only machines that have small GPUs. The task termination issue has occurred on both DGX Spark and Strix Halo machines, using both the q4 and q6 quants of 3.8 27b. So, this is not an isolated issue with just one version, one machine, or one configuration.

The other issue is that the 27b model is slow for a smallish model. The q4 quant on Strix Halo and the q6 quant on DGX Spark both run at just over 12 tokens per second. That's usable, but no where near the 50-60 tps I typically get with the Qwen 3.6 35a3 MOE model, on those same hardware platforms (and on faster hardware such as RTX 5090 and RTX 6000, the 3.6 MOE model runs over 200 tps).

I understand that historically, the dense 3.6 27b (dense) model is generally expected to produce higher quality output than the 3.6 35a3 (MOE) model, but for the sorts of tasks I've used Qwen 3.6 to accomplish, the MOE version has always done a great job, and at many times the speed. What that means in practice is that I can perform more iterations in less time, and craft better output overall, using the MOE model. Honestly, I haven't needed any better quality than I get from the MOE model. It's fantastic. When I need more world knowledge and capability, I use Deepseek V4 Flash, and for vision, Mimo 2.5. The Gemma 4 models, Hy3, and Stepfun 3.7 Flash also get thrown into the local AI mix. Orchestrating tasks with that stable of brains has been effective for such a wide variety of workflows.

One thing I think important to note is that Qwen 3.5 122b (MOE) performs more than twice as fast as 3.8 27b (dense) on both Strix Halo and DGX Spark machines. That bigger class of model has much more world knowledge than 27 billion parameters can ever contain. So, I'm very eager to see if Alibaba will release a 120ish billion parameter version of 3.8. I have a sense that a new 100b+ sized MOE model from Qwen could likely be a Deepseek V4 Flash killer.

So for now, I'm disappointed with Qwen 3.8 27b. I'll be extremely eager to try any mixture of experts Qwen 3.8 model. I expect a 35b or 122b size MOE model will likely be a serious improvement over other comparable options. I'll also try other harnesses to see if fewer task termination issues are experienced - and/or perhaps an update to Pi will help reduce the problems I'm currently seeing with 3.8 27b. Either way, I think we'll find that in practice, 3.8 27b is not the magic pill so many people hoped it would be. I'll keep my ears open, and continue testing...

Version 5Aug 18, 2026 at 13:19

Note from the future: see the recommended sampling setting the post below - they make a difference

When you watch all the first review videos, everyone's response to Qwen 3.8 27b is as expected - it's being called the best model that can run on small consumer GPUs. And the benchmark hype all seems to show that it produces better output than 3.6, in most tests.

The only problem is that Qwen 3.8 27b thinks a lot. In fact, I'm experiencing significant issues with it terminating before completing tasks, because tasks are running so long. Perhaps another harness may help with that issue (the problematic terminations all occurred in Pi), but I've seen many reviews expressing a similar problem - even small software generations end up using nearly all the 256k context. Results are great, when the job is complete, but I'm regularly seeing jobs not completing.

To be clear, I'm not experiencing long context termination issues with only machines that have small GPUs. The task termination issue has occurred on both DGX Spark and Strix Halo machines, using both the q4 and q6 quants of 3.8 27b. So, this is not an isolated issue with just one version, one machine, or one configuration.

The other issue is that the 27b model is slow for a smallish model. The q4 quant on Strix Halo and the q6 quant on DGX Spark both run at just over 12 tokens per second. That's usable, but no where near the 50-60 tps I typically get with the Qwen 3.6 35a3 MOE model, on those same hardware platforms (and on faster hardware such as RTX 5090 and RTX 6000, the 3.6 MOE model runs over 200 tps).

I understand that historically, the dense 3.6 27b (dense) model is generally expected to produce higher quality output than the 3.6 35a3 (MOE) model, but for the sorts of tasks I've used Qwen 3.6 to accomplish, the MOE version has always done a great job, and at many times the speed. What that means in practice is that I can perform more iterations in less time, and craft better output overall, using the MOE model. Honestly, I haven't needed any better quality than I get from the MOE model. It's fantastic. When I need more world knowledge and capability, I use Deepseek V4 Flash, and for vision, Mimo 2.5. The Gemma 4 models, Hy3, and Stepfun 3.7 Flash also get thrown into the local AI mix. Orchestrating tasks with that stable of brains has been effective for such a wide variety of workflows.

One thing I think important to note is that Qwen 3.5 122b (MOE) performs more than twice as fast as 3.8 27b (dense) on both Strix Halo and DGX Spark machines. That bigger class of model has much more world knowledge than 27 billion parameters can ever contain. So, I'm very eager to see if Alibaba will release a 120ish billion parameter version of 3.8. I have a sense that a new 100b+ sized MOE model from Qwen could likely be a Deepseek V4 Flash killer.

So for now, I'm disappointed with Qwen 3.8 27b. I'll be extremely eager to try any mixture of experts Qwen 3.8 model. I expect a 35b or 122b size MOE model will likely be a serious improvement over other comparable options. I'll also try other harnesses to see if fewer task termination issues are experienced - and/or perhaps an update to Pi will help reduce the problems I'm currently seeing with 3.8 27b. Either way, I think we'll find that in practice, 3.8 27b is not the magic pill so many people hoped it would be. I'll keep my ears open, and continue testing...

Version 4Aug 18, 2026 at 12:56

When you watch all the first review videos, everyone's response to Qwen 3.8 27b is as expected - it's being called the best model that can run on small consumer GPUs. And the benchmark hype all seems to show that it produces better output than 3.6, in most tests.

The only problem is that Qwen 3.8 27b thinks a lot. In fact, I'm experiencing significant issues with it terminating before completing tasks, because tasks are running so long. Perhaps another harness may help with that issue (the problematic terminations all occurred in Pi), but I've seen many reviews expressing a similar problem - even small software generations end up using nearly all the 256k context. Results are great, when the job is complete, but I'm regularly seeing jobs not completing.

To be clear, I'm not experiencing long context termination issues with only machines that have small GPUs. The task termination issue has occurred on both DGX Spark and Strix Halo machines, using both the q4 and q6 quants of 3.8 27b. So, this is not an isolated issue with just one version, one machine, or one configuration.

The other issue is that the 27b model is slow for a smallish model. The q4 quant on Strix Halo and the q6 quant on DGX Spark both run at just over 12 tokens per second. That's usable, but no where near the 50-60 tps I typically get with the Qwen 3.6 35a3 MOE model, on those same hardware platforms (and on faster hardware such as RTX 5090 and RTX 6000, the 3.6 MOE model runs over 200 tps).

I understand that historically, the dense 3.6 27b (dense) model is generally expected to produce higher quality output than the 3.6 35a3 (MOE) model, but for the sorts of tasks I've used Qwen 3.6 to accomplish, the MOE version has always done a great job, and at many times the speed. What that means in practice is that I can perform more iterations in less time, and craft better output overall, using the MOE model. Honestly, I haven't needed any better quality than I get from the MOE model. It's fantastic. When I need more world knowledge and capability, I use Deepseek V4 Flash, and for vision, Mimo 2.5. The Gemma 4 models, Hy3, and Stepfun 3.7 Flash also get thrown into the local AI mix. Orchestrating tasks with that stable of brains has been effective for such a wide variety of workflows.

One thing I think important to note is that Qwen 3.5 122b (MOE) performs more than twice as fast as 3.8 27b (dense) on both Strix Halo and DGX Spark machines. That bigger class of model has much more world knowledge than 27 billion parameters can ever contain. So, I'm very eager to see if Alibaba will release a 120ish billion parameter version of 3.8. I have a sense that a new 100b+ sized MOE model from Qwen could likely be a Deepseek V4 Flash killer.

So for now, I'm disappointed with Qwen 3.8 27b. I'll be extremely eager to try any mixture of experts Qwen 3.8 model. I expect a 35b or 122b size MOE model will likely be a serious improvement over other comparable options. I'll also try other harnesses to see if fewer task termination issues are experienced - and/or perhaps an update to Pi will help reduce the problems I'm currently seeing with 3.8 27b. Either way, I think we'll find that in practice, 3.8 27b is not the magic pill so many people hoped it would be. I'll keep my ears open, and continue testing...

Version 3Aug 16, 2026 at 14:40

When you watch all the first review videos, everyone's response to Qwen 3.8 27b is as expected - it's being called the best model that can run on small consumer GPUs. And the benchmark hype all seems to show that it produces better output than 3.6, in most tests.

The only problem is that Qwen 3.8 27b thinks a lot. In fact, I'm experiencing significant issues with it terminating before completing tasks, because tasks are running so long. Perhaps another harness may help with that issue (the problematic terminations all occurred in Pi), but I've seen many reviews expressing a similar problem - even small software generations end up using nearly all the 256k context. Results are great, when the job is complete, but I'm regularly seeing jobs not completing.

To be clear, I'm not experiencing long context termination issues with one of my machines with a small GPU. This issue has occurred on both DGX Spark and Strix Halo machines, using both the q4 and q6 quants of 3.8 27b. So, this is not an isolated issue with just one version, one machine, or one configuration.

The other issue is that the 27b model is slow for a smallish model. The q4 quant on Strix Halo and the q6 quant on DGX Spark both run at just over 12 tokens per second. That's usable, but no where near the 50-60 tps I typically get with the Qwen 3.6 35a3 MOE model, on those same hardware platforms (and on faster hardware such as RTX 5090 and RTX 6000, the 3.6 MOE model runs over 200 tps!).

I understand that historically, the dense 3.6 27b (dense) model is generally expected to produce higher quality output than the 3.6 35a3 (MOE) model, but for the sorts of tasks I've used Qwen 3.6 to accomplish, the MOE version has always done a great job, and at many times the speed. What that means in practice is that I can perform more iterations in less time, and craft better output using the MOE model. Honestly, I haven't needed any better quality than I get from the MOE model. It's fantastic. When I need more world knowledge and capability, I use Deepseek V4 Flash, and for vision, Mimo 2.5. The Gemma 4 models, Hy3, and Stepfun 3.7 Flash also get thrown into the local AI mix.

One thing I think important to note is that Qwen 3.5 122b (MOE) performs more than twice as fast as 3.8 27b (dense) on both Strix Halo and DGX Spark machines. That bigger class of model has much more world knowledge than a 27b can have. So, I'm very eager to see if Alibaba will release a 120ish billion parameter version of 3.8. I have a sense that a new 100b+ sized MOE model from Qwen could likely be a Deepseek V4 Flash killer.

So for now, I'm disappointed with Qwen 3.8 27b. I'll be extremely eager to try any mixture of experts Qwen 3.8 model. I expect a 35b or 122b size MOE model will likely be a serious improvement over other comparable options. I'll also try other harnesses to see if fewer task termination issues are experienced - and/or perhaps an update to Pi will help reduce the problems I'm currently seeing with 3.8 27b. Either way, I think we'll find that in practice, 3.8 27b is not the magic pill so many people hoped it would be. I'll keep my ears open, and continue testing...

Version 2Aug 16, 2026 at 14:33

When you watch all the first review videos, everyone's response to Qwen 3.8 27b is as expected - it's being called the best model that can run on small consumer GPUs. And the benchmark hype all seems to show that it produces better output than 3.6, in most tests.

The only problem is that Qwen 3.8 27b thinks a lot. In fact, I'm experiencing significant issues with it terminating before completing tasks, because tasks are running so long. Perhaps another harness may help with that issue (the problematic terminations all occurred in Pi), but I've seen many reviews expressing a similar problem - even small software generations end up using nearly all the 256k context. Results are great, when the job is complete, but I'm regularly seeing jobs not completing.

To be clear, I'm not experiencing long context termination issues with one of my machines with a small GPU. This issue has occurred on both DGX Spark and Strix Halo machines, using both the q4 and q6 quants of 3.8 27b. So, this is not an isolated issue with just one version, on machine, or one configuration.

The other issue is that the 27b model is slow for a smallish model. The q4 quant on Strix Halo and the q5 quant on DGX Spark both run at just over 12 tokens per second. That's usable, but no where near the 50-60 tps I typically get with the Qwen 3.6 35a3 MOE model, on those same hardware platforms (and on faster hardware such as RTX 5090 and 600, the 3.6 MOE model runs over 200 tps!).

I understand that the dense 3.6 27b (dense) model generally produces higher quality output than the 3.6 35a3 (MOE) model, but for the sorts of tasks I've used Qwen 3.6 to accomplish, the MOE version has always done a great job, and at many times the speed. Honestly, I haven't needed any better quality than I get from the MOE model. It's fantastic. When I need more world knowledge and capability, I use Deepseek V4 Flash, and for vision, Mimo 2.5. The Gemma 4 models, Hy3, and Stepfun 3.7 Flash also get thrown into the local mix.

One thing I think important to note is that Qwen 3.5 122b (MOE) performs more than twice as fast as 3.8 27b (dense) on both Strix Halo and DGX Spark machines. That bigger class of model has much more world knowledge than a 27b can have. So, I'm very eager to see if Alibaba will release a 120ish billion parameter version of 3.8. I have a sense that a new 100b+ sized model could likely be a Deepseek V4 Flash killer.

So for now, I'm disappointed with Qwen 3.8 27b. I'll be extremely eager to try any mixture of experts Qwen 3.8 model. I expect a 35b or 122b size MOE model with likely be a serious improvement over other comparable options. I'll also try other harnesses to see if fewer issues are expected - and/or perhaps an update to Pi will help reduce the problems I'm currently seeing the 3.8 27b. Either way, I think we'll find that 3.8 27b is not the magic pill so many people hoped it would be. I'll keep my ears open, and continue testing...

Version 1Aug 16, 2026 at 14:27

When you watch all the review videos, everyone's response to 3.8 is as expected - Qwen 3.8 27b is the best model that can run on small consumer GPUs. The benchmarks show that it produces better output that 3.6 in most tests.

The only problem is that it thinks a lot. In fact, I'm experiencing significant issues with it terminating before completing tasks, because they are running so long. Perhaps another harness may help (the problematic terminations all occurred in Pi), but I've seen many reviews expressing the same issue - even small software generations end up using nearly all the 256k context.

To be clear, I'm not experience the long context termination issues with one of my small GPUs. It's happened on both the DGX Spark and Strix Halo machines, using both the q4 and q6 quants. So, this is not an isolated issue with just one version, machine, or configuration.

The other issue is that the 27b model is slow for a smallish model. The q4 quant on Strix Halo and the q5 quant on DGX Spark both run at just over 12 tokens per second. That's usable, but no where near the 50-60 tps I get with the Qwen 3.6 35a3 MOE model, on those same hardware platforms (and on faster hardware such as RTX 5090 and 600, that model runs over 200 tps).

I understand that the dense 27b 3.6 model generally produces higher quality output than the MOE 35b model, but for the sorts of tasks I've used Qwen 3.6 to accomplish, the MOE version has always done a great job, and at blinding speed. Honestly, I haven't needed any better quality than I get from the MOE model. It's fantastic. When I need more world knowledge and capability, I use Deepseek V4 Flash, and for vision, Mimo 2.5. The Gemma 4 models, Hy3, and Stepful 3.7 Flash also get thrown into the local mix.

One thing that I think is important to note is that Qwen 3.5 122b (MOE) performs more than twice as fast as 3.8 27b (dense) on both the Strix Halo and DGX Spark machines. That bigger class of model has much more world knowledge than a 27b can have. So, I'm very eager to see if Alibaba will release a 120ish billion parameter version of 3.8. I have a sense that a new 100b+ sized model could likely be a Deepseek V4 killer.

So for now, I'm disappointed with Qwen 3.8. I'll be extremely eager to try any mixture of experts 3.8 model. I expect a 35b and/or 122b size MOE model with likely be serious improvement over other options. I'll also try other harnesses to see if fewer issues are expected - and/or perhaps an update to Pi will help reduce the problems I'm currently seeing the 3.8 27b.