Post History

Current version by Nick Antonaccio

Current VersionSep 27, 2026 at 00:33

This comparison got buried a while back in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread, with some new examples and more detailed results.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware. They were all created by the same initial prompt: 'Please create a rubiks_modelname.html 3d rubiks cube that the user can interact with, and which has the option to start with a randomly mixed up cube, and can visually show the steps to solve, as a 3D motion demo'

Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model which actually created a real Rubik's cube solver. In the application it created, you can scramble the cube, or enter any random moves, and the software reasons a complete working solution, from the ground up, using solving logic derived from first principles. All the other applications simply record moves which have been entered in the UI, and play them in reverse. That's a cheat which doesn't include any real logic that can be used to solve actual cubes. The Qwen 3.8 Flash Next model took a long time to complete this solution, but wow, it accomplished everything from a single prompt, basically completely unattended, without any manual iterations (the server timed out a few times so Pi paused, but a simple 'please continue' was all the model required to complete the task). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:

https://com-pute.com/nick/rubiks_cube_solver--qwen38fnxt--pi-session-2026-09-21T16-08-57-242Z_01a0c4ba-4697-7370-978e-5fd491b6623c.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a human approaching the problem might be expected to work. It created the simplest possible working solution, and then I worked with it in steps, to add features. It tends to break down problems into engineering steps, and confidently completes reasonable improvements with interactive guidance. That's a style which feels to me very controllable, with more human intention involved. Mimo 2.6 flash has now been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE LLM runs very quickly on laptops which I've purchased for as little as $800, and which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without any iterations. That's truly a fantastic outcome for such a small model.

GLM 5.3 Flash and Deepseek V4 Flash have been my most used, most trusted self-hosted models, but for this task, Hy3 was most efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only requires a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's an Hy3 Pi session export for one of the Hy3 generations - note that this session was from a full precision version of Hy3 hosted on Openrouter. I no longer have the session export which used the locally hosted version - all other sessions in this case study were from locally hosted model generations:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

Here is the session export for the Deepseek V4 Flash Antirez DS4 mixed 2-4 bit quant, running on a single DGX Spark. The 4 bit quant which requires 2 clustered DGX Spark machines is actually significantly more capable and reliable, so this result is more impressive than it appears at first. The 2-bit model worked fine, but did require some feedback:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

Here's the export for the GLM 5.3 Flash session, which used the IQ3_XXS quant on 2 clustered DGX Spark machines. It didn't require any iterations:

https://com-pute.com/nick/rubiks-cube--glm5.3-flash--pi-session-2026-08-30T14-25-45-086Z_01a0530f-e27e-7c43-bc40-9c30a4003386.html

And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Here's a version created by Muse Spark 1.3 Contributor, on the API (not on a locally hosted machine):

https://com-pute.com/nick/rubiks_modelname--musespark13contributor.html

Previous Versions
Version 18Sep 27, 2026 at 00:33

This comparison got buried a while back in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread, with some new examples and more detailed results.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware. They were all created by the same initial prompt: 'Please create a rubiks_modelname.html 3d rubiks cube that the user can interact with, and which has the option to start with a randomly mixed up cube, and can visually show the steps to solve, as a 3D motion demo'

Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model which actually created a real Rubik's cube solver. In the application it created, you can scramble the cube, or enter any random moves, and the software reasons a complete working solution, from the ground up, using solving logic derived from first principles. All the other applications simply record moves which have been entered in the UI, and play them in reverse. That's a cheat which doesn't include any real logic that can be used to solve actual cubes. The Qwen 3.8 Flash Next model took a long time to complete this solution, but wow, it accomplished everything from a single prompt, basically completely unattended, without any manual iterations (the server timed out a few times so Pi paused, but a simple 'please continue' was all the model required to complete the task). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:

https://com-pute.com/nick/rubiks_cube_solver--qwen38fnxt--pi-session-2026-09-21T16-08-57-242Z_01a0c4ba-4697-7370-978e-5fd491b6623c.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a human approaching the problem might be expected to work. It created the simplest possible working solution, and then I worked with it in steps, to add features. It tends to break down problems into engineering steps, and confidently completes reasonable improvements with interactive guidance. That's a style which feels to me very controllable, with more human intention involved. Mimo 2.6 flash has now been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE LLM runs very quickly on laptops which I've purchased for as little as $800, and which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without any iterations. That's truly a fantastic outcome for such a small model.

GLM 5.3 Flash and Deepseek V4 Flash have been my most used, most trusted self-hosted models, but for this task, Hy3 was most efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only requires a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's an Hy3 Pi session export for one of the Hy3 generations - note that this session was from a full precision version of Hy3 hosted on Openrouter. I no longer have the session export which used the locally hosted version - all other sessions in this case study were from locally hosted model generations:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

Here is the session export for the Deepseek V4 Flash Antirez DS4 mixed 2-4 bit quant, running on a single DGX Spark. The 4 bit quant which requires 2 clustered DGX Spark machines is actually significantly more capable and reliable, so this result is more impressive than it appears at first. The 2-bit model worked fine, but did require some feedback:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

Here's the export for the GLM 5.3 Flash session, which used the IQ3_XXS quant on 2 clustered DGX Spark machines. It didn't require any iterations:

https://com-pute.com/nick/rubiks-cube--glm5.3-flash--pi-session-2026-08-30T14-25-45-086Z_01a0530f-e27e-7c43-bc40-9c30a4003386.html

And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Here's a version created by Muse Spark 1.3 Contributor, on the API (not on a locally hosted machine):

https://com-pute.com/nick/rubiks_modelname--musespark13contributor.html

Version 17Sep 26, 2026 at 20:32

This comparison got buried a while back in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread, with some new examples and more detailed results.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware. They were all created by the same initial prompt: 'Please create a rubiks_modelname.html 3d rubiks cube that the user can interact with, and which has the option to start with a randomly mixed up cube, and can visually show the steps to solve, as a 3D motion demo'

Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model which actually created a real Rubik's cube solver. In the application it created, you can scramble the cube, or enter any random moves, and the software reasons a complete working solution, from the ground up, using solving logic derived from first principles. All the other applications simply record moves which have been entered in the UI, and play them in reverse. That's a cheat which doesn't include any real logic that can be used to solve actual cubes. The Qwen 3.8 Flash Next model took a long time to complete this solution, but wow, it accomplished everything from a single prompt, basically completely unattended, without any manual iterations (the server timed out a few times so Pi paused, but a simple 'please continue' was all the model required to complete the task). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:

https://com-pute.com/nick/rubiks_cube_solver--qwen38fnxt--pi-session-2026-09-21T16-08-57-242Z_01a0c4ba-4697-7370-978e-5fd491b6623c.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a human approaching the problem might be expected to work. It created the simplest possible working solution, and then I worked with it in steps, to add features. It tends to break down problems into engineering steps, and confidently completes reasonable improvements with interactive guidance. That's a style which feels to me very controllable, with more human intention involved. Mimo 2.6 flash has now been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE LLM runs very quickly on laptops which I've purchased for as little as $800, and which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without any iterations. That's truly a fantastic outcome for such a small model.

GLM 5.3 Flash and Deepseek V4 Flash have been my most used, most trusted self-hosted models, but for this task, Hy3 was most efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only requires a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's an Hy3 Pi session export for one of the Hy3 generations - note that this session was from a full precision version of Hy3 hosted on Openrouter. I no longer have the session export which used the locally hosted version - all other sessions in this case study were from locally hosted model generations:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

Here is the session export for the Deepseek V4 Flash Antirez DS4 mixed 2-4 bit quant, running on a single DGX Spark. The 4 bit quant which requires 2 clustered DGX Spark machines is actually significantly more capable and reliable, so this result is more impressive than it appears at first. The 2-bit model worked fine, but did require some feedback:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

Here's the export for the GLM 5.3 Flash session, which used the IQ3_XXS quant on 2 clustered DGX Spark machines. It didn't require any iterations:

https://com-pute.com/nick/rubiks-cube--glm5.3-flash--pi-session-2026-08-30T14-25-45-086Z_01a0530f-e27e-7c43-bc40-9c30a4003386.html

And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Version 16Sep 26, 2026 at 12:09

This comparison got buried a while back in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread, with some additional information.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware. They were all created by the same initial prompt: 'Please create a rubiks_modelname.html 3d rubiks cube that the user can interact with, and which has the option to start with a randomly mixed up cube, and can visually show the steps to solve, as a 3D motion demo'

Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model which actually created a real Rubik's cube solver. In the application it created, you can scramble the cube, or enter any random moves, and the software reasons a complete working solution, from the ground up, using solving logic derived from first principles. All the other applications simply record moves which have been entered in the UI, and play them in reverse. That's a cheat which doesn't include any real logic that can be used to solve actual cubes. The Qwen 3.8 Flash Next model took a long time to complete this solution, but wow, it accomplished everything from a single prompt, basically completely unattended, without any manual iterations (the server timed out a few times so Pi paused, but a simple 'please continue' was all the model required to complete the task). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:

https://com-pute.com/nick/rubiks_cube_solver--qwen38fnxt--pi-session-2026-09-21T16-08-57-242Z_01a0c4ba-4697-7370-978e-5fd491b6623c.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a human approaching the problem might be expected to work. It created the simplest possible working solution, and then I worked with it in steps, to add features. It tends to break down problems into engineering steps, and confidently completes reasonable improvements with interactive guidance. That's a style which feels to me very controllable, with more human intention involved. Mimo 2.6 flash has now been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE LLM runs very quickly on laptops which I've purchased for as little as $800, and which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without any iterations. That's truly a fantastic outcome for such a small model.

GLM 5.3 Flash and Deepseek V4 Flash have been my most used, most trusted self-hosted models, but for this task, Hy3 was most efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only requires a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's an Hy3 Pi session export for one of the Hy3 generations - note that this session was from a full precision version of Hy3 hosted on Openrouter. I no longer have the session export which used the locally hosted version - all other sessions in this case study were from locally hosted model generations:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

Here is the session export for the Deepseek V4 Flash Antirez DS4 mixed 2-4 bit quant, running on a single DGX Spark. The 4 bit quant which requires 2 clustered DGX Spark machines is actually significantly more capable and reliable, so this result is more impressive than it appears at first. The 2-bit model worked fine, but did require some feedback:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

Here's the export for the GLM 5.3 Flash session, which used the IQ3_XXS quant on 2 clustered DGX Spark machines. It didn't require any iterations:

https://com-pute.com/nick/rubiks-cube--glm5.3-flash--pi-session-2026-08-30T14-25-45-086Z_01a0530f-e27e-7c43-bc40-9c30a4003386.html

And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Version 15Sep 22, 2026 at 18:16

This comparison got buried a while back in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread, with some additional information.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware. They were all created by the same initial prompt: 'Please create a rubiks_modelname.html 3d rubiks cube that the user can interact with, and which has the option to start with a randomly mixed up cube, and can visually show the steps to solve, as a 3D motion demo'

GLM 5.3 Flash and Deepseek V4 Flash have been my most used, most trusted self-hosted models, but for this task, Hy3 was most efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only requires a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's an Hy3 Pi session export for one of the Hy3 generations - note that this session was from a full precision version of Hy3 hosted on Openrouter. I no longer have the session export which used the locally hosted version - all other sessions in this case study were from locally hosted model generations:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model which actually created a real Rubik's cube solver. In the application it created, you can scramble the cube, or enter any random moves, and the software reasons a complete working solution, from the ground up, using solving logic derived from first principles. All the other applications simply record moves which have been entered in the UI, and play them in reverse. That's a cheat which doesn't include any real logic that can be used to solve actual cubes. The Qwen 3.8 Flash Next model took a long time to complete this solution, but wow, it accomplished everything from a single prompt, basically completely unattended, without any manual iterations (the server timed out a few times so Pi paused, but a simple 'please continue' was all the model required to complete the task). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:

https://com-pute.com/nick/rubiks_cube_solver--qwen38fnxt--pi-session-2026-09-21T16-08-57-242Z_01a0c4ba-4697-7370-978e-5fd491b6623c.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a human approaching the problem might be expected to work. It created the simplest possible working solution, and then I worked with it in steps, to add features. It tends to break down problems into engineering steps, and confidently completes reasonable improvements with interactive guidance. That's a style which feels to me very controllable, with more human intention involved. Mimo 2.6 flash has now been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE LLM runs very quickly on laptops which I've purchased for as little as $800, and which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without any iterations. That's truly a fantastic outcome for such a small model.

Here is the session export for the Deepseek V4 Flash Antirez DS4 mixed 2-4 bit quant, running on a single DGX Spark. The 4 bit quant which requires 2 clustered DGX Spark machines is actually significantly more capable and reliable, so this result is more impressive than it appears at first. The 2-bit model worked fine, but did require some feedback:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

Here's the export for the GLM 5.3 Flash session, which used the IQ3_XXS quant on 2 clustered DGX Spark machines. It didn't require any iterations:

https://com-pute.com/nick/rubiks-cube--glm5.3-flash--pi-session-2026-08-30T14-25-45-086Z_01a0530f-e27e-7c43-bc40-9c30a4003386.html

And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Version 14Sep 22, 2026 at 18:11

This comparison got buried a while back in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread, with some additional information.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware. They were all created by the same initial prompt.

GLM 5.3 Flash and Deepseek V4 Flash have been my most used, most trusted self-hosted models, but for this task, Hy3 was most efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only requires a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's an Hy3 Pi session export for one of the Hy3 generations - note that this session was from a full precision version of Hy3 hosted on Openrouter. I no longer have the session export which used the locally hosted version - all other sessions in this case study were from locally hosted model generations:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model which actually created a real Rubik's cube solver. In the application it created, you can scramble the cube, or enter any random moves, and the software reasons a complete working solution, from the ground up, using solving logic derived from first principles. All the other applications simply record moves which have been entered in the UI, and play them in reverse. That's a cheat which doesn't include any real logic that can be used to solve actual cubes. The Qwen 3.8 Flash Next model took a long time to complete this solution, but wow, it accomplished everything from a single prompt, basically completely unattended, without any manual iterations (the server timed out a few times so Pi paused, but a simple 'please continue' was all the model required to complete the task). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:

https://com-pute.com/nick/rubiks_cube_solver--qwen38fnxt--pi-session-2026-09-21T16-08-57-242Z_01a0c4ba-4697-7370-978e-5fd491b6623c.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a human approaching the problem might be expected to work. It created the simplest possible working solution, and then I worked with it in steps, to add features. It tends to break down problems into engineering steps, and confidently completes reasonable improvements with interactive guidance. That's a style which feels to me very controllable, with more human intention involved. Mimo 2.6 flash has now been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE LLM runs very quickly on laptops which I've purchased for as little as $800, and which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without any iterations. That's truly a fantastic outcome for such a small model.

Here is the session export for the Deepseek V4 Flash Antirez DS4 mixed 2-4 bit quant, running on a single DGX Spark. The 4 bit quant which requires 2 clustered DGX Spark machines is actually significantly more capable and reliable, so this result is more impressive than it appears at first. The 2-bit model worked fine, but did require some feedback:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

Here's the export for the GLM 5.3 Flash session, which used the IQ3_XXS quant on 2 clustered DGX Spark machines. It didn't require any iterations:

https://com-pute.com/nick/rubiks-cube--glm5.3-flash--pi-session-2026-08-30T14-25-45-086Z_01a0530f-e27e-7c43-bc40-9c30a4003386.html

And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Version 13Sep 22, 2026 at 14:01

This comparison got buried a while back in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread, with some additional information.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware. They were all created by the same initial prompt.

GLM 5.3 Flash and Deepseek V4 Flash have been my most used, most trusted self-hosted models, but for this task, Hy3 was most efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only requires a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's an Hy3 Pi session export for one of the Hy3 generations - note that this session was from a full precision version of Hy3 hosted on Openrouter. I no longer have the session export which used the locally hosted version - all other sessions in this case study were from locally hosted model generations:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model which actually created a real Rubik's cube solver. In the application it created, you can scramble the cube, or enter any random moves, and the software reasons a complete working solution, from the ground up, using solving logic derived from first principles. All the other applications simply record moves which have been entered in the UI, and play them in reverse. That's a cheat that doesn't include any real logic that can be used to solve actual cubes. The Qwen 3.8 Flash Next model took a long time to complete this solution, but wow, it accomplished everything from a single prompt, basically completely unattended, without any manual iterations (the server timed out a few times so Pi paused, but a simple 'please continue' was all the model required to complete the task). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:

https://com-pute.com/nick/rubiks_cube_solver--qwen38fnxt--pi-session-2026-09-21T16-08-57-242Z_01a0c4ba-4697-7370-978e-5fd491b6623c.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a human approaching the problem might be expected to work. It created the simplest possible working solution, and then I worked with it in steps, to add features. It tends to break down problems into engineering steps, and confidently completes reasonable improvements with interactive guidance. That's a style which feels to me very controllable, with more human intention involved. Mimo 2.6 flash has now been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE LLM runs very quickly on laptops which I've purchased for as little as $800, and which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without any iterations. That's truly a fantastic outcome for such a small model.

Here is the session export for the Deepseek V4 Flash Antirez DS4 mixed 2-4 bit quant, running on a single DGX Spark. The 4 bit quant which requires 2 clustered DGX Spark machines is actually significantly more capable and reliable, so this result is more impressive than it appears at first. The 2-bit model worked fine, but did require some feedback:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

Here's the export for the GLM 5.3 Flash session, which used the IQ3_XXS quant on 2 clustered DGX Spark machines. It didn't require any iterations:

https://com-pute.com/nick/rubiks-cube--glm5.3-flash--pi-session-2026-08-30T14-25-45-086Z_01a0530f-e27e-7c43-bc40-9c30a4003386.html

And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Version 12Sep 22, 2026 at 13:59

This comparison got buried a while back in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread, with some additional information.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware. They were all created by the same initial prompt.

GLM 5.3 Flash and Deepseek V4 Flash have been my most used, most trusted self-hosted models, but for this task, Hy3 was most efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only requires a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's an Hy3 Pi session export for one of the Hy3 generations - note that this session was from a full precision version of Hy3 hosted on Openrouter. I no longer have the session export which used the locally hosted version - all other sessions in this case study were from locally hosted model generations:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model which actually created a real Rubik's cube solver. In the application it created, you can scramble the cube, or enter any random moves, and the software reasons a complete working solution, from the ground up, using solving logic derived from first principles. All the other applications simply record moves which have been entered in the UI, and play them in reverse. That's a cheat that doesn't include any real logic that can be used to solve actual cubes. The Qwen 3.8 Flash Next model took a long time to complete this solution, but wow, it accomplished everything from a single prompt, basically completely unattended, without any manual iterations (the server timed out a few times so Pi paused, but a simple 'please continue' was all the model required to complete the task). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:

https://com-pute.com/nick/rubiks_cube_solver--qwen38fnxt--pi-session-2026-09-21T16-08-57-242Z_01a0c4ba-4697-7370-978e-5fd491b6623c.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a human approaching the problem might be expected to work. It created the simplest possible working solution, and then I worked with it in steps, to add features. It tends to break down problems into engineering steps, and confidently completes reasonable improvements with interactive guidance. That's a style which feels to me very controllable, with more human intention involved. Mimo 2.6 flash has now been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE LLM runs very quickly on laptops which I've purchased for as little as $800, and which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without any iterations. That's truly a fantastic outcome for such a small model.

Here is the session export for the Deepseek V4 Flash Antirez DS4 mixed 2-4 bit quant, running on a single DGX Spark. The 4 bit quant which requires 2 clustered DGX Spark machines is actually significantly more capable and reliable, so this result is more impressive than it appears at first. The 2-bit model worked fine, but did require some feedback:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

Here's the export for the GLM 5.3 Flash session. It didn't require any iterations:

https://com-pute.com/nick/rubiks-cube--glm5.3-flash--pi-session-2026-08-30T14-25-45-086Z_01a0530f-e27e-7c43-bc40-9c30a4003386.html

And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Version 11Sep 22, 2026 at 13:58

This comparison got buried a while back in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread, with some additional information.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware. They were all created by the same initial prompt.

GLM 5.3 Flash and Deepseek V4 Flash have been my most used, most trusted self-hosted models, but for this task, Hy3 was most efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only requires a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's the Hy3 Pi session export (note that this session was from a full precision version of Hy3 hosted on Openrouter - I no longer have the session export which used the locally version - all other sessions in this case study were from locally hosted models):

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model which actually created a real Rubik's cube solver. In the application it created, you can scramble the cube, or enter any random moves, and the software reasons a complete working solution, from the ground up, using solving logic derived from first principles. All the other applications simply record moves which have been entered in the UI, and play them in reverse. That's a cheat that doesn't include any real logic that can be used to solve actual cubes. The Qwen 3.8 Flash Next model took a long time to complete this solution, but wow, it accomplished everything from a single prompt, basically completely unattended, without any manual iterations (the server timed out a few times so Pi paused, but a simple 'please continue' was all the model required to complete the task). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:

https://com-pute.com/nick/rubiks_cube_solver--qwen38fnxt--pi-session-2026-09-21T16-08-57-242Z_01a0c4ba-4697-7370-978e-5fd491b6623c.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a human approaching the problem might be expected to work. It created the simplest possible working solution, and then I worked with it in steps, to add features. It tends to break down problems into engineering steps, and confidently completes reasonable improvements with interactive guidance. That's a style which feels to me very controllable, with more human intention involved. Mimo 2.6 flash has now been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE LLM runs very quickly on laptops which I've purchased for as little as $800, and which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without any iterations. That's truly a fantastic outcome for such a small model.

Here is the session export for the Deepseek V4 Flash Antirez DS4 mixed 2-4 bit quant, running on a single DGX Spark. The 4 bit quant which requires 2 clustered DGX Spark machines is actually significantly more capable and reliable, so this result is more impressive than it appears at first. The 2-bit model worked fine, but did require some feedback:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

Here's the export for the GLM 5.3 Flash session. It didn't require any iterations:

https://com-pute.com/nick/rubiks-cube--glm5.3-flash--pi-session-2026-08-30T14-25-45-086Z_01a0530f-e27e-7c43-bc40-9c30a4003386.html

And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Version 10Sep 22, 2026 at 13:56

This comparison got buried a while back in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread, with some additional information.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware. They were all created by the same initial prompt.

GLM 5.3 Flash and Deepseek V4 Flash have been my most used, most trusted self-hosted models, but for this task, Hy3 was most efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only requires a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's the Hy3 Pi session export (note that this session was from a full precision version of Hy3 hosted on Openrouter - I no longer have the session export which used the locally version - all other sessions in this case study were from locally hosted models):

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model which actually created a real Rubik's cube solver. In the application it created, you can scramble the cube, or enter any random moves, and the software reasons a complete working solution, from the ground up, using solving logic derived from first principles. All the other applications simply record moves which have been entered in the UI, and play them in reverse. That's a cheat that doesn't include any real logic that can be used to solve actual cubes. The Qwen 3.8 Flash Next model took a long time to complete this solution, but wow, it accomplished everything from a single prompt, basically completely unattended, without any manual iterations (the server timed out a few times so Pi paused, but a simple 'please continue' was all the model required to complete the task). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:

https://com-pute.com/nick/rubiks_cube_solver--qwen38fnxt--pi-session-2026-09-21T16-08-57-242Z_01a0c4ba-4697-7370-978e-5fd491b6623c.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a human approaching the problem might be expected to work. It created the simplest possible working solution, and then I worked with it in steps, to add features. It tends to break down problems into engineering steps, and confidently completes reasonable improvements with interactive guidance. That's a style which feels to me very controllable, with more human intention involved. Mimo 2.6 flash has now been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE LLM runs very quickly on laptops which I've purchased for as little as $800, and which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without any iterations. That's truly a fantastic outcome for such a small model.

Here is the session export for Deepseek V4 Flash. It worked fine, but required feedback:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

Here's the export for the GLM 5.3 Flash session. It didn't require any iterations:

https://com-pute.com/nick/rubiks-cube--glm5.3-flash--pi-session-2026-08-30T14-25-45-086Z_01a0530f-e27e-7c43-bc40-9c30a4003386.html

And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Version 9Sep 22, 2026 at 13:52

This comparison got buried in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware. They were all created by the same initial prompt.

GLM 5.3 Flash and Deepseek V4 Flash have been my most used and most trusted self-hosted models, but for this task, Hy3 was the most efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only requires a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's the Hy3 Pi session export (note that this session was from a full precision version of Hy3 hosted on Openrouter - I no longer have the session export which used the locally version - all other sessions in this case study were from locally hosted models):

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model which actually created a real Rubik's cube solver. In the application it created, you can scramble the cube, or enter any random moves, and the software reasons a complete working solution, from the ground up, using solving logic derived from first principles. All the other applications simply record moves which have been entered in the UI, and play them in reverse. That's a cheat that doesn't include any real logic that can be used to solve actual cubes. The Qwen 3.8 Flash Next model took a long time to complete this solution, but wow, it accomplished everything from a single prompt, basically completely unattended, without any manual iterations (the server timed out a few times so Pi paused, but a simple 'please continue' was all the model required to complete the task). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:

https://com-pute.com/nick/rubiks_cube_solver--qwen38fnxt--pi-session-2026-09-21T16-08-57-242Z_01a0c4ba-4697-7370-978e-5fd491b6623c.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a human approaching the problem might be expected to work. It created the simplest possible working solution, and then I worked with it in steps, to add features. It tends to break down problems into engineering steps, and confidently completes reasonable improvements with interactive guidance. That's a style which feels to me very controllable, with more human intention involved. Mimo 2.6 flash has now been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE LLM runs very quickly on laptops which I've purchased for as little as $800, and which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without any iterations. That's truly a fantastic outcome for such a small model.

Here is the session export for Deepseek V4 Flash. It worked fine, but required feedback:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

Here's the export for the GLM 5.3 Flash session. It didn't require any iterations:

https://com-pute.com/nick/rubiks-cube--glm5.3-flash--pi-session-2026-08-30T14-25-45-086Z_01a0530f-e27e-7c43-bc40-9c30a4003386.html

And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Version 8Sep 22, 2026 at 13:49

This comparison got buried in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware. They were all created by the same initial prompt.

GLM 5.3 Flash and Deepseek V4 Flash have been my most used and most trusted self-hosted models, but for this task, Hy3 was the most efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only required a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's the Hy3 Pi session export:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model which actually created a real Rubik's cube solver. In the application it created, you can scramble the cube, or enter any random moves, and the software reasons a complete working solution, from the ground up, using solving logic derived from first principles. All the other applications simply record moves which have been entered in the UI, and play them in reverse. That's a cheat that doesn't include any real logic that can be used to solve actual cubes. The Qwen 3.8 Flash Next model took a long time to complete this solution, but wow, it accomplished everything from a single prompt, basically completely unattended, without any manual iterations (the server timed out a few times so Pi paused, but a simple 'please continue' was all the model required to complete the task). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:

https://com-pute.com/nick/rubiks_cube_solver--qwen38fnxt--pi-session-2026-09-21T16-08-57-242Z_01a0c4ba-4697-7370-978e-5fd491b6623c.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a human approaching the problem might be expected to work. It created the simplest possible working solution, and then I worked with it in steps, to add features. It tends to break down problems into engineering steps, and confidently completes reasonable improvements with interactive guidance. That's a style which feels to me very controllable, with more human intention involved. Mimo 2.6 flash has now been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE LLM runs very quickly on laptops which I've purchased for as little as $800, and which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without any iterations. That's truly a fantastic outcome for such a small model.

Here is the session export for Deepseek V4 Flash. It worked fine, but required feedback:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

Here's the export for the GLM 5.3 Flash session. It didn't require any iterations:

https://com-pute.com/nick/rubiks-cube--glm5.3-flash--pi-session-2026-08-30T14-25-45-086Z_01a0530f-e27e-7c43-bc40-9c30a4003386.html

And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Version 7Sep 22, 2026 at 13:45

This comparison got buried in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware. They all were created by the same initial prompt.

GLM 5.3 Flash and Deepseek V4 Flash have been my most used and most trusted self-hosted models, but for this task, Hy3 was very efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only required a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's the Hy3 Pi session export:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model which actually created a real Rubik's cube solver. In the application it created, you can scramble the cube, or enter any random moves, and the software reasons a complete working solution, from the ground up, using solving logic derived from first principles. All the other applications simply record moves which have been entered in the UI, and play them in reverse. That's a cheat that doesn't include any real solving logic. The Qwen 3.8 Flash Next model took a long time to complete the solution, but wow, it accomplished everything from a single prompt, basically fully unattended, without any manual iterations (the server timed out a couple times so Pi paused, but a simple 'please continue' was all the model required to complete the task). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:

https://com-pute.com/nick/rubiks_cube_solver--qwen38fnxt--pi-session-2026-09-21T16-08-57-242Z_01a0c4ba-4697-7370-978e-5fd491b6623c.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a human approaching the problem may work. It creates the simplest possible working solution, and then you can work with it to add features during iterations. Basically, it breaks up each engineering step, and confidently progresses forward through reasonable steps with guidance. They're something more controllable about that approach, with more human intention involved. Mimo 2.6 flash has been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE LLM runs very quickly on laptops which I purchased for $800, and which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without any iterations. That's truly a fantastic outcome for such a small model!

Here is the session export for Deepseek V4 Flash. It worked fine, but required feedback:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

Here's the export for the GLM 5.3 Flash session. It didn't require any iterations:

https://com-pute.com/nick/rubiks-cube--glm5.3-flash--pi-session-2026-08-30T14-25-45-086Z_01a0530f-e27e-7c43-bc40-9c30a4003386.html

And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Version 6Sep 22, 2026 at 13:38

This comparison got buried in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware. They all were created by the same initial prompt.

GLM 5.3 Flash and Deepseek V4 Flash have been my most used and most trusted models, but for this task, Hy3 was very efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only required a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's the Hy3 Pi session export:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model that actually created a real solver. You can scramble the cube, or enter any random moves, and it will reason a complete working solution, from the ground up, using actual solving logic. All the other applications simply record the moves which have been entered, and play them in reverse. They don't include any logic. The Qwen 3.8 Flash Next model took a very long time to complete the solution, but wow, it accomplished everything from a single prompt, without any manual iterations (the server timed out a couple times and pi paused, but a simple 'please continue' was all the model required). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:

https://com-pute.com/nick/rubiks_cube_solver--qwen38fnxt--pi-session-2026-09-21T16-08-57-242Z_01a0c4ba-4697-7370-978e-5fd491b6623c.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a person approaching the problem may work. It creates the simplest possible working solution, and then you can ask it to add features. Basically, it breaks up each engineering step, and confidently progresses forward through reasonable iterations. They're something more controllable about this approach. Mimo 2.6 flash has been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE model runs very quickly on laptops I purchased for $800, which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without iterations. That's truly fantastic for such a small model

Here is the session export for Deepseek V4 Flash:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Version 5Sep 22, 2026 at 13:22

This comparison got buried in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware.

GLM 5.3 Flash and Deepseek V4 Flash have been my most used and most trusted models, but for this task, I think Hy3 did the best job out of the gate:

Hy3 got it done surprisingly quickly, with the fewest iterations and issues, and it built a nice UI. Hy3 only required a single machine, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's the Hy3 Pi session export:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

And as always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE model runs very quickly on laptops I purchased for $800, which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks.

Version 4Sep 21, 2026 at 18:49

This comparison got buried in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

I think GLM 5.3 Flash and Hy3 did the best job out of the gate:

Hy3 got it done surprisingly quickly, with the fewest iterations and issues, and it built a nice UI. Hy3 only required a single machine, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's the Hy3 Pi session export:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

And as always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE models runs very quickly on laptops I purchased for $800, which have only 16GB VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks.

Version 3Sep 21, 2026 at 17:13

This comparison got buried in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

I think GLM 5.3 Flash and Hy3 did the best job out of the gate:

Hy3 got it done surprisingly quickly, with the fewest iterations and issues, and it built a nice UI. Hy3 only required a single machine, where the GLM 5.3 result was created with a 2 DGX Spark cluster. Here's the Hy3 Pi session export:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

And as always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE models runs very quickly on laptops I purchased for $800, which have only 16GB VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks.

Version 2Sep 21, 2026 at 17:12

This comparison got buried in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread.

I performed a telling comparison between Hy3, Deepseek V4 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver. I think Hy3 did the best job out of the gate:

Hy3 got it done quickly, with the fewest iterations and issues, and it built a nice UI. Here's the Pi session export:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

I've got to note that Mimo 2.5 also did a great job, very quickly:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

And as always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. It runs quickly on laptops I purchased for $800, which have only 16GB VRAM (those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000).

Version 1Sep 21, 2026 at 16:49

This comparison got buried in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread.

I performed a telling comparison between Hy3, Deepseek V4 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver. I think Hy3 did the best job out of the gate:

https://com-pute.com/nick/rubikshy3.html https://com-pute.com/nick/rubiks_cube.html (Deepseek) https://com-pute.com/nick/qwenrubiks.html https://com-pute.com/nick/rubiksmimo25.html https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail.html

Hy3 got it done quickly, with the fewest iterations and issues, and it built a nice UI. Here's the Pi session export:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

I've got to note that Mimo 2.5 also did a great job, very quickly:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

And as always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. It runs quickly on laptops I purchased for $800, which have only 16GB VRAM (those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000).