This comparison got buried a while back in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread, with some new examples and more detailed results.
I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.
These examples were all performed on local hardware. They were all created by the same initial prompt: 'Please create a rubiks_modelname.html 3d rubiks cube that the user can interact with, and which has the option to start with a randomly mixed up cube, and can visually show the steps to solve, as a 3D motion demo'
- https://com-pute.com/nick/rubikshy3.html
- https://com-pute.com/nick/rubiks--glm53f-iq3xxs.html
- https://com-pute.com/nick/rubiks_cube.html (Deepseek)
- https://com-pute.com/nick/qwenrubiks.html (Qwen 3.6 35a3 MOE)
- https://com-pute.com/nick/rubiksqwen38fnxt.html
- https://com-pute.com/nick/rubiksmimo25.html
- https://com-pute.com/nick/rubiks--mimo26f.html
- https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail.html
Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model which actually created a real Rubik's cube solver. In the application it created, you can scramble the cube, or enter any random moves, and the software reasons a complete working solution, from the ground up, using solving logic derived from first principles. All the other applications simply record moves which have been entered in the UI, and play them in reverse. That's a cheat which doesn't include any real logic that can be used to solve actual cubes. The Qwen 3.8 Flash Next model took a long time to complete this solution, but wow, it accomplished everything from a single prompt, basically completely unattended, without any manual iterations (the server timed out a few times so Pi paused, but a simple 'please continue' was all the model required to complete the task). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:
I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a human approaching the problem might be expected to work. It created the simplest possible working solution, and then I worked with it in steps, to add features. It tends to break down problems into engineering steps, and confidently completes reasonable improvements with interactive guidance. That's a style which feels to me very controllable, with more human intention involved. Mimo 2.6 flash has now been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:
As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE LLM runs very quickly on laptops which I've purchased for as little as $800, and which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without any iterations. That's truly a fantastic outcome for such a small model.
GLM 5.3 Flash and Deepseek V4 Flash have been my most used, most trusted self-hosted models, but for this task, Hy3 was most efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only requires a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's an Hy3 Pi session export for one of the Hy3 generations - note that this session was from a full precision version of Hy3 hosted on Openrouter. I no longer have the session export which used the locally hosted version - all other sessions in this case study were from locally hosted model generations:
Here is the session export for the Deepseek V4 Flash Antirez DS4 mixed 2-4 bit quant, running on a single DGX Spark. The 4 bit quant which requires 2 clustered DGX Spark machines is actually significantly more capable and reliable, so this result is more impressive than it appears at first. The 2-bit model worked fine, but did require some feedback:
Here's the export for the GLM 5.3 Flash session, which used the IQ3_XXS quant on 2 clustered DGX Spark machines. It didn't require any iterations:
And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:
Here's a version created by Muse Spark 1.3 Contributor, on the API (not on a locally hosted machine):
https://com-pute.com/nick/rubiks_modelname--musespark13contributor.html