Both projects are different in scope. Think of slotstream as optimizing for memory and for this specific model for now, my intention is not to build an inference engine the same as oMLX
Of course you can use the native UI of all the apps in your ecosystem, the biggest feature of Hermes for me personally is that I can run any task in any of my 30 or so self hosted tools from a single chat interface (matrix), which is also quite secure. No longer do I need 30 open tabs and lots of clicking around, one sentence in my favorite chat app (even on the go in the phone), and many tasks can be executed at once. Unification of control.
And despite the enormous capital expenditure, Chinese models are nipping at their heels at what must be a fraction of the cost. Sometimes constraints are healthy for inducing creative solutions.
You add requirements and make previous tests invisible to see how pigeon brained the model is - Sol and Fable seem to rank the same as Opus tends to fall behind
This seems very cool, but I'm not sure I understand exactly what it's doing. Are they making a new speculative drafter for Qwen 3.8 27B? Maybe they're optimizing the MLX code for the decoder itself? Thank you in advance.
reply