syntaxing 10 hours ago

Step models were IMO the first local model you can run on 128GB shared memory that worked well. Really excited to see how it compares to Qwen Flash Next.

Edit: bummer, didn’t know it’s 600B-A27B. No way to run that on 228GB.

robertlane0 11 hours ago

Well, according to Artificial Analysis (which I'll admit I've been using as a bit of a mental crutch to avoid comparing models myself, so YMMV), it's smarter and slightly cheaper than Gemini 3.8 Flash, which has been my benchline for "cheap and smart enough", I'll give it a try on OpenCode for the week but I'm not sure I'll be compelled enough to switch from Muse Spark 1.3.

hosel 16 hours ago

Without the pelicans I don’t know what to think

-------------

https://postimg.cc/c6L2MjSt

12m 0s and $0.28

>Draw a Hacker News-style comment thread. Top comment by a user named "pelican_enjoyer": "Without the pelicans I don't know what to think." Reply from "minimaxir" in a grumpy tone: "Since people keep doing it: no, you don't have to make an allusion to Simon's pelicans every time a Hacker News thread about a new LLM pops up. It's a lower-effort joke than even Reddit memes." Beside the thread, show a pelican riding a bicycle, looking smug.

  • behole 15 hours ago

    paging Simon Wilson!

    • JSR_FDED 6 hours ago

      He’ll probably respond quicker if you get his name right

  • minimaxir 15 hours ago

    Since people keep doing it: no, you don't have to make an allusion to Simon's pelicans every time a Hacker News thread about a new LLM pops up. It's a lower-effort joke than even Reddit memes.

    • behole 14 hours ago

      Pelican Police. Cry more.

    • hypfer 14 hours ago

      I'd just read it as social friction, just like I'd read your comment as that very same thing.

      This is the consensus mechanism doing its job, essentially.

      __

      Though to be fair, the way I frame it assumes no connections between nodes and independent choices, when in reality, we have groups supporting each other.

      So it's not necessarily the best mechanism, as social cohesion and other such dysfunctions might be steering away from the objectively correct solution through not necessarily rational biases.

      Or rather not necessarily rational when viewed in just the specific context, but possibly rational when zooming out and considering whole-subsystem health.

      • minimaxir 14 hours ago

        Low-effort content is generally discouraged on Hacker News because it lowers the quality of the site. It's not a deep commentary on socioeconomics.

        • hypfer 14 hours ago

          Like posting the exact same personal brand-building under each new model you mean?

          As said, there is no right or wrong here. Well, technically there is and it is my opinion (obviously), but if we take a step back, it's exactly what I just described and you (unfortunately) discarded through bulldozing.

          The """"thought leaders"""" will have to live with the fact that some people just don't think that their work is adding all that much value. It's a bit unpleasant for the ego of course, but that's kinda the trade when making money with fluff.

    • minimaxir 13 hours ago

      Re: the edit where the OP added an image: I'm more bemused than offended.

    • tomrod 5 hours ago

      I love the pelicans almost as much as the penguins. One day we shall see the max positive of minimaxir ;)

jstummbillig 14 hours ago

Why is that interesting?

  • theturtletalks 13 hours ago

    Stepfun made a close to SOTA model with 3.5. They got overtaken quickly but have shown enough to be given attention when a new model releases.

    It’s also an open model. Even if you only use Opus and Sol, these open models help push them and the frontier.

    • UncleOxidant 6 hours ago

      Seems like it was an under-noticed model back when it came out because there were so many new Qwen models coming out around the same time. I tried it some on my Strix Halo box, but it was right on the edge of fitting.

      • theturtletalks 3 hours ago

        3.5 Flash was decent, but I was using it on OpenRouter. What model are you running on your Strix Halo nowadays?

        • Iolaum 3 hours ago

          Qwen-3.8-flash-next @Q4

  • serf 12 hours ago

    well, it's free for now , beyond that eh.

  • ssivark 2 hours ago

    Back in their 3.5/3.7 era [1], they had a small+fast model, specifically eschewing knowledge, and instead focused on "general intelligence" like managing tool calls, orchestrating sub-agents, etc. Could be a very good substrate for running a personal assistant (time will tell what's the right approach); so I'm curious to try that out and see how they've come along since then.

    [1] https://news.ycombinator.com/item?id=47069179

lykr0n 15 hours ago

Wonder if this was the space bunny alpha model

  • minimaxir 15 hours ago

    After some quick tests, it likely isn't. Additionally, Space Bunny Alpha had hints it was a Minimax-variant model.

  • james2doyle 11 hours ago

    Its been on ZenMux for at least 2 weeks so I doubt they are hiding

MisterMunchkin 13 hours ago

I like trying new models but I wish we’d get something actually new. Like a new architecture or something. LLMs are just so sloppish. We can do better.

  • cyanydeez 13 hours ago

    welcome to the sigmoid...

  • famouswaffles 12 hours ago

    Talk is cheap. You have to actually beat transformers first. All alternate architectures that have sprung up are sidegrades at best.

esafak 15 hours ago

It doesn't look competitive along any dimension: https://artificialanalysis.ai/models/step-5#intelligence-com...

Better luck next time.

  • RussianCow 13 hours ago

    It's fast. The average speed is 115 tokens/sec according to OpenRouter. I haven't tested the model to see how it is in practice, but I'd certainly pay a little extra for faster inference.

    Edit: Though the average latency of 1.5s isn't very low, so it might not be that fast in practice for agentic work. Also, I don't know how much thinking it does, as that's generally been the drawback to Chinese models.

    • celrod 13 hours ago

      Using Artificial Analysis

      Model | Reasoning | Intelligence Index | Artificial Analysis million output tokens for the intelligence index -|-|-|- Step 5 | ? | 44 | 160 GLM 5.3 | Max | 45 | 210 MiMo V2.6 Pro | ? | 46 | 140 Kimi K3 | Max | 44 | 160 Qwen Max 0902 | ? | 45 | 190 DeepSeek 4.1 Flash | Max | 39 | 250 GLM 5.3-flash | Max | 42 | 180 GPT-6 Astra | Low | 46 | 10

      It looks reasonable by open model standards. This doesn't capture the fact that DeepSeek and Step 5 have much higher token/s than the rest, other than MiMo Ultraspeed. MiMo V2.6 is either slow but cheap, or fast but expensive. Just based on these numbers, it looks good. Astra-low is one of the fastest and cheapest because it doesn't use many tokens, but I've never tried it. I liked DeepSeek and GLM when I used them.

  • ssivark 2 hours ago

    Artificial Analysis has some very specific biases or perspectives on what they are measuring, so I wouldn't take their benchmarks as the final word on model quality.

    Composite indices are only useful for model companies which want to build one horizontal capability for N use cases, or naive users who don't want to bother with the effort of carefully pairing models with use cases. If you're a power user looking to understand and make deliberate choices, then you want pointed evaluations -- not general composite indices. You wouldn't hire the same person to do your taxes and mow your lawn, so why is it any different with LLMs? Only if you come at the problem with the folklore around "AGI" do you start making composite benchmarks.

    As I mentioned in a sibling comment... Back in their 3.5/3.7 era [1], Stepfun had a small+fast model, specifically eschewing knowledge, and instead focused on "general intelligence" like managing tool calls, orchestrating sub-agents, etc. Could be a very good substrate for running a personal assistant (time will tell what's the right approach); so I'm curious to try that out and see how they've come along since then.

    [1] https://news.ycombinator.com/item?id=47069179