manlymuppet 1 day ago

Am I hearing this right, that they made a decision model based on Typesafe's new paradigm, and actually made a model better than Jev based on Typesafe's own ranking?

And it's only been a few weeks.

  • TeMPOraL 1 day ago

    It's not a "new paradigm", it's a low-hanging fruit that's been lying around for years; Typesafe were the first to bother to stop and pick it up, and market the shit out of it. But it was still a low-hanging fruit.

    There are many, many of those left around, because AI frontier is moving forward so fast, everyone is racing ahead. Which is why I laugh when people say AI is not transformative and LLMs are a dead end (and my favorite, "what are we going to do with all those GPUs when the bubble pops?"). Even if SOTA LLMs hit a hard capability limit tomorrow and never advanced again, there's a good decade of growth and advancement to be extracted just from all the low-hanging fruits that were left unpicked along the way.

    • seizethecheese 1 day ago

      Name a few of these low hanging fruit left around.

      • murkt 1 day ago

        Easy to reach doesn’t automatically mean “easy to see”.

      • TeMPOraL 1 day ago

        Jev is one.

        Diffusion transformers are not "easy" but underfunded.

        Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware - opens up so many possibilities I'm probably unable to imagine half of them.

        E.g. Imagine spellcheck/predictive text (or code autocomplete) where the model is able to process a whole paragraph + surrounding application/system context in between keystrokes. Or an OS being able to reliably guess what you're doing in real-time, in between your UI interactions, and offer actually helpful contextual reactions.

        Or imagine finally funding some decent studies into exploring the models as computational artifacts - studying their latent spaces, how they form and how they model reality internally.

        Or imagine automated sliding doors that don't suck.

        --

        [0] - Or anything substantially better than BERT-level models used in Jev or that demo from the company doing inference ASICs, that has a chatbot online that does 14 kilotokens per second.

        • aeve890 1 day ago

          >Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware

          That's low hanging for you?

          • blurbleblurble 1 day ago

            It's likely quite close. There are so many papers proving concepts that would bring this, they just haven't been combined in production.

          • msdz 1 day ago

            Maybe they meant in the sense of “untapped potential”, because so far a lot of the focus has been on increasing model capabilities, not necessarily performance/power budget.

            • TeMPOraL 1 day ago

              Yes. Point is, it's untapped only because everyone is running in the race (even if out of curiosity), and there's just not enough people with means to tap into these side threads. For the past few years, there's been many interesting papers that circulated the industry, got recognized as worthwhile pursuits, and then dropped because running behind the Big Vendors had massively better ROI.

          • ekabod 1 day ago

            That's a high hanging fruit, not low.

          • guyomes 1 day ago

            If we throw in hardware dedicated to a specific LLM, it seems to be a rather low hanging fruit. Especially considering that this is already happening for vision models [1].

            [1]: "FPGA-based CNN Acceleration using Pattern-Aware Pruning" https://inria.hal.science/hal-04689673/document

            • mdp2021 1 day ago

              > hardware dedicated to a specific LLM

              That wording screams "Taalas". Which, importantly, is not the only player trying to abate the distance between data and arithmetics...

          • TeMPOraL 1 day ago

            Yes. It's well within realm of possibility, but so far wasn't pursued because the Big Vendors went all-in into capability growth (rightfully testing "the bitter lesson" to its limits) and got themselves stuck in an arms race, while everyone else is barely keeping up and/or starstruck with fascination, exploring what these models can do.

            This got everyone racing forward and right now there is not enough human attention left in the world to productionize this, or any of the other "side threads". When the race slows down, people will catch up, branch out, and loop back.

            • blurbleblurble 1 day ago

              Just like renewable energy and so many other things. Hyperconcentration of capital is really tragic. I hope things turn around.

              • TeMPOraL 1 day ago

                They will. That's the fallacy of the "S-curve" everyone likes to commit these days actually giving a positive outlook.

                Assuming it won't get to full RSI, the current approach will burn out - most likely economically. The race slows down, people branch out, look back, start picking up the "untapped potential"/low-hanging fruits, and you have new S-curves launching in place of the one that just tapered off (hence a fallacy - a stack of S-curves adds up to continuing exponential growth).

                In other words: it comes and goes. Hyperconcentrated capital will eventually deconcentrate.

          • mdp2021 1 day ago

            > That's low hanging for you

            An important part of the industry is studying that: it is built-up effort. Sooner or later, the fruits will be harvested. The targeted preparation has been there for years now.

          • Twirrim 1 day ago

            We already have examples of LLMs running 16k+ tokens a second using custom ASICs.

            It's down at the moment (Not sure if it'll return?) but Chat Jimmy[0] produced by Taalas[1] was powered by an ASIC running Llama 3.1 8B, and hitting 17,000 tokens/sec. It was amazing to use, you'd no sooner have hit enter than you had a full response back. I actually found its speed to be a problem for interactive stuff, as every answer got several paragraphs I'd then wade through, vs a model populating text closer to my reading speed.

            I appreciate, there are differences between an 8 billion parameter model and something GPT-4-ish, but we're currently in the middle of a race between a half dozen or so companies to produce the next best frontier model, which requires their infrastructure to be dynamic.

            We really don't always need newer better faster stronger models, there's quite a lot of room for "good enough" where getting 17kt/s at significantly lower power would be amazing.

            [0] https://chatjimmy.ai/ [1] https://taalas.com/

            • AshamedBadger56 1 day ago

              >I actually found its speed to be a problem for interactive stuff, as every answer got several paragraphs I'd then wade through, vs a model populating text closer to my reading speed.

              It would be interesting to pair the super fast model with a normal speed model. Have the super fast one do all the background research, code writing, etc. The normal model would just relay the needed info to you at a more reasonable pace.

              • drewstiff 11 hours ago

                Why do you even need a second model for that? There are many simpler ways to slow down the rate that the near-immediate response is pseudo-typed onto the screen.

        • flipping_beacon 1 day ago

          Definitely agree with edge computation, although inference extensively researched and funded if SOTA LLMs hit a dead end tomorrow,there is still a lot to explore and research in inference and edge computation

        • blurbleblurble 1 day ago

          Diffusion models combined with these new looping techniques are gonna change the whole conversation about efficiency. Imagine control net but in one or more conceptual latent spaces.

          But also harnesses and more generally new insights on "the control flow problem" could end up squeezing a ton of performance out of small models.

        • Amekedl 1 day ago

          yeah your reply, nobody can predict the future.

          Enough stuff can happen, software use itself might change, and that could really cause anything. "What will we do with all the gpus" might become a question if for a magnitude of tech and reasons leaked-opus-9 runs on a macbook m6 or 7

        • seizethecheese 1 day ago

          I commend you for actually answering, independent of what I think of the answers.

          • dominotw 1 day ago

            not a very good answer though

            • tomrod 18 hours ago

              Yes yes, most comments and responses fail to contribute to conversation (though I disagree with your particular assessment); it thus remains important to continue to share our own thoughts as we battle through thesis and antithesis to arrive at synthesis.

        • mdp2021 1 day ago

          There are "low hanging fruits" - easier to achieve goals -, and there are super-fruits, milestone-fruits.

          Among the most important ones:

          -- the long-known Problem of Transparency, applied to the apparent emergent intelligence in NNs. Why does it happen - in detail?

          -- then, a Theory of Apparent Intelligence through NNs. Transforming the results achieved into a Science. Which allows to do what we are doing - but in a lean and targeted way.

          -- then, a General Theory of Intelligence, that includes the above to go beyond current architectures and get those features of Intelligence we expect and still not have.

          The long-term direction we got into must lead to this.

          (You note a ponderant detail of the above when you note the importance of explaining the emergence of a World Model from a Language Model.)

          • TeMPOraL 1 day ago

            Those are the absolutely fascinating parts, and I sincerely hope AI won't get out of control before we're able to tackle some of these.

          • patcon 1 day ago

            If anyone is interested, following Dr Michael Levin's Thoughtforms.life podcast is the cutting edge of where all this previously fuzzy stuff is becoming more concrete. So long as you can tolerate distinguished scientists flailing about as they discuss consciousness and life and developmental biology (and other less-obviously living things, like algorithms) as involving "free lunches" and "ingressing patterns from the platonic realm" :)

      • hobofan 1 day ago

        Closely connected to decision models: A good library to do ranking based on pairwise ranking on multiple attributes. By using a decision model (especially one that can make decisions on multiple fields at the same time) this becomes a lot faster and more powerful. Could make for a pretty nice search reranker as well as prioritizer for many problems.

        Of course you can also do ranking one-off with a decision model, but this likely less stable, and by doing pairwise ranking you can also relatively quickly do incremental inserts to the list.

      • sarkarghya 1 day ago

        I can imagine advancements on making smaller models work together better instead of a generalized core. Imagine a community or city having a https://pirateface.co/ so that the shard of the model that you need can be streamed in with minimal latency with your box only holding the minimal version (say deepseek v4 flash as orchestrator) of the model that you use on day to day basis.

        We have overcome split brain problems before so this wont be our first

        • esseph 1 day ago

          You're putting a lot of trust in uncorrupted, untrusted, unknown models (potentially).

      • CamperBob2 1 day ago

        An example I like to use is: compare the quality and scope of games released with a brand-new console to the ones released for that console towards the end of its life, when everyone has learned how to take advantage of whatever weird, wacky hardware Sony invented for that console generation.

        There is still a lot we don't know about how to get the most out of existing LLM components from a speed or cognitive-performance perspective. People could easily spend the next decade studying and refining what's been built so far, even if no new, original approaches ever arrive.

        • randomNumber7 9 hours ago

          I never enjoyed playing ps4 when it is constantly louder than my vacuum cleaner.

      • sroussey 1 day ago

        Just look at all the model type on hugging face. LLMs are a small percentage.

    • btown 23 hours ago

      Synthetic data is key here! Compared to 5 years ago, we now have oracles that can generate massive data sets of perfectly labeled multimodal data, practically for free. The number of architectures that can benefit from that is innumerable, and far beyond just LLMs themselves. On top of this, LLMs can implement any architectural ideas you have, and write custom tools to manage training and evaluation.

      Whether or not LLMs can self-improve their frontier capabilities, they can absolutely create a wake for themselves that accelerates everything else that's training on their synthetic data. We'll see every architecture of the past 40 years suddenly show leaps and bounds.

    • dzonga 21 hours ago

      my take there's a lot of low hanging fruit in applying small models to knowledge economy workflows.

      then vision & robotics.

      while everyone's chasing the frontier.

    • outofpaper 15 hours ago

      Yup, one output token, and read the logprobs of the posible tokens that could have been thos 1st token. Plenty of systems already do this; Typesafe's fundraising and marketing just made it visible.

      Good to see interest broadening beyond "just extend thinking." More approaches in the toolbox means fewer problems get treated as nails.

  • slopnt 1 day ago

    They have to have decision models already in production. Part of their business is detecting bots, DDoSers and spammers.

    • alightsoul 1 day ago

      Yeah that's a decision tree, random Forest or some other machine learning classifier. They have existed for a long time

      • smallmancontrov 1 day ago

        I'm all for rebranding discriminative models as decision models, though.

        "Discriminative" always had pointlessly bad optics, but I knew it was over when I started seeing prominent machine learning researchers who p=100% knew better describe discriminative models as generative because that was the buzzword of the year. "Decision model" sells the value proposition much better and doesn't sound like an anti-woke crusade.

        • tomrod 18 hours ago

          Decision models have long been called `choice` models to may understanding. Huge statistical literature on dynamic discrete choice models and inference!

      • eastdakota 20 hours ago

        I think you just called me old.

  • segmondy 1 day ago

    A lot of people claim to have made better than jev, there's a jev benchmark, I have tried many of those models and they eventually end up failing, a non trivial task which doesn't seem like much but reminds me of the svg pelican bench is games, have one of these decision/classifier models play a game, hook it up to the input, most of the ones that are supposedly on jev level end up playing a terrible game, showing that they are very narrow. Cloudflare doesn't compare to the top open bench alternatives, I just finished downloading it and will compare it to jev for non trivial tasks tonight.

    • SebastianSosa 23 hours ago

      Public benchmarks are easy to cheat, if I am typesafe I would also release a public benchmark to distract otherwise competent people in overfitting to a benchmark instead of making something actually useful. Diogo very much is against public benchmarks ;)

      • Foobar8568 10 hours ago

        You take Qwen3.6 35b on a 5090rtx, and here you get a higher score than Jev, for 2sec more latency on average. So yeah it's not subsecond, but I am sure that if I had VC money, I could too get within 500ms too!

    • verdverm 23 hours ago

      watching Jev play Pokemon demonstrated this too, more hype than meat

      - I'd like a potion, are you sure, no, repeat

      - in and out of doors on loop

      - sisyphean effort in the cave

      - jev-ish level grinding

      It was impressive, beat pokemon for less than $2, but not all that interesting. People asking how different Math.Random plays pokemon would be, and at the other end, regular llms playing games.

    • indoor47 10 hours ago

      Well, it makes sense:

      "Clef builds upon this concept, but uses a different base model as the backbone. We currently use Qwen as the base model and post-trained it to suit decision model use cases. "

  • heliosAtwork 22 hours ago

    There was some parallel independent work from Sep 28. They mention it in the blog post. But I am sure Jev has opened a lot of eyes on the possibilities.

    "In the same week that Jev came out, we posted about some experiments [1] we had with our own homegrown decision model."

    [1] https://x.com/michellechen/status/2101091012559151480

  • lofaszvanitt 22 hours ago

    Race to the bottom. The one with the most resources wins.

  • fwip 9 hours ago

    From their blog post, it sounds like it is still 10x as slow as Jev.

    As far as I've seen, all of the Jev-compatible projects simply take an LLM, hack off a layer or two at the end, and call it good. Some of them spend more work than others trying to back-estimate in accurate probabilities.

    • rahimnathwani 5 hours ago

      Yeah I think people were already using either of these as part of traditional software workflows):

      A) Structured outputs from LLMs (doesn't need fine tuning but can be expensive)

      B) Classification output from fine-tuned BERT-like or GLiNER models (is calibrated well and has cheap/fast inference)

      What Jev did is combine the advantages of both A and B into one model/product, and create a really good API.

      They claim that a key innovation is how they've trained the model using what they call RLCD (RL from calibrated decisions). So it's not just that you can get the outputs (which is easy to add to any LLM) but that the different primitives they expose (Choice, Score, Noul) have each been calibrated. For example, they claim that if you use the Score primitive (which gives you probabilities along a bunch of choices representing a continuum) that's not just using the more general 'Choice' primitive under the hood. It's been calibrated separately.

      I don't know how many of the Jev-like things we've seen do that. For example Cloudflare offers a Jev-like model with the same API, and which they say was trained with RLCD. But I don't know whether Choice and Score are different under the hood, or whether Score is just sugar on top of Choice. (Should be easy to test this, but I haven't done it.)

  • nater5000 8 hours ago

    >And it's only been a few weeks.

    You make it sound like this is some noteworthy timeframe. I'd be more surprised if Jev wasn't immediately made obsolete within days of release (which was effectively the case), especially when backed by a company like Cloudflare lol

    ML moves quick. Add the extreme hype and cash floating around in this space and you can expect that anything resembling something novel and relatively untapped is going to be pounced on and turned over basically immediately.

  • xnx 4 hours ago

    What's the hype with Jev? Hasn't Gemini had structured outputs since November 2025? https://ai.google.dev/gemini-api/docs/structured-output

    • manlymuppet 2 hours ago

      Traditional LLMs doing structured output is like trying to fit a square peg in a round hole. It's just not the right tool for the job. They can do structured output, but it's janky.

      With Jev, structured output is its native format. Jev is just way faster, cheaper, and outright better for a lot of things.

      Note that while all of this is great, this is nowhere near a sort of "ChatGPT moment". It's a cool new thing, and it's way better at certain tasks, which could be big. That's all though.

djray 7 hours ago

Unrelated to the subject matter, but something that I never fail to appreciate is just how well-written the majority of Cloudflare's posts are. (There was an article on Merkle Tree Certificates which was phenomenal.) There is real skill in creating an informative work which condenses complex technical subject matter into something that a technically-inclined non-expert can understand.

buildbuildbuild 1 day ago

Open weights, not open source.

The weights have permissive licensing, but the data and training pipeline are not published to reproduce them from their proprietary Qwen starting points. Weights are not "source."

  • jMyles 1 day ago

    Came directly to comments hoping not to see this one.

    <sad trombone sound>

    Surely someone will soon do what the title of this post makes it seem like cloudfare did. Truly modular open source training and inference logic, along with a totally open corpus and weights, will eventually out-compete the closed ecosystem.

    • ainch 1 day ago

      There are some groups doing it for LLMs - like the Allen Institute for AI's Olmo models and Eluether AI's Pythia.

  • dang 22 hours ago

    Ok, we've put weights instead of source in the title, at least until someone comes along and says that's not right either!

vulture916 1 day ago

Jev = $0.042/m input, output free Clef = $0.24/m input, no output price listed

At 300 tokens per call, you'd get:

One million decisions on Jev cost about $12.60. One million decisions on Clef cost about $72.

Would probably make sense to self-host Clef, if you have the capability/resources. If not...

  • jampekka 1 day ago

    Output for decisions have so few output items (not really tokens here) that they are negligible anyway. Jev hyping "free output" is almost lying by omission.

    • verdverm 23 hours ago

      framing it as "free output" is a decision /s

  • scronkfinkle 1 day ago

    > no output price listed

    It's weird to think of these kinds of models as having "output tokens". Cross-encoder approaches like Laya add a [MASK] marker per option, but nothing is generated the way an autoregressive transformer generates. It's one bidirectional pass over your input, then a small head scores each option, so you wouldn't really pay for output as much as only input

  • strangescript 23 hours ago

    Because Clef is a larger model, larger context, can process images. Plus its open weights, so you can run it yourself.

    Since it benches better than Jev, Jev is probably smaller and easier to host. They could also be losing a lot of money.

    • wongarsu 13 hours ago

      Considering Jev is named after Jevon's Paradox (making things cheaper increases spending) I'm willing to bet that it's the former. Smaller model, possibly an inference pipeline optimized for a Jev-like workload (large shared context that is used for multiple questions in parallel)

      And while Jev has a lot of latency, that latency stays very flat with larger inputs. So it's probably related to their inference pipeline, not model size

      • andrewingram 10 hours ago

        It's anecdata, but I ported a Jev-based prototype i've been working on to Clef, and it ended up being significantly slower and the results were worse.

  • Havoc 14 hours ago

    >Would probably make sense to self-host Clef

    For privacy perhaps but on pricing you're unlikely to come out ahead versus datacentres with scale and industrial power pricing

    • otabdeveloper4 12 hours ago

      That's what they originally said about AWS too.

      Turns out AWS is actually about 10 times more expensive than renting your own datacenter compute.

      • Havoc 8 hours ago

        You’re not really paying for just the compute in AWS though. You’re paying for compute that comes with 200 integrated ready to go other products and stuff like compliance and certifications.

        It’s like comparing pricing of a ribeye at the butcher and on restaurant menu. It’s apples and orange.

        Certainly AWS is making good money too though. And yeah if you just need a VM then AWS isn’t the way

dgacmu 8 hours ago

Since it supports vision, I tested it (an 8 bit quant) against one of my pet problems of coin classification (US coins, but not beauty shots of them, just lots of cell phone images on a piece of paper). It did .. horribly. About 41% accuracy. My local gemma4:27b gets 53% and runs 4x faster. Not too surprising but seemed worth poking at it.

  • leopoldj 8 hours ago

    The model supports both instruction and RL fine-tuning. It looks like people have started training it already [1]. The RL tuning may be only available to the Cloudflare customers, I couldn't be sure.

    1. https://huggingface.co/solanaclawd/clef-solana-research

    • ActivePattern 7 hours ago

      The biggest upside to using a zero-shot classification model is that it's actually zero-shot, i.e., you don't need to construct task-specific training sets.

      If it needs to be fine-tuned, then why not fine-tune smaller and cheaper models that are only as big as the task requires?

      • leopoldj 7 hours ago

        I hear what you're saying. But this is not such a model. The article says that various internal teams had to fine-tune Clef to suit their specific needs. They're even rolling out a RL training product line for it.

        • dgacmu 7 hours ago

          Fair. Though in this case my actual classifier (dino v3 features fed into a very shallow DNN) works surprisingly well and is insanely faster than using an LLM -- so it really was the one shot performance I was hoping for.

agrippanux 20 hours ago

I'm a big fan of Cloudflare products. I recently stuck Jev in front of a Cloudflare-hosted Ollama model for chat/username moderation, so I was excited to test out Clef. The setup is user send a message -> Jev does first pass to see if it's toxic/hate speech/profane, if Jev is unsure then Ollama on Workers AI takes a deeper look.

Clef was 2-3x slower and worse (it caught less hate speech) than Jev. Overall disappointing.

  • teleforce 19 hours ago

    In "How we trained Clef" section Cloudflare mentioned that they initially based their Jev-like system on DiffussionGemma (DJev) model but then changed to Qwen with post training for Clef, but never mentioned any reason and justification for the change.

    Although they included the benchmark against DJev and Clef is better, perhaps if you can test it to see the real-world performance.

    • nostrebored 18 hours ago

      Diffusion Gemma sucks to train if you are not already mostly in distribution. The reason diffusion Gemma is fast is a learned denoising that balances quality and speed. So the further your data is from what already happens, the slower it gets. In almost every case I’ve tried, you might as well just train Nemotron.

  • wongarsu 13 hours ago

    It's likely no coindicence that they only show "median latency", not how latency scales with input size.

    For small inputs Jev is slow, but its latency curve is very flat. Fine-tuning a decently-sized llm (like this 27B model) gives you something that's faster on small input sizes, but even with moderate contexts quickly becomes much slower than Jev. Characterizing it as "faster than Jev" is very misleading, unless you know all your questions have tiny context (less than 1k tokens or so)

    • tomrod 10 hours ago

      True. But do I understand that Clef is multimodal and Jev is only text?

      • wongarsu 8 hours ago

        yes. Jev is text-only, while Clef and a lot of the other alternatives are fine-tunes of multi-modal models, so you get image-input basically for free. Actual decision quality based on images is a bit up in the air though, I am not aware of any benchmarks testing that

ssiddharth 1 day ago

Pricing is $0.24/million input tokens which is ~6x compared to Jev. Clef-flash is at $0.09 which is way more competitive.

  • CBLT 1 day ago

    Yeah I also thought it was strange their pareto frontier didn't include cost.

bityard 1 day ago

Clef is based on Qwen3.8-27B and Clef-flash is based on Qwen3.8-9B (edit: actually Qwen3.5-9B). So, similar in spirit to Kev by my understanding, but based on a newer model.

  • NitpickLawyer 1 day ago

    > and Clef-flash is based on Qwen3.8-9B

    There is no official qwen 3.8 9b

    From the model card:

    > Clef-Flash is post-trained from Qwen/Qwen3.5-9B. See Clef for the larger variant.

    • ddarolfi 1 day ago

      It's based on Qwen3.5-9B, maybe a typo

      • verdverm 23 hours ago

        I've been guilty of the same wishful projection

    • bityard 1 day ago

      Thanks, I missed that. Fixed my comment.

  • okpatil 1 day ago

    Atom is 60M Param (around 133x to 400x smaller).

    16ms latency. And locally run.

    https://at0m.pienomial.com/

    Why go big when you can go small ?

    • kamranjon 1 day ago

      Cause it's not open?

      • okpatil 1 day ago

        Good point.

        To counter, most of the AI is not open. So is none of Microsoft Products. As long as they work, we keep using them.

        • verdverm 23 hours ago

          counter point, I've stopped using all closed models and harnesses as a life choice

          Ai is too important and transformational to let Big Ai dominate in a closed ecosystem, thankfully the Chinese have a different mindset and approach

          • okpatil 6 hours ago

            That's an absolutely fair way. More power to you !

    • ricardobeat 13 hours ago

      Because with a tiny model you're skipping all the intelligence and world knowledge that makes it useful without fine tuning. `typed-decisions` is almost entirely text classification tasks.

      • okpatil 6 hours ago

        Don't assume. It has generic world knowledge. It is performing well on the benchmarks we didn't even train it on.

iugtmkbdfil834 7 hours ago

I am playing with it now trying to make it play adom ( mostly to avoid issues with timing; using local variants so speed is a consideration ). By itself, it plays in ways you can likely imagine.. not well. Adding big Qwen on top for strategic decisions helps, but makes it fail in unexpected ways. Will be adding smaller non-qwen in the middle. WIP so actual details are changing fast.

mrkn1 1 day ago

For smaller scale decision model that runs on CPU, check https://news.ycombinator.com/item?id=49923223

  • fastball 1 day ago

    tbh saying all these dumb decision models are similar to Jev is like saying markov chains weren't far from GPT-2.

    The value isn't really in the I/O shape, it is in the intelligence combined with the output shape. Every extra ounce of intelligence in these models unlocks additional use-cases. But the converse is also true: a dumb decision model is going to be less useful than using a more intelligent standard LLM.

    That is the appeal of Jev: for certain usage it has more intelligence than some small SOTA LLMs. It is the first decision model that actually feels intelligent (to me).

    • mrkn1 1 day ago

      That makes sense, I agree. I still think there might be use cases when people don't want to use the Jev API, and end on a different trade-off.

yipinwong 1 day ago

A question someone not trainined in AI/ML field, Is a decision model that easy to crete that there are floods of these JEV alternatives already?

Or are companies/people already building this based on say an arXiv docs? n

---

The pricing is ... hm more expensive but not at the point I won't give it a try due to the embeded vision encoding

  • XCSme 1 day ago

    You can make a basic one in minutes based on existing open-source models.

    Latency won't be that good, but could still work similarly. Simply force the structured output of a LLM to the given schema.

    Probably also easy to train because we can use stronget LLMs to generate input/output data, or even synthetic data is easy to generate.

    It's not really a new technology, it's more like a new use-case.

    • sigbottle 1 day ago

      What even are these new "decision models?" Take an existing LLM, feed it a prompt, force it to pick a choice; decode is 1 token (or rather, the whole logit set for only that last token; token implies selecting one logit) so you made a choice. That's it?

      • orbital-decay 1 day ago

        Yes but optimized specifically for the purpose. Using that for "decision making" is also not a new use case, but turned out to be new to many people. Which is great, I hope they make something cool with it!

      • redox99 1 day ago

        Yes, although you probably want to calibrate your model if you want the probabilities to actually be meaningful.

  • redox99 1 day ago

    Yes it's very easy if you have fairly basic ML knowledge.

  • nico 1 day ago

    The basics are pretty simple. And depending on what your specific need is, the model can be really really basic, fast and super effective (ie. run on a mobile device and process thousands of requests in <100ms)

    I've been playing with this for the last year or so. Started with a personal email classifier, also did benchmarks with some public datasets, then created a couple classifiers that could play Doom, and now I've been trying out some other experiments, like a request proxy/router to automatically choose a classifier and fallback to LLM to handle unseen requests

    Jev did a great job at creating hype, but also at shaping the concept and space of "decision engine" or "decision model". People were already doing this with LLMs, which is very inefficient for most tasks like that, and the Jev guys figured there was a market there. It seems like they were right, and now there's a rush to flood the space, taking advantage of the hype window

  • calebkaiser 1 day ago

    There is a bunch of stuff to tease apart.

    In general, training a general purpose classifier is something lots of people have worked on for a long time. Large Transformer models themselves are typically "generalists" already, so structured generation and constrained decoding have given you the ability to use an LLM as a general classifier for years. It's an incredibly common pattern for working with LLM judges or any sort of branched decision making workflow.

    A lot of people who are a bit less familiar with the field saw the hype around Jev and presumed that the reason it was so exciting was that it was a fundamentally new interface for working with an LLM. And that additional excitement drove even more attention to Jev. But fundamentally, TypeSafe's announcement was that they found a particular architecture/training paradigm that resulted in a model for this particular interface that had incredible accuracy, very low latency, and for which they could offer inference at a super low cost.

    I've not kept up with the flood of Jev clones that have been released, but I think this is just typical for any new component in deep learning that gets popular. There are an absurd number of open source autoregressive LLMs and fine tunes you can use. The thing that makes one more popular than the other is typically the general performance of the individual model.

    But training a model for this purpose, or emulating the procedures described in Jev's papers, isn't something that would be beyond the capabilities of any lab. It's not an entirely alien architecture or approach.

    The bigger question for TypeSafe as a company would be if other teams are producing Jev-like models that win on performance or cost. Like I said, I haven't followed the reports super closely, so no idea if that's the case or not.

    • tomrod 1 day ago

      > The bigger question for TypeSafe as a company would be if other teams are producing Jev-like models that win on performance or cost. Like I said, I haven't followed the reports super closely, so no idea if that's the case or not.

      If CF's benchmark is representative and sufficient, Clef outperforms Jev!

      Models by themselves don't guarantee market capture. Rather, its how they integrate. I think a lot of folks are burned by the closed nature of many models.

  • TeMPOraL 1 day ago

    Yes, it's easy. The thing people are missing (especially those believing AI is a "dead end" and "not transformative") is that the field has been advancing so fast in the past few years, that there's lots of such unexplored avenues, unpicked low-hanging fruits, that everyone just raced past. We've barely begun exploring the capabilities ML brought us - patterns, applications, and architectures.

    Now that we're hitting against the hardware supply limits of global economy, I expect more people to go back and revisit the things left along the way in the mad rush to "just throw more compute at it / make a bigger model" - and thus many more cases like Jev to show up in the next few years.

  • orbital-decay 1 day ago

    Yes it's easy for an established shop, all they need to do is to tweak the post-training workflow. "Decision model" is the same kind of marketing as "LRM" attempted by OpenAI when RL CoT was new (to hyped up crowd). It's still fundamentally a classifier used for "decision making", games and RP were using generalist models and constrained outputs to do what the DOOM demo does for years.

  • janalsncm 1 day ago

    The interesting part is also the easy part. The model and architecture are not hard for an experienced machine learning engineer to build.

    The hard part is the data and evaluation. Sure, it’s not that hard to build a fast model with good predictive power. But fast at doing what? You probably don’t care about classifying whether a hotdog is a sandwich (which is the Jev demo).

open592 1 day ago

2 years in stealth...

fooker 1 day ago

This is awesome.

I bet the competition will result in research into how to make these decision models several more orders of magnitude faster and cheaper.

Here's a challenge problem - look at a 1M context window and produce N decisions (different queries) from it in 50-100ms.

amluto 1 day ago

I’ll go out on a limb and suggest that I don’t think a Jev-like model is particularly useful unless you can fine tune it. The Jev API has zero ability to pass in a prior [0], and, if you can neither pass in a prior nor fine tune for your system, you will get an output that may be almost meaningless.

I’d love to see someone build a model of this sort that can actually accept priors and do something intelligent with them.

[0] You can feed Jev a prior as text. I’ve tried it. It works poorly.

  • brokensegue 1 day ago

    I think better than priors would be a closed loop where you tell it what the right answer was (or some signal) and they monitor and fine-tune for you

  • okpatil 1 day ago

    We were able to completely automate 20,100 token prompts with At0m[https://at0m.pienomial.com/].

    We believe entire compliance workflows (even multilingual) could be automated.

    Would you like to get a demo ?

    • sheepscreek 1 day ago

      You’re coming on a bit strongly - a couple of comments with a link is sufficient. Before trying to sell, try to genuinely further the conversation, provide some useful knowledge in return for the reader’s attention.

      • okpatil 1 day ago

        Point taken. Let me explain if you allow me.

        It is possible with deterministic decision models, such as At0m, to gauge the probabilities at every decision. This behavior in addition to hard coded logic, it is possible to completely replicate a prompt's logic.

        Using Fable 5.1, it is a matter of minutes.

        I believe that most of the compliance check documents will be a solved problem, 3-6 months in future.

        None of the LLMs can do it.

        Hence I asked to the comment poster if he would want to demo, so that I can show it to him, how to do it step by step. By bad, if it came out too strongly.

  • sheepscreek 1 day ago

    Also one of the more interesting features of Jev is the confidence rating that hardly any Jev-cc talks about.

    • amluto 1 day ago

      It seems interesting to me only in the sense of being useless. From the horse’s mouth:

      > Confidence is derived from the probabilities

      https://docs.typesafe.ai/confidence

      (Why is it much easier to find AI-slop websites quoting this than it is to find the actual documentation?)

      My inner Bayesian would like for Jev to provide something resembling “evidence”, although I admit that one might ask Jev questions that are somewhat awkward to treat as typical Bayesian questions. If I ask “will this PR be merged”, it’s kind of strange to contemplate the probability of a PR conditioned in that PR being merged in the future. But I bet there is a way to formalize a prior-free classifier in a way that makes Bayesians and non-Bayesians happy, possibly involving actual learned probabilities and confidence levels. If you read the literature on scoring rules, you will find that classifier scores do somewhat naturally decompose into a few interpretable terms.

  • mikeocool 1 day ago

    It seems like jev's major advantage over existing classifiers is that I dont have train it.

    If I have to gather and tag data to fine-tune Jev, I can probably just train an "old school" classifier model and make it even cheaper, faster, and just as accurate.

  • jdthedisciple 1 day ago

    I suppose a sort of prior-proxy can be encapsulated by a carefully written system prompt.

    • amluto 18 hours ago

      The straightforward approach of literally staying a numeric prior has some effect but not the correct effect.

  • AnthusAI 23 hours ago

    You can't fine-tune Jev itself but you can train an ML model that uses outputs from Jev as inputs. Which does enable you to 'fine-tune' your overall model.

    You can also improve your Jev classifications based on your ongoing data if you're labeling it continuously, especially if you're explaining the reasoning in the feedback labels. You can identify new elements of the rubric and add them to the list of classifications that Jev produces, and then those become new features for your ML model.

    Two levers of control for using data to make a Jev-based classifier model continuously better-aligned.

ricardobeat 13 hours ago

In my experiments decider-4B performs better than Kev with significantly lower latency. It's remarkably good for it's size, shame it wasn't included in the benchmarks. Laya on the other hand shouldn't even be featured - despite being 'the original' decision model, it can only do simple text classification and is nowhere near usable performance for anything else.

johnbatch 18 hours ago

“we’ve been testing our new Clef model on our Threat Intelligence team to help us classify website domains” Not sure if any of this testing is in production but their website classifications need lots of work. I use the cloudflare ZTNA agent on my work laptop and often run into uncategorized websites. Just last night I could not get into my insurance company’s website because it was uncategorized so I got a block page. If Clef can decide this in 2 seconds I shouldn’t have to enter a support request and wait hours for it to get categorized. Especially the main page for a Fortune 500 company.

ksymph 1 day ago

With all these new Jev-like models popping up, has anyone actually started building anything with them yet? It's odd how quickly they've multiplied despite being relatively niche in their use cases, as far as I can tell. I suppose they're simple and cheap enough to make that it's a sort of 'why not' thing for a lot of these companies.

  • reassess_blind 20 hours ago

    I'm trialling using it to scan user signups for signs of gambling spam, phishing etc. My current process uses an LLM, but seems like a good fit for these classifiers.

meander_water 1 day ago

> This contrasts with the world of Large Language Models (LLMs), which are largely non-deterministic, but are open-ended enough to reason...

This seems misleading. Decision models do not produce deterministic output. Repeated calls can product different decisions just like an LLM with structured outputs.

  • lmc 10 hours ago

    Also, insinuating it's not related to an LLM is bonkers. Does anyone that made this actually know what they're doing?

cakoose 1 day ago

> This means that a human does not necessarily need to be in the loop for agentic decisions anymore — agents can programmatically gather context, make decisions, and take actions on tasks, or defer to a human when needed.

1. Humans are already not in the loop for lots of LLM agent actions. Isn't that just a function of how much you trust it and not some completely new paradigm? Am I missing something?

2. How can it gather context if it just outputs a single decision?

One guess: Maybe it's decision can be "gather more context and re-run me"? But an LLM can be much more expressive about what context it needs.

  • lmc 11 hours ago

    Yeah reading that made me question the credibility of the whole thing.

pdlug 21 hours ago

I love what Cloudflare is doing generally so I was excited to try Clef in my evals on a real task vs Jev: should an agent's knowledge-base write go to human review?

Quality: close (recall 0.98 vs 1.00) Hosted p50: Clef ~850ms, Jev ~110ms Clef-flash: over-escalates

Data + script: https://github.com/nicia-ai/admission-decision-eval

sauercrowd 7 hours ago

I'm a bit confused, why is clef-flash outperforming clef in quite a few of the benchmarks?

ranyume 1 day ago

I found the paragraph about how much networking data they have weird. I mean, if you already have all that data why didn't you train your models already using that? Why did you need clef to begin with?

Jeeetendra 11 hours ago

small decision models make way more sense than calling a giant llm for every yes/no call. the fine-tuning part is the interesting bit

croemer 1 day ago

Is there a Jev-like model I can run on my Mac? Something like Ollama? Or what's the best way to play with it? Is there a cheap/free API service eg on OpenRouter?

Wazzymandias 20 hours ago

It's amusing that this blog post explained Jev's underlying design far more clearly than their onslaught of bull posts and marketing

This entire time they could have just said "decision model" but they kept using vague, flowery wording. I have no idea why

tough 12 hours ago

Interesting how this is downstream of their acquisition of Replicate

jasfi 1 day ago

Related: an intelligence cache for decision model data: https://cachev.dev

I built this for my own needs, and thought others might find it useful too.

6thbit 1 day ago

I wonder if a good usecase for this would be cloudflare's WAF rules. Give broader request context to the decider and let it pick type of challenge/block traffic directly.

Perhaps that may be too costly atm

warkdarrior 1 day ago

Can someone explain how so many folks managed to build decision models within days or weeks after Typesafe came out with Jev? Is this concept of decision models been in the works for a while? Is it easy to copy?

  • didibus 1 day ago

    You can use already trained large transformer models to make one, so it doesn't require the kind of high-scale compute, high quality data, data cleanup, reinforcement, and so on training that say an LLM does.

  • petercooper 1 day ago

    Smaller models have been able to do these sorts of tasks, but a little slower, for a while now. Give a small Qwen 3.8 model a classification task and force a structured output, and it'll do a good job. I've used Qwen 0.8b for basic image classification in <500ms on my local machine for a while now.

    There are a few technical details that can reduce the latency significantly (covered in the post) but the real insight has been from watching the reaction to Jev and seeing that there's enough of a market interest to offer it as a distinct thing. The underlying concept/approach was already there.

    • theapadayo 1 day ago

      Not just structured output. Dropping down to logprobs, prompting the model to emit one word as the answer, and then ranking the output tokens to pick your answer works great on small Qwen & Gemma models.

      The fascinating part to me is that Jev seems like this technique plus post-training to get multiple independent confidence values for each possible answer.

  • woah 1 day ago

    Transformers output a set of probabilities over outputs. For ChatGPT etc, those are predictions of what the next token will be. But it can also be a structured list of options or classes. Jev mostly innovated on the interface, API, and product concept around this, and made it click for a large number of people. Unfortunately for Jev, it's very easy to copy an API, and any pretrained LLM can be adapted to work in this way.

    • ford 1 day ago

      I think Jev also innovated on data & algorithms, but it remains to be seen if it's enough to be meaningfully better than traditional LLMs + a few tweaks.

  • zitterbewegung 1 day ago

    You just have to fine tune an LLM like Qwen on some synthetic data to do so. There was even someone that had a model that was exactly like Typesafe and published their work a year before Jev (but wasn't marketed as heavily since it was academic).

  • ramoz 1 day ago

    Jev created accessible/programmatic ergonomics around a general purpose classifiers, and did it very well; ie intuitive api and structured data approach.

    Anyone can copy that and apply to an array of models - stripped down LLMs or already slim/highly performant traditional classification architectures (just wrap inference with an api that inputs/outputs the same structured data).

    Jev, I think, would say their advantage is the intelligence of their models and training data including calibration: https://medium.com/code-applied/calibrated-classifiers-makin... (which i still struggle with in the general application... there's no free lunch with these things).

  • 233mhz 1 day ago

    What's new is "smart" decision models than you can supposedly use on anything without additional training.

    If you have a very narrow use case you can train a BERT based decision model on a laptop an hour if you have good data to train it on. It'll answer faster than the roundtrip to clef/jev and use <1gb memory

    • conmod278 1 day ago

      If you have a very intelligent swiss army knife like hammer, that hammer will adapt to almost any nail, which is a good thing.

  • porridgeraisin 1 day ago

    They are not too difficult to train if you already have infra to train regular LLMs. You can typically replace a few layers train them alone and you're off to the races.

    Getting training data that works well for calibrated classification objectives is difficult.

    I hear conflicting opinions (including my own) about how well calibrated each of these are. Jev seems to be the best.

    But the jev release made obvious the PMF for these models, and the underlying reality is that calibration really doesn't matter much when you're replacing usecases where people were using damn LM head softmax probabilities before, which are nowhere near calibrated.

    So now everyone simply finetunes qwen and makes a compared-to-regular-LLM vastly cheaper decision model. And it works for majority of usecases. People mostly only care about accuracy, not confidence.

  • pizzafeelsright 1 day ago

    The question of AI in automation is "can it make decisions in a consistent and predictable manner, with near 100% determinism?"

    Many people seem to have run into the same question and started working out the answer.

  • giancarlostoro 1 day ago

    It's not a new concept, it just took someone adding on to the approach and refining it. I never deep dove it, but I assume JEV is sort of like how Sora works? They had a blog post about how it has a sort of tiny LLM, which OpenAI's small LLMs are insanely good and well defined. I think any lab tackling this with a from-scratch model could yield affordable alternatives that are highly competitive.

    It seems insanely obvious at least to me, that JEV is the new hot thing for the AI field since they give you stronger output that isn't... flat out wrong, that alone is impressive.

  • nico 1 day ago

    Most answers explain the LLM-based approach to these models, which is also what Typesafe did with Jev. However, depending on what you need, there are far simpler classification models, and for a lot of use cases, these models can be way faster and more accurate than Jev

    But, for these adhoc models, you need to understand the task more, collect some data and train the model (on CPU, no need for GPU). So Jev-like models are a great way of getting a hosted general decision model, but if you have a very narrow task or set of tasks, you might be better off with some more basic models that you can run on the same server you run other things or even on your laptop

aryabakh 1 day ago

it's great to see Cloudflare releasing consumer edge level models.

radium3d 22 hours ago

Can these decision models be used to enhance the intelligence of video game adversary and friendly NPC "AI"?

alex7o 1 day ago

Oldy enough I tired this 2h ago as I was testing jev on cf and was like oh this should be a better replacement but it takes 3s which is useless to me

pcthrowaway 9 hours ago

> First, it has a vision encoder so it’s able to take in images and classify visual content.

Wait, does this mean this is the first (non-generative) model to be able to do https://xkcd.com/1425/ !?

damsta 1 day ago

Competition in this area is great and kudos for releasing something that we can try out today.

schainks 1 day ago

AMAZING, thanks, Cloudflare!

swingboy 1 day ago

It allows image input. Nice!

  • ttul 1 day ago

    That's probably driven by their own internal need to show the model images of emails and webpages to detect phishing, despite obfuscation of the underlying HTML.

nikcub 1 day ago

in a quick mini-bench here n=250 of clef vs jev, clef came out 5.2x more expensive, a lot slower (p50 of 350ms vs 1.9s) with only marginally better results (78.6% vs 79.8%)

_superposition_ 1 day ago

Commodities get commoditized. Nothing to see here.

DesaiAshu 1 day ago

brb while I build my entire cloud stack on Cloudflare

swe_dima 1 day ago

Would love to see benchmarks on visual tasks.

gitghxst 1 day ago

that's good. I believe decision focused models will explode in the next few months

wakeywakeywakey 1 day ago

jev can ask jev if they should return investor money and shut down jev

ralusek 1 day ago

Just tested clef-flash vs jev:

- Jev/TypeSafe: 230 ms median, 254 ms mean

- Jev/OpenRouter: 237 ms median, 267 ms mean

- Clef Flash: 661 ms median, 806 ms mean

What gives?

reexpressionist 19 hours ago

> "This means that a human does not necessarily need to be in the loop for agentic decisions anymore"

That's only true in a practical sense if you can actually rely on the probabilities estimated by the model. That's a non-trivial problem for multiple reasons, among them: 1. What is the particular quantity you seek to estimate (marginal, approximately conditional, etc.)? 2. What is the reference class for that quantity? 3. What method are you going to use to estimate that quantity? 4. What is the error in your method to estimate that quantity (e.g., as via accounting for the effective sample size)?

A further practical challenge is that the output logits of neural networks are in effect a highly lossy compression of the epistemic (reducible) uncertainty. Even if your estimates are well-calibrated (for some definition of well-calibrated) on in-distribution data using the output logits, those estimates can be grossly uncalibrated in the presence of covariate shifts, and the logits themselves are not reliable signals of such shifts, nor of being out-of-distribution. Informally, the output logits themselves do not encode a good sense of what they do[n't] know.

Additionally, ideally the probability estimates are interpretable in the sense that there is some instance-wise connection to the training/calibration data. If the estimates are being used for decision-making, you need to be able to post-hoc audit the estimates to be able to modify the data for future decision-making, if needed.

Growing evidence in ML/NLP/Stats from the last few years is that with neural networks, as a starting point for constructing reliable estimates of the predictive uncertainty, we need to control for metric-learner signals over the support/training set (e.g., the L^2 distance to the nearest training instance and depth-matches into training). Once you have that, then you can choose your desired quantity of interest (e.g., class- and prediction-conditional accuracy at least some given value). Concretely, here's a tutorial (along with Apache-2.0 code) that steps through a simple, illustrative example: https://reexpressai.github.io/reexpress_sdm/tutorials/gettin...

More context is in the link, but at a high-level from an engineering perspective, just as dense vector matching is used by RAG for information retrieval, we can also use dense vector matching in this way to estimate the predictive uncertainty, getting around the limitations of the output logits.

MisterMunchkin 1 day ago

Imagine making your whole company on one model and then being cucked by everyone within a week. I don't think I've ever seen anything like it.

  • RGS1811 1 day ago

    If everyone else can spin up their own version of your product in under a month, there probably wasn't much product there.

  • hansonkd 1 day ago

    Yeah, these AI companies have some weird paradox that if they actually had a model that was super efficient and could arbitrage cost/intelligence of other inferior models, they would keep everything about it secret. If an intelligence research group had something groundbreaking, they would just dump their own money into the magic money machine.

    Instead, to make up for the lack of economic viability of their models, they are forced to release publicly to get marketing to get others to pay based on hype.

    • globular-toast 1 day ago

      If they were actually useful they'd just make money doing the useful thing and wouldn't even talk about models or AI.

AmazingTurtle 15 hours ago

quick reminder: this is a one shot pydantic token guess. if you try to use a model that was trained on reasoning? you're not getting any of the reasoning.

zwaps 1 day ago

No mention of calibration. Is it just another llm finetune?

  • kflansburg 1 day ago

    > Our post-training utilizes label-smoothed cross-entropy for valid schema outputs paired with a Brier loss to refine probability calibration.

winddude 1 day ago

Now we just need some new NER models.

esafak 1 day ago

I feel bad for the Jev guys. I wonder if they anticipated this much competition?

hbcdbff 1 day ago

“Urgency” of “yes”?

  • alashow 9 hours ago

    It's "Is it urgent?" Not "Urgency"

mococa 1 day ago

I bet this's 100% slop.

okpatil 1 day ago

Why give cloudflare your data ?

At0M: A 60M local Jev at 16 ms latency and 79% accuracy on Typed Decision

https://at0m.pienomial.com/ https://news.ycombinator.com/item?id=49920350

  • dcastm 1 day ago

    What’s the context window?

  • cootsnuck 1 day ago

    If it's 60M local, is it open source as well?

    • okpatil 1 day ago

      It is a single rust executable, with model embedded inside of it. Current evaluation API is being run by the same.

      We wanted to stress test the system before the V1 release.

      • Ciph 20 hours ago

        It runs locally, but will weights be available?

        • okpatil 19 hours ago

          Of course yes. They will be inside the executable.

  • kamranjon 1 day ago

    This account was created only a couple weeks ago and seems to be spamming this closed model, seems it might be a bot?

    • okpatil 1 day ago

      Actually it was created around 12 years ago. We just launched the model around 3 days ago.

      Apologies if it is too much of a bother.

  • Transformanshen 1 day ago

    It hasn't been released yet, and local versions aren't always convenient

    • okpatil 1 day ago

      The same engine is is available as an evaluation API.

      curl -s -X POST https://at0m.pienomial.com/decide/v0 \ -H 'Content-Type: application/json' \ -d '{ "state": "Charged twice for the same card payment this morning.", "questions": { "queue": {"type":"choice", "instructions":"Which team should handle this?", "criteria": {"billing":"invoices, charges, refunds", "technical":"outages, bugs, deploys", "fraud":"unauthorised or suspicious activity"}}, "urgent": {"type":"noul", "instructions":"Needs action today."}}, "email_id": "you@example.com"}'

      If it fits your use case, you are welcome to use it.

      When it is a rust standalone rust executable, as it is powering the API, it becomes just plug and play. No dependencies needed.

johnecheck 1 day ago

Wow, Cloudflare is definitely buying some goodwill from me. Just consistently interesting new releases alongside and solid products at great prices. Seems nearly too good to be true.