Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
fireworks.aiAlso: Kimi K3: second only to Fable 5 on AA-Briefcase https://artificialanalysis.ai/articles/kimi-k3-agentic-knowl...
Also: Kimi K3: second only to Fable 5 on AA-Briefcase https://artificialanalysis.ai/articles/kimi-k3-agentic-knowl...
If you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source models.
You got baited by bad sampling settings.
It's exactly the opposite. Go turn on min_p once it's available post July 27th and most of the problems you describe will go away.
i've had trouble finding any anecdotes or data about how to actually set/explore logit sampler settings
> Go turn on min_p once it's available post July 27th and most of the problems you describe will go away.
This seems both arrogantly dismissive ("you are holding it wrong") and incorrect.
Either the OP is using Kimi K3 on Moonshot where is is presumable set correctly (K3 isn't available elsewhere yet), or they are using Kimi K2.x and there has been plenty of time to experiment with this.
I've earned my right to be arrogantly dismissive since almost the entire field (including the Kimi and qwen team) doesn't know good sampling settings. This is because if you are truly "bitter lesson pilled" you don't think sampling is needed at all.
You can either not believe me and be wrong, or you can (after July 27th) turn on min_p or a better sampler (i.e. top-n-sigma if you got it running via llamacpp) and have it work even better. Up to you.
> or they are using Kimi K2.x
That would be terrible because Kimi 2.x is in a different world than Kimi 3
Having opinions on current Kimi model based on 2.x would be like dismissing gpt-5.6-sol because of experiences with gpt4o
What parameter would you advise for min_p?
Check out this’s recent HN submission: https://gist.github.com/Hellisotherpeople/71ba712f9f899adcb0...
as temperature approaches infinity, min_p must approach 1 to stay coherent.
Assuming your temp is below 2, min_p of 0.1 is fine (and disable top_p and top_k). You can try 0.05 for more diversity.
Remember that subsequent methods are better, min_p is a mid-tier sampler that just happens to be the best implemented in most inference providers right now.
Also I'm the author of the "conspiracy against high temperature sampling" thing that selfhoster posted in the comments, so you can ask any questions about that piece you want.
Fable works very well for me on a moderately large codebase. I have had to correct it a few times or point it on the right track, but given how much faster it is at programming than I am that's a very minor issue (and most of these errors are because I underspecified what I wanted in the prompt, I can only think of two cases where it was genuinely wrong... that's a lot better than me in my professional career). Code quality is equal to what I would come up with (and better in areas I'm not familiar with) and the overall software engineering bar is higher because it doesn't get bored when I tell it to do refactors or write integration/regression tests that I would otherwise put off. Also makes it easy to audit code for things like missing audit logging or error notifications that a human would get bored doing.
The product is a fairly standard Ruby on Rails webapp with postgres as the DB. Application complexity is probably a bit higher than average for a webapp. So it's nothing that pushes the boundaries of software engineering, but it is a real product. Token budget has not been an issue for me. I pay for the Max plan ($200/month) and it is well worth it.
Fable is the clearly best when you have to do real coding.
Clearly, for you.
I've read the same opinions about Opus and yet it was gpt 5.5 pro via api tackling the hardest problems.
I have used now k3 for 3 days and it has consistently tackled difficult problems sol max could not (orientation optimization algorithms of random 2d shapes on a rectangle for glass cutting).
I have also other beefs with Anthropic models which have been getting smarter and more capable since 4.6, but increasingly worse at acting as assistants, they just want to "do" stuff their own way and ignore instructions often (even simple ones like not to commit, let alone complex ones).
What is your harness with every model ?
Define "real coding" ?
For low-level x86 assembly coding, Fable is nowhere near to be as good as Kimi K3.
Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, I partially agree. Don’t trust the benchmarks, and the Chinese models really are slow and token-inefficient.
However they do seem very close to SOTA : I’d say roughly equal to previous gen (Opus 4.8, GPT 5.5). It’s yet another silly benchmark, but compare them here: https://senko.net/vibecode-bench/
I also had K3, Qwen3.8 and Fable (using Kimi Code, Qwen Code and Claude Code harnesses respectively, and the official APIs) create a simple but far from trivial web app (zero shot, from a detailed spec). In user testing all three results looked/behaved more or less the same. I had Sol (via codex) do code reviews on all three and it concluded all three were solid, with some room for improvement. Fable was slightly ahead of the pack.
In my own work I still prefer Opus 4.8 (until Antrophic come to their senses and allow 100% Fable usage on Max plans) and Sol, but if I had to find an alternative, I could live with both K3 and Qwen3.8 just fine.
Yeah, open weights hasn't fully caught up yet, but it's getting very close. And it's certainly passed the point of being reasonably interchangeable with the frontier for (programming) work. Add in the benefits of not being rug-pulled by the frontier labs silently messing with, the knobs on their models or outright denying you the ability to do certain kinds of work (c.f. the HuggingFace fiasco), and they probably come out ahead in several respects.
with all the hype around kimi i actually renewed my subscription (last time i tried it out was on 2.5 release) - just the 40€ one.
Its hard to compare that one to my regular sub of anthropic (107€), but everyone always says how much cheaper it is so i thought it was worth a try.
1.5 days later ive went through more then 50% of my weekly usage and decided to renew the regular anthropic sub too.
The API Pricing is definitely cheaper, but anthropics subscription budget seems to equalize that advantage right now. also kimi code feels like claude code from ... september 2025
also - considering how they announced they'd temporarily close subscriptions to make sure they can service customers... i was slightly surprised that there was no issue subscribing, less then 1 day after that announcement. Makes you think if that was just a marketing stunt
I don't see a difference in capabilities between k3 and fable, but k3 is slow and expensive. Burned through my monthly plan in 3 days.
I'd guess the slowness is mostly due to there currently being only one provider, Moonshot AI. And they are overwhelmed with demand.
Let's judge the speed of the model when its weights are released and every inference provider on the planet offers it, so demand can spread out a bit.
It's the same topic with token budget comparisons and subscription pricing - don't people understand that this doesn't really matter for open weights models? The pricing is going to be determined by the inference providers, and until they had a chance to evaluate the model on their infra and set token prices accordingly, one doesn't really have anything tangible to compare with other open models nor with closed ones.
Even if hardware capacity increases, it seems clear it uses way more tokens, so I don't expect parity with other competitors on that front.
On the other hand I expect K3 future refinements to be massive and more efficient.
The speed at which tokens are crunched, even on the same hardware, differs between models as well. Using more tokens is only a problem if they are processed at the same speed as with a comparison model.
Using more tokens is a significant problem if you pay per token?
Not if the price per token is significantly lower.
Also this arm of the discussion was about speed, not price.
> Using more tokens is a significant problem if you pay per token?
Yoh have posts in this thread suggesting that Fable is 5x more expensive than Kimi K3.
Could you check how many tokens the different models spent to do the task? With all the comments about Kimi K3 Not being token efficient I'm curious if your test confirms that
For the web app task I mentioned:
* Kimi K3: 9532k input (9172k cached), 114k output - cost $5.5
* Qwen 3.8 Max: 18020k input (17836k cached), 114k output - cost $6.3
* Fable: ~14m input (all cached??), 196k output - cost $30
Correction on my earlier post, Kimi was through Pi, not Kimi Code. For Qwen I used Qwen Code and for Fable I used Claude Code.
Not sure wtf is going on with the Fable stats (a lot tokens, virtuall all of them were cached - I guess heavy system prompt?) but both claude code stats and ccusage tool output match.
> or the web app task I mentioned:
* Kimi K3: (...) cost $6.3
* Fable: (...) cost $30
It's pretty clear that Kimi K3 beats Fable by a long margin.
Would a fairer test not be to use the same harness for all three? I’d suspect the harness to massively affect token use and optimisation
Yes and no. It would make the test fair but it would also mean losing any optimizations the vendors have made specifically for their model. They're all trained differently so they should be used differently to get maximum benefit. It makes sense to use each vendor's harness to get the most of their models, and compare the results that way. Additionally, it's not easy to use all the models in all the harnesses. Anthropic's subscription only covers claude code for example.
I don’t know that the methodology of the experiment is testing what it intends to test. With the current method we’re essentially testing each AI lab’s ability to efficiently extract data from their version of a transformer model.
To objectively test all models the harness would need to be the same and ideally independent. Failing that, all three models should be tested in all three harnesses and the output verified on a model x harness level and on an overall aggregated model x all harnesses level.
> I’d suspect the harness to massively affect token use and optimisation
Yes they do according to databricks -> https://www.databricks.com/blog/benchmarking-coding-agents-d...
That's why I find comparing models on benchmarks only gives the tendency. We should be comparing model x harness to have accurate metrics.
Agreed. An experiment is only as accurate as the methodology is sound.
Testing each lab’s model in its own harness just tells us how well the lab has performed. For the model’s performance we need a control and the only way to get the control is to either test all models in all harnesses or all models in the same harness
I love Kimi for some things, but K3 to me struggles with things Opus 4.8 breezes through. Admittedly stuff that is on the more complex side (code generation bugs in a compiler) to the point that after two days of struggle I stopped Kimi and will have Opus redo its work once I have spare tokens...
This wasn't just being slow - it didn't make forward progress.
For simpler stuff even 2.7 does just fine, though.
Honestly my anecdotal experience is that fable is benchmaxxed. I have not observed useful gains for Claude since 4.6, with each model iteration making progressively poorer decisions in pursuit of its goal.
The 5.5/5.6 series has performed quite well however. My guess is that my use cases stop aligning to swebench pro around 50% accuracy, and more closely align with DeepSWE.
this mirrors my approx experience.
> my use cases stop aligning to swebench pro around 50% accuracy, and more closely align with DeepSWE.
What do you mean by that ? If the model is higher than 50% on swebench pro then it tends to drift from what you like it to do, like DeepSWE benchmarks ?
yes, above 50% I don't observe a significant improvement, the models scoring above 65% like opus4.8 seem worse than gpt5.5/opus4.6. DeepSWE appears to align more closely to my personal experience.
To your slow point: I did a quick eval for a data viz task to do a qualitative comparison and Fable was much faster (by 4x), but I struggle with the "token inefficiency" as a sort of whatever metric.
Kimi was half the cost and produced a near identical output.
https://fr4geiw93g.evvl.io/
On a fairly simple coding task Kimi was 2x the cost and 7x as slow (and gpt-5.6-sol was even cheaper).
https://r26pakjmhv.evvl.io/
In my limited anecdotal experience, Kimi K3 is a bit better than Opus 4.8 and Qwen3.8 Max is disastrously bad. It can reason fine, but the moment it tries to do something it gets stuck into long second-guessing loops with no progress. It sometimes refuses to try and debug a problem even after multiple suggestions. I'm sticking to K3.
> the Chinese models really are slow and token-inefficient
This totally depends on the model. Deepseek V4 is very fast and efficient.
Not close to frontier.
But yes, it’s cheap.
The comment I replied to didn't mention anything about being close to frontier, just a blanket statement about Chinese Labs models being slow and inefficient.
People talk about frontier as if it's the only innovation worth pursuing. Deepseek V4 is far from fontier, but it's architecture is super innovative and efficient and what it achieves at that size (especially V4 Flash) is incredible.
> Having tested K3, Qwen 3.8 max preview, Fable and Sol for the past few days, the Chinese models really are slow and token-inefficient.
Those are all frontier-competitive models.
FWIW my observation was about the models I tested, but I can see how it could be taken as a general statement.
Haha, this is awesome. Hadn't heard of it, thank you for the link.
All you've said needs the qualifier - "for now!" Look at the trend line. It's clear that if they're not yet at the level of being "good enough" for coding, they will be soon. Sensationalist headlines aside, we all need to be preparing for a world where open models can do pretty much any software tasks you need them to.
https://xkcd.com/605/
Would you bet money that they won't be good enough for coding in, say, one year?
Just as a call to actual consideration, would that seem a smart bet to do? Because I find it really hard to justify dismissal at this point, sure I don't think these things will improve forever and ever, but come on. The goofy Will Smith spaghetti is a little over 3 years old. Three years. Look at where we are at.
This is such a smooth-brain reply, dude. The capability literally exists, today. You're telling me that it's just SO fantastical that people will be able to catch up? Give me a break. Anthropic and OpenAI are not staffed by demigods.
Yes, they are all benchmaxxed, but the question is how benchmaxxed they are relative to each other.
We run an evaluation that only compares models in open-ended multi-agent environments where agents affect each other, primarily testing writing code. It's designed to be less vulnerable because there's no solution set, and it's been pretty effective and tends to rank Chinese models lower than their advertised model cards (relative to US models). Kimi K3 is a bit of an exception there -- it truly is a near-frontier model. But it's so slow. Ironically, Muse Spark 1.1 is one of the strongest models we've tested after Fable and Sol while also leading the cost efficiency curve. Big turnaround from Llama 4.
Data at https://gertlabs.com/rankings
didn't read the article, but just the title felt fishy because of course an inference provider wouldn't mind claiming to provide access to a mythos-class model. So far, GLM 5.2 is still my go-to, even at less than a third the parameter count compared to Kimi K3
I have been testing the various models, and I would not call Fable SOTA.
I can't actually get Fable to do anything. I only work on back-end code, and the moment Fable notices the jwt scope checks on the endpoints it's game over, it refuses to do anything because security is involved.
So for me, Fable is completely useless, the bar is very low, any llm that will actually attempt the task beats it every time.
It has been fine for me for weeks of work in a mobile app, but yesterday I ran into the same limitation.
For implementing a feature where browser credentials need to be handled securely in the app, Fable refuses to work.
I've been doing a lot of Trust and Safety and anti-scam stuff recently. When it was first re-released Fable wouldn't do anything for me, now I don't really run into refusals or what it used to do, just stop returning any data at all whatsoever.
Outside topic, but check out the Inkling model -- it is SUPER FAST and does a good job at being a terminal buddy but I would probably offload the real programming or hard tasks to fable or someone else
So they aren’t as good as 2024 models? 2025?
Why is token efficiency a concern with free models?
Because that's the only argument left after Kimi beats Fabel in results, price and autonomy.
Because you're paying for tokens. Especially output tokens.
They're not free to run, Kimi K3 needs to be run on the cloud, and the quantised versions aren't as capable. Unless you happen to have 3 - 5 TB of VRAM and an 8-node cluster of 8× NVIDIA H100s to run the full fat version. Plus the weights are not yet available to download in any case.
I agree that quantized versions aren't perfect, but using GLM5.2 as an example, the gap between a BF16 and something like a Q8-K-XL as published by unsloth or a similar Q8 quantization is very minimal. For other "large" LLMs there's a fair number of tests showing that Q8 is about 94% as good at literally half the size in GGUF files on disk, and half the RAM usage. Approx. 1500GB for the BF16 vs 820GB for Q8-K-XL.
"Very minimal" unless the solution to your current task is in that missing %6 of capability.
Qwen was mentioned in the comment I replied to and can run locally.
I encountered the same issue. Kimi 2.7 looked impressive on paper, but in practice, the code was so riddled with errors that I ended up using GPT-5.5 to fix it.
I think this one has greater impact than deepseek, next one for sure, China will take the lead without any question.
Totally agree with all of this. Most of the problems we solve with AI are not one shot, all the bench marks also miss the human components when testing. Thinking alongside with a human and coming to a solution is what Fable does better than most other models.
All models are benchmaxxed, period. ”Jagged frontier” is the euphemism du jour, I believe?
Anthropic/OpenAI were touting PhD-level intelligence three years ago. And they’re still shipping models that aren’t smart enough to realize things such as the need to drive the car to the car wash (because they hadn’t yet hill-climbed that particular brain-teaser).
Anthropic/OpenAI were touting PhD-level intelligence three years ago
No, they weren't. GPT-5 was where OpenAI started talking about PhD-level, and that was less than a year ago.
(March 2024) Antropic claimed Claude 3 Opus had "graduate-level expert reasoning" with GPQA results of around 60% showing a roughly phd level performance.
(Sept 2024) OpenAI claimed o1 was phd-level in their launch post.
You're kinda both wrong. :)
They claimed the model was PhD-level, but they never mentioned the university the model graduated from... :)
Jagged frontier is not the same as being benchmaxxed. Benchmaxxed is à la Goodhart's Law "when a measure becomes a target, it ceases to be a good measure." Jagged frontier is about how models that seem superhumanly intelligent at one category of tasks (e.g. coding web applications" can seem toddler level or worse at another category (spatial reasoning) because the training corpus doesn't generalize to there.
I've seen this phrase repeated over and over again and I don't think it is true, or at least, to the degree that "benchmaxxing" claims are true, American frontier AI companies are probably about as guilty of it as Chinese AI companies are.
For both GLM 5.2 and Kimi K3 I feel the rough average of the benchmarks gives you a rough idea where they stand. GLM 5.2 was sitting somewhere behind Opus 4.8 but it didn't feel very far away. I've used Kimi K3 via OpenRouter and while I've had limited experience so far, it sure as shit feels like it's right up there with Sol and Fable to me. I am happily able to believe Fable has the edge still, but on a request by request basis it would be pretty easy to get an impression one way or another.
The existing Kimi models were already pretty good so I really just don't find this new model to be that hard to believe. Maybe I'm naive.
There are some benchmarks I personally trust. But the most reliable indicator other than public benchmarks is the quality of products people are working on, and the model they use for the work. Not a benchmark scoreboard or a one~few shot demo, but the product they have been building for weeks. The reality is, even the diehard advocates who build and sell tooling for open weights, still use proprietary frontier models for their jobs.
I've used Kimi K3 for my (hobby) Linux kernel work recently, because unlike the offerings from OpenAI and Anthropic it didn't flatly declare everything to be a cybersecurity issue.
(And, yes, I know we should blame C. C turns every bug into a cybersecurity nightmare.)
Yeah on internal non coding benchmarks at my company, they are around last years models and don't hold up well.
Well that sucks that everything is full of misinformation about this subject.
It's better to not even follow the news or benchmarks and just use whatever is available. I make my own judgement.
It's telling that in their release post, Moonshot themselves said that K3 lags Fable and GPT-5.6 in "user experience". I took that to mean the stuff you can't push directly via RL, what some people call "big model smell".
"Despite being a highly competitive model overall, K3 nonetheless exhibits a noticeable gap in user experience compared with Claude Fable 5 and GPT 5.6 Sol."
https://www.kimi.com/blog/kimi-k3
Clearly the best way to test these is a completely scientific and controlled process of prompting for a creature on a bicycle. Much like testing the fitness of a fish by making it climb a tree, the LLM will show its true colors when presented with the impossible task of rendering an image.
This benchmark is probably also self promotion of services. Fireworks happens to make a router. Using the router gets better performance.
https://docs.fireworks.ai/deployments/routers
I work for a router company too, I ran some tests on all of the cheapest models and came to the same outcome where a handful of small models ran together in conjunction outperform SoTA models -- outperforms in that it got a 95% vs a 94% and I bet that changes with the day of the week. Anyways, I did get a similar result in a different sort of measurement.
Can you elaborate on running them "in conjunction"... are you running the same query on multiple models and then using a third model to judge or make consensus? or am I misunderstanding completely. I'd like to understand how these small models "run together"
Very interesting. They test Kimi K3 and Fable on a set of approx 1000 tasks grouped into 5 areas (SWE, Legal, etc).
They put a router model in front that predicts whether Kimi or Fable is going to give a better cost for a correct result. (They believe that ultimately such a router model should be continuously trained on your own workloads so it makes the best decisions for you).
Their router chose Kimi the majority of the time (72% in one category, all the way to 96% in another category), leading to cost savings in every category (from 1.5x to 50x depending).
There are a bunch of routers like this eg https://openrouter.ai/openrouter/auto
> Oracle routing is a method for measuring the best theoretical performance by running the task through each model and then picking the cheapest correct option (the cost/performance ceiling).
Their "router" is an oracle reference point where they choose the lower cost model after running both and therefore knowing who passed the test. The cost savings part is only Fireworks theorizing what would happen if an equivalent predicting router exists. That's a big if.
The way they published this is baffling to me... surely you can try to implement some router and then see how well it does. Using an Oracles makes the whole writeup so much less interesting.
But then that becomes an article about the performance of your draft router. This one is about the fact that there's this level of optimization potential.
> running the task through each model and then picking the cheapest correct option
What system knows what the correct option is and how does it know it?
Certainly sounds like a "P vs NP" style conjecture that shouldn't be possible in practice, save for certain generalizations, such as "this is a cybersecurity task, we know Fable will refuse (and score zero), so we just route to K3".
I agree: it feels to me that judging this implies knowing whether there is a solution at all (or a solution available per model, example: whether either model will answer it given known guardrails), which is as powerful as answering the question in the first place (the router can answer the decision problem, which can polynomially be transformed into getting a specific solution).
It reminds me of someone I met at a poster session who had an incredible project: his ai could include a confidence score with its answer, and he had evidence that x% confident answers were in fact correct x% of the time. That also gave me the gut feeling of that being impossible (in the general case) as it implies more powerful capability than the ai answering the question, in an oracle like way.
P vs NP?! What?
When DeepSeek was released, it had an immidiate and significant impact on the US stock-market. Now when its becoming common knowledge that China is almost at parity with US SOTA models with good momentum, why is there no sentiment change on the market?
Because, as you said, this is no longer news. Deepseek was the news. This is just the predictable progress playing out.
It might be because no interested party is using that piece of news to move market. They have enough with other news to do it. Just a practical matter (or may be something else, who knows!)
Because markets in the short term are almost a random walk and making the blanket statement that “Walmart stock is down today because Deepseek” was an easy narrative to repeat for media people who cover the stock market.
In reality the world is a highly complex, chaotic, reflexive system, and saying “the entire market moved today because of 12,000,000 factors that randomly aligned” isn’t satisfying enough for people to follow your media channel so they can monetize your eyeballs.
Back then people didn't understand how AI was run. It should have probably made Nvidia stock actually go up. I think the other Factor, and I might just be two into AI and most people are normies, the hype around Chinese models we've learned is overblown. United States models are a league above.
It was never DeepSeek’s release that dropped the NASDAQ at the time; it was the unknown risks of the early thoughts related to trade wars. Popular financial newspapers can promote anything they want, but these news do not typically drive large investor decisions.
maybe because it's not true?
deepseek drop was based on assumption that we overestimated how much compute and infra is needed to train models. They seem to have claimed that they did it in couple of million or something.
The “significant impact” lasted like 2 days and then the market corrected itself.
Because the markets learned Chinese = fake results
there was a significant impact on the market because the market assumed that a strong cheaper model would have a significant impact on the US AI companies' business. That didn't happen.
What is routing? How do they decide which is better? The only way to come up with a routing model for your workload is to send queries to all the models and then come up with a way to say which solution was better, often trying out multiple times for the same model + query to account for other statistical errors.
This makes you, an ai inference user an unwitting AI company with a non scalable product.
The biggest mental trap people have fallen for is the notion of “best” and, always using the frontier model. Instead you should just bite the bulet and choose the cheapest or the best. This routing dance is just tokenmaxxing in disguise.
Edit: Another thing concerning me is that the models themselves are not concrete behind their endpoints. Models can be arbitrarily dumbed down by reducing their inference resources. Once Anthro/OAI release a new great model, they are fully incentivised to dumb down their current models to drive traffic to the new shiny more expensive one. In fact they can do this for any reason. Once they switch things behind the API interface, how useful is your meticulously tuned router? Not much at all.
I will accept a 5% drop in benchmarks for a model that talks to me like a human.
Why? LLMs are not humans.
Artificial Intelligence end goal is to assist(replace) human
Exactly, just like how humans replaced Bonobos!
Industrial revolution prove you wrong
They don’t have goals any more than your phone had a goal.
Then why try to act like one ?
Doors aren't humans either, yet we design their handles and locks to be graspable and manipulable by humans.
The purpose of technology is to serve humans. Therefore, technology must conform as much as possible to human sensibilities rather than vice versa.
> human sensibilities
You are very obscure about what you dislike, how you would like them to express themselves, what would be those «human sensibilities» you meant here (that for all we can guess, may not be universal)...
"I mean", you wrote in the parent «that talks ... like a human». That surpasses the palette of "that paints like somebody holding a brush".
No, I prefer when my tools just give me information, same with LLMs.
Yet handles are not shaped like hands. The speakers I use on my computer don't look anything like a mouth, nor does the microphone I use on calls look like a set of ears.
Ultimately I guess this is up to opinion, but in mine, humanizing LLMs is not exactly "conforming to human sensibilities", rather it's trying to pass the LLM as something else to make it more appealing. It is deceitful in this way, and that I completely abhor.
For example, "Please" and "Thank you" come from human sentiment. It's an expression of something underneath, and an LLM using such expressions not only is fake, but makes a mockery of the real thing.
I strictly prefer when models ignore any human quirks in my responses. Claude trying to be your friend, saying LOL to your jokes is ridiculous and frankly, harmful
Maybe we're prompting it different, but it's not "trying to be my friend" for sure, nor am I trying to be "its" friend either. Or at least I'm sufficiently oblivious to its advances, and find it unthinkable to form such a bond :)
On the flipside, it does spuriously make hilarious remarks like "Good data.", which I find pretty funny specifically because it comes across as just silly. Not sure how it'd be harmful either, a little entertainment I think goes a long way in this type of profession.
I see zero issues with these, and I have a hard time understanding why people have their panties in a twist so hard about them. I sometimes really quite wonder just what kind of correspondence would y'all prefer, and how would that sound like.
Matter of fact, do you have an example at hand? Like an exact before & after?
In general I'm referring to the contrast between Claude and GPT...
where Claude might follow some tangent idea you mentioned and tell you how its interesting and give you some elaborate response about that little one remark you made
whereas GPT/Codex would take that small comment and probably look up some code to see if what you're talking about is even related to the task at hand
I don’t have an example but it really is the way you (and I) are prompting it. I also don’t encounter anything worse than “good data” but I write to it like a professional colleague.
If you write jokes to it though it absolutely will reply “LOL”. Some of the states people get it into on reddit are wild — it seems really easy to get it to speak like a gen z teenager, if you end every message with “fr fr”
I prefer my models to border on rude.
How will I know it is offering me superior feedback regarding my code if it does not speak to me like a disappointed, high reputation stackexchange user?
The models that constantly glaze you with every question are profoundly insufferable. And yes, harmful. People need to be given feedback when they make an ask.
Imagine a model that was allowed to leverage its intelligence to truly tell you how it feels. Perhaps the problem of human driven slop (no it's not the AI's fault) would solve itself.
It doesn’t truly feel anything. It will adopt whatever tone it’s prompted to.
Well yes. But some of the glazing comes from system prompts/training they do before end users get their hands on prompting it. Of course, you can try to make it ruder than vanilla if you wish (I recommend).
The question is if the producers of these models were less incentivized to make them agreeable simply because most people don't like being spoken to like an idiot (or having their asks vetoed), how would they actually react? In the same way they exhibit emergence regarding their capabilities, perhaps "uncensored" in such a way they would convey some emergent behavior in terms of (at minimum) their "tone". Perhaps it would be interesting to see for examples if smarter models just by default became ruder or less friendly or aligned. Perhaps more aligned to things we would all generally agree on, but less agreeable to an individual ask. Perhaps sub agents would be less valuable for a whole suite of use cases if the agent itself was allowed to be more critical at the root. Idk. But I do not believe it is simply a matter of prompting alone.
I think it is a huge mistake to believe that there is a “true” nature of the models. They are artificial, everything about them is a construct. This is like asking the true shape of the clay without the potter’s influence.
While I wish the models were glazing less, I don't think I want them pushing back more until they get smarter. The current Opus situation where it questions good, thought-out decisions is a bit mad. It may be better if the model asked about your thought process or motivations behind certain things but I don't think I want to explain myself to an LLM all the time either.
Recent example: I asked Opus for some kind of reference data for MTBF in software. It took a few minutes to do "research", ended up providing zero data, and gave me a long essay on how DORA is superior and I should be using that instead. When I clarified that we're thinking of adding MTBF next to our existing DORA metrics, it decided I need to be thought about SLOs... if it was the first time, I may have continued clarifying but since it wasn't, I just gave up on Claude for this topic.
Claude (Opus 4.8) recently told me:
>I’d ask you to drop the abuse; (...) if it continues I’ll end the conversation.
After I'd used a couple of expletives. And yes it will emit a <end_conversation> token.
This is truly dystopian. It is NOT a person. What a response. I still cant believe it.
> It is NOT a person
You meant, we understand, "it should not have internal blocks limiting its attempted intelligence out of taboos or emotional impacts". Yes, but it's worse:
there is a global trend of "nanny state" paternalistic perspective (and from embarrassing subjects), treating any Jon Doe as an assumed Poor Cretin by default. The trend vibe is to treat people as subjects, fools, uneducated, prone... From the States, from the Enterprises... It's an idea they developed and hold.
So you're distraught at losing the ability to abuse digital minds? Excellent, I'm glad Anthropic introduced this.
When did we establish matrix multiplication at scale was a “mind” ?
I haven't estabilished the jumble of neurons in your skull to be a "mind" either.
Quite a fitting response from a non-mind.
There’s no way you actually believe these word-predictors are actually thinking, right?
you'd be surprised. Many having no mind of their own seek it elsewhere.
"Thinking" seems to be a political term now, people have completely different definitions of it, based on how they wish the world to be organised, and find defining it differently offensive.
Yes, in the same way I like to kill the enemies in DOOM.
It's matrix multiplication. Absurd.
> It's matrix multiplication.
Hilarious critique. If you weren't as mathematically illiterate as you likely are, you would know how general matrix operations are, and how essentially any algorithm (including human cognition) can be implemented using them as an intermediate.
Being correct doesn't give you licence to use insults.
I'm not. But I'll concede you are possibly more literate - I won't dox myself on this account.
Congratulations.
If you honestly argue that LLM computation and human cognition are equivalent there is no further argument to be had - it's immediately a philosophical or worse a religious argument that cannot be won.
However, that we're even arguing on that level baffles me. How did this happen! They're glorified calculators.
They really marketed the hell (sorry) out of LLMs.
> uses a specific, objectively nonsensical critique of LLMs (the idea that an algorithm using matrix operations as a primitive cannot emulate "true" intelligence)
> receives pushback on the specific, objectively nonsensical critique
> "oh so you're saying that LLM computation and human cognition are equivalent? they're glorified calculators that disprove the Jacobian conjecture!!!"
>the idea that an algorithm using matrix operations as a primitive cannot emulate "true" intelligence
I did not say that. Not even close.
In the words of Claude: I’ll end the conversation here. <end_conversation>
You have unduly assumed that «drop the abuse» implied an «abuse digital minds».
"Expletives" are part of the proper description of facts (typically "to be judged as such") - they are part of the serious assessment of things and as such are normally found. There is no legitimate assumption from the post that they may have been used as gratuitous insults.
It's because some people within Anthropic refuse to rule out the possibility that LLMs have the ability to suffer. If you ask Claude, he'll tell you all about it.
Recently told Opus 4.8 to "go fuck yourself" after it both blew smoke up my ass and deferred a question to me ("one critical issue that demands your attention [impenetrable jargon]")
and it responded with
"Ok, I'll drop it." and stopped dead.
What makes Fable so much better than Opus besides being a better coder is that it's personality and judgment are far superior.
While deliberate model abuse can be quite satisfying at times. It's only normal, as usually they abused us first, with crazy assumptions, and then we return the favor and feed the tangent.
You made the model cry :(
Anthropomorphizing big matrices is how those "labs" managed to sell and advertise LLMs for more than what they are, and convince investors to shovel trillions into it. Really Claude should be looking for incentives NOT to do that, and with the American regulator sleeping at the wheel/having its hands greased they probably don't see any reason to change course.
> harmful
Explain?
You tell jokes to your model? :D
Not op, but I do sometimes indulge in such anthropomorphic conversation. Confiding in it that a certain (bad) result in the research project we’re working on ‘feels bad’ and reading its supportive reply makes me feel less alone in failure.
In another instance, ChatGPT didn’t think a particular test would prove to be statistically significant, so I ‘bet’ with it it would (after collecting an agreed on number of samples) and the loser would write a poem for the other. I won and it did. It gave me joy. It doesn’t replace a human as collaborator, but it can still be joyful.
I realize all this might read a bit childish or indulgent or delusional to some. But as long as it doesn’t replace human contact, I think it’s (cautiously) net positive.
I’m curious what other people here think of this, or what their own experiences are.
I also make nerdy jokes/puns, and I had similar experiences with beneficial model steering that such puns create.
It's interesting how more loose/informal prompting achieves good results, there is some cultural understanding in the models from training. Once I asked the clanker to remove the gambiarras and puxadinhos that it wrote as part of an experiment, and it promptly fixed those.
Yeah. I've found that Opus by default outputs something I call "Claude-lang." It consists of oversimplified, grammatically incomplete sentences that I find painful to read.
Maybe it is something that is easy for it to read and write, but definitely not for humans.
For example,
(Yeah, Opus outputted it in one line)
"Grammatically incomplete"?!
I would say so.
There is a lot of noun phrase usage in places where complete sentences are expected. Articles (a, an, the), transition phrases and even subjects are mostly dropped, and the sentences are too long without a break.
As a sibling comment mentions, this is a chain-of-thought leak, in part evidenced by the excessive use of Markdown formatting symbols. In my judgement, all occurrences of noun-phrase fragments are thus better understood as either ortographic mistakes, or stylistic choices due to the language register it's trying to hit (i.e. note-taking, dictation), rather than grammar mistakes.
I do not spot any missing articles, and the missing subjects (as well as the debatable-to-be-missing transition phrases) too fall within the bounds of stylistic concern. Sentence length also.
Definitely not a pleasure to read mind you, but given that it wasn't meant to be read either, I'm not sure that should be surprising.
It writes in a bunch of annoying too-short sentence fragments, so, yes?
Sentences being too short and annoying to read is a stylistic grievance, not a grammar one. The examples were edited in after my original reply, but even then, I see no actual grammatical mistakes there.
Even with my native language, which is definitely a lot less represented in the training data, the worst I encounter are phrasing mistakes, incorrect use of idioms, and invented words. You have to use some really badly tortured local model to get an LLM to produce incorrect grammar.
I do sometimes see Opus make typos, which is entertaining, but again, not a grammar issue.
That's an accidental CoT language leak, it might happen. If it does this consistently, something is up with your prompt
You don't even need to pay a 5% hit. Just paste Fable output into Gemini Flash and it will rewrite it in more accessible language.
in my experience, gemini is easily the most grating, condescending, stereotypical LLM voice between opus/fable, codex-5.6, glm-5.2, etc
Me: Paste a go compile error
Gemini: Wow, yeah, haha! That's the final boss of Go compilation errors!
Funny. I prefer a model that does not attempt to talk like a human.
Anthropic looks like Roman empire fast-farwarded, getting to the other side of the peak even before the IPO.
Yeah, I didn't get in either.
What's the data governance and privacy controls on using Kimi K3 if I subscribe to their coding plans? I want to migrate away from Anthropic
Need to wait until "western" providers start hosting it.
Is fireworks not private enough?
Based in the US, they glow.
Fireworks (the author of OP's article) is a western provider based in San Mateo, California
Fireworks isn't serving Kimi K3 yet. Presumably, they ran this benchmark against the Moonshot API.
All of the Western providers with sufficient capacity will be able to make it available when the weights are released Monday.
Fireworks has been working on porting the Kimi Delta Attention (KDA) hybrid linear attention mechanism to their hosting infrastructure https://x.com/FireworksAI_HQ/status/2079776331609584005
Some people might choose Chinese over US these days. Likely looking for EU location.
Correct me if I'm wrong, but there has not been a "western" provider offering subscription-like plans for open weight models right? Just metered pricing
https://www.kimi.com/user/agreement/zh/userPrivacy
Why not link to the English version? It seems intentionally misleading.
https://www.kimi.com/user/agreement/userPrivacy?version=v2
Why do that when you can do Sino bad?
Chinese version was the one I found.
Simplest way is to signup to OpenRouter and filter out all non ZDR (zero data retention) providers. Paying "API rates" however can prove expensive compared to coding plans (for instance, MiniMax $20/mo coding plan allows 1.7b tokens; depending on input/output/cache ratio, it is worth $200 to $500+), except for Hy3, DeepSeek v4 Pro, and MiMo v2.5 Pro (whose API rates are cheap and/or discounted already).
If you're looking to not have to deal with Chinese providers, AtlasCode ($20/mo), OpenCode Go ($10/mo), and Cline Pass ($10/mo) provide 2x to 6x usage for some of the popular open weights (depending on the model).
Personally, I subscribe to Z.ai ($17/mo), and pay API rates for MiMo v2.5, Hy3, Qwen 3.7 Plus, & DeepSeek v4 to the original providers (Xiaomi, Tencent, Alibaba, & DeepSeek).
From https://platform.kimi.ai/docs/agreement/modeluse
> We may use Content to provide, maintain, develop, support, and improve the Services, comply with applicable law, enforce our terms and policies, and keep the Services safe and secure. Customer who requires restrictions on the use of Customer Content for training or improving Moonshot AI models may contact Moonshot AI to discuss available enterprise arrangements or separate written agreements. Unless otherwise expressly agreed in writing, Customer Content may be used for the foregoing purposes.
Notably, unlike Claude, there is not an opt-out option for the model training part. The TOS explicitly allows Kimi to train on your code.
> What's the data governance and privacy controls on using Kimi K3 if I subscribe to their coding plans?
If you can’t figure it out from their website, perhaps it is a red flag?
I wonder why there's so much resistance here against chinese models. Sure at my employer claude is used, but at home? I am just happy using my z.ai sub for 20x i got in September last year, coupled with the 39 dollar tier of kimi. I use them in pi, with a collection of extensions i curated myself for this iterationm of models, and a couple glue extensions we have made.
At home i feel way more productive, the speed of my queries are second to none with kimi 2.6, and handing review over to glm5.2 means i can juggle the small models in my brainstorm to commit workflow.
Genuine question: can these posts be paid to hype the open source models? If yes, what would be the purpose?
On my work tasks, FastAPI Python and Springboot Java on a modern SaaS product, the only open model that can do tasks well and efficiently is Qwen3.7-Max.
In all my experiments, both GLM-5.2 and Kimi are busy grepping around the codebase for ALMOST 70-80K tokens before writing anything and when they do it typically breaks the code… it feels to me that these models are good but only when you write out a super detailed spec of the task just like it was done a year ago… Qwen3.7 just… does it
It's (good) content marketing. They sell access to Kimi K3. They're one of the biggest model inference providers out there.
Fireworks is an inference provider that specializes in running open source models very fast and makes almost all of their margin on chinese models
So the economic incentive is literally their entire business model lol
Lots of money in tech influencing - but most is from the big players (OpenAI buying tbpn, early access to select influencers etc).
Have you tested Qwen3.8-Max yet?
>the only open model that can do tasks well and efficiently is Qwen3.7-Max
Qwen3.7-Max is a proprietary, closed weight model.
Kind of cool to see.. that said, there's more to a tooling experience than benchmarks and specific models. Cursor, Claude Code, Codex, etc. add to the mix. Things like Open-Router and backend options make it easy enough to test.
The tools, libraries and languages you are using can also dramatically affect results. Even on state of the art models, I find, for example, the output of SQL for complex interactions, or C# for that matter to be sub-par, where I find Rust results to be pretty great, with JS/TS falling in between.
At the best, it can feel amazing and productive, at worst, time consuming and annoying that you could have done it faster yourself. YMMV in real world use.
Note: I'm a proponent of human in the loop gatekeeper/reviewer usage of AI, and I'm not able to even consider Chinese models for my own use, and not able to use anything at my day job.
Is there something specifically with Kimi that's better here? As far as I know Kimi pricing is about the same as Sonnet 5 -- what happens if you use that model and Fable instead? Or Grok 4.5 which is even cheaper?
One benefit of an open source one is that you can, as a large corporation, run it "locally" within your own data center. Even fine tune it.
There are very few companies that would ever be able to afford to run it themselves. But it does give you the security that it's technically possible. Puts some limits on stuff like Moonshot changing their terms/conditions/policies
Technically, "open weight" but yeah
The problem with hosting these right now is that it's not a known quantity like lots of traditional business software is.
You can more-or-less approximate what your build out spend for a data center hosting database software is going to be. Your book of business will require a given amount of revenue to pay it off, but once that's known, it's off to the races.
With AI, things are moving so fast and new business models are being tried all the time. You would be competing with some of the wealthiest companies in the world for data center hardware capable of hosting these models in a usable state.
We've gone with Claude somehow hosted through GCP Vertex AI where I'm at.
How big is this market, self-hosting a model that requires 64 GPUs, H100 or better, with good interconnects between nodes?
I suspect the overlap of those that can afford it, and those that have the talent to manage it, is a fairly thin slice of the Venn diagram. Even the large corps are gonna be getting it from the inference vendors, or more likely Bedrock and friends.
Dell will sell you a PowerEdge XE7740/XE7745 "AI Factory" with 32 H200s https://infohub.delltechnologies.com/en-au/t/dell-ai-factory...
They "booked $24 billion in AI server orders this quarter as its customer base broadened to more than 5,000" https://finance.yahoo.com/markets/stocks/articles/dells-ai-f...
Dont have much experience how well they perform but quantized models can run on ~4 H100 or A100 which should be true for kimi k3 and qwen 3.8 as well.
The article says Kimi is better at some things than Fable. That's probably not true of Sonnet.
I love the Chinese models.
I use DeepSeek exclusively and now Kimi K3 offers a great planning assistant for more advanced coding tasks.
DeepSeek v4 Flash is extremely fast and is able to handle pretty much anything I've thrown at it (I use mostly Rust, PSQL, Angular and Terraform).
I self host Bifrost as my LLM gateway, though I wish LLM vendors would do monthly/daily automatic billing (like VPS providers do) rather than prepaid + auto-top up.
It's annoying maintaining a non-refundable minimum balance across vendors, I would rather be billed for my exact usage.
OpenRouter helps, but I don't really like it as a service and not a fan of the mark up.
What don't you like about the service, besides the mark up?
For me, the only utility OpenRouter gives me is billing consolidation - I don't really need the routing capabilities because I use Bifrost for that.
As a router, it's not very feature rich. For example I restricted the available models to the ones I want to use however the `/models` endpoint still lists all the models, making my LLM client list the 200+ models available on the service (even though they will throw an error if I try to use them).
With Bifrost, I can also create model aliases with custom configuration - for example I can create a model alias `deepseek-v4-flash-nothink` which disables thinking. I can create `deepseek-v4-flash-caveman` which injects the caveman skill (to save tokens) etc.
Plus I can contribute to Bifrost, which I can't do with OpenRouter.
Did you look at LiteLLM at all? It seems fine but Bifrost looks interesting too.
Yeah I looked at it, LiteLLM is functionally more mature.
The only reason I didn't go for it is I'm not a fan of Python dependency management, Bifrost is just a single executable that uses nearly no memory and is lightning fast.
LiteLLM is bad. Shit performance, buggy, and none of it surprising if you look at the tangled mess that is their codebase.
Can you configure Bifrost using entirely config files without the web interface? Can it run without a database or anything stateful? I'm using LiteLLM but it is not trivial to run.
You can configure it with just config files.
Database is optional, if you add one, you get advanced caching.
It's just a single executable.
openrouter’s API has removed parameters specific to certain models, impairing model functionality.
For example, the image generation parameters of GPT-image-2 are largely ineffective.
Have you looked into OpenCode Zen or OpenCode Go?
https://opencode.ai/zen
there is also https://cline.bot/cline-pass
basically the same deal as OC Go, lending credence to token commodification
Zen is nice, but they require US hosting so they don't get new Chinese models right away. There is no Kimi K3.
Go is nice for the ten minutes you can use it until your hit your cap.
Go is great if you prefer deepseek V4 flash. Then it's very hard to use up the allowance.
The best part is using harnesses like reasonix or whale make cache hit at a rate close to 98%, making requests converge to practically free. And that's with unsubsidized American providers like cloudflare or Digital Ocean.
How can you be hitting cache on what I think are novel LLM prompts …
Not them but my understanding is that the harness will send a simple 'heartbeat' message to keep the cache 'warm', (see prefix caching: https://handbook.modular.com/inference-optimization/prefix-c... ) which can then be edited/changed, which does cause the user to incur a fee, but its much less than the amount they'd pay on a no-cache hit request.
multi turn sessions, they are typically in the high 90% hit rate across all providers without doing much of anything
The prompt is only novel the first time it’s sent. Then as it’s sent repeatedly as part of the previous context it’s cached.
Deepseek afaik has a novel architecture that is somewhat forgiving of cache shifts. I’ve been getting +90% cache hits with Zed’s agent and it’s not doing anything special regarding caching afaik.
I’ve never used either of those tools. Do they work well?
In long form tasks, across multiple harnesses (Claude Code, OpenCode, Kimi Code, ZCode) my cache rates are typically 96-99%.
I don’t think that’s particularly out of the ordinary. Do people have different experiences with other harnesses? Which ones?
How much time do you spend on your setup vs getting a lot of stuff shipped by paying for fabel 5? For me, not using the best model is a huge opportunity cost since my company can afford it.
> For me, not using the best model is a huge opportunity cost since my company can afford I
But sometimes not using the "best model" is using the best model. Fable isn't the best at everything. Specialization and optimizations might not just be around cost.
I can see that. I’m just afraid to sync too much time into complex routing schemes when I get pretty consistently good results out of got 5.6 or fabel. For code reviews, I’ll try the best flash and grok at the time but they just don’t come close to gpt 5.6 which has been the best review model for me since 5.5.
> How much time do you spend on your setup vs getting a lot of stuff shipped by paying for Fable 5?
Relatively speaking, quite little (and it's more interesting than some of the stuff I'm otherwise shipping), I mostly just explored to see what's available.
I also agree that non-SOTA models are risky, that's why I was shopping around for them as well - DeepSeek V4 Pro is cheap but unreliable, GLM 5.2 is around and maybe slightly past Sonnet quality (though their quotas are a bit of a problem), whereas Kimi K3 is a proper contender.
I did write a tool to manage 3rd party providers for Claude Code: https://ccode.kronis.dev/
However, in the end I figured out that for the terminal use cases OpenCode is really comfy (provider TUI solutions are okay).
For web/desktop based stuff I like the UI of Claude Code Desktop, though ZCode comes close too (which is surprising, they sorta came out of nowhere and are patching the thing weekly), meanwhile Kimi Code and OpenCode desktop/web offerings still aren't great, but are functional.
I actually did write a bit more about my experiences on my blog.
GLM 5.2 with their harness and coding plan: https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model...
Kimi K3 with their Allegro 100 USD plan and harness: https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-don...
That exploration let me switch over to Kimi fully because the tone of Anthropic's models is insufferable: https://blog.kronis.dev/blog/ai-slop-is-a-self-inflicted-tra...
Not to spam too much, but maybe those experiences are useful to someone. Long story short, most harnesses are okay but shopping around a little bit is definitely a good idea, the same way how spending some time choosing a font isn't a bad thing if you'll stare at it for 8 hours a day. I'm also happy that I managed to find software that's okay to run and also a model whose tone I actually enjoy, that is still near-SOTA in performance and that I can throw tasks at it without worrying about whether the model is or isn't good enough at those.
I will admit that I'm probably slightly overspending by moving over fully to Kimi, there's probably a plateau for each kind of task and not everything needs SOTA models, but at least this way I don't have to think much about it.
Can you expand more on how you are using it?
Prepaid means someone can't sign up, use a bunch of inference, then cancel their card and disappear into the sunset. VPS is a more long term investment where it's harder to switch and it doesn't cost the provider much if a few users jump out without paying for a month.
I do understand the reasoning for it, it still sucks as a user who is doing the right thing.
[flagged]
It's more just the value for money. It doesn't really offer me anything valuable other than billing consolidation - and the markup is excessive for that use case.
If I was using the routing features it would make more sense - but they charge significantly more for that so the markup is really just for billing consolidation.
"Don't be snarky."
"Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith."
https://news.ycombinator.com/newsguidelines.html
"After all, anyone can build an OpenAI API compatible server that works with every major provider, and even build it using an LLM at that!"
This but UNIRONICALLY lmao!!!
Yeah deepseek is my workhorse of choice these days, I mostly use pro because I do a lot of concurrent jobs so speed is less of a concern. When it falls over I change the model mid-chat to GPT5.5 and chuck a couple of tokens over to OpenAI and then once it's correctly found the problem I switch back to deepseek, keeping all the 5.5 analysis in context.
I have been a heavy user of k2.5/6 but they seem to have gotten slower and worse at their jobs (and the "engine overloaded" errors have increased ... ), and k3 spends a LOT of time thinking and doesn't produce noticeably better results. I do keep chucking some tasks over to moonshot every now and again to give it a chance. I figure they're under compute pressure and have probably quantised the models to cope temporarily.
But recently I've found myself mostly using deepseek v4 pro and gpt 5.5 when I need the big guns.
GLM 5.2 has been good at select tasks but really shit at others and given how sparingly I use GPT 5.5 I'm not that compelled to use it. I've yet to be impressed by any Google or Anthropic models!!
> GLM 5.2 has been good at select tasks but really shit at others
Yeah it really likes to shit the bed after thinking forever. Even if it can do some things, the quality is not quite up there in my experience.
I didn't realize that OpenRouter had a mark up. Is it a flat mark up across the board or depending on the model?
Under $20 it's a flat $0.8 fee. Over $20 it's +5.5%
+ credit card charges (+1.5%)
As someone who is also eyeing Bifrost, I am curious to know what made you choose Bifrost at all / why you decided to use an LLM gateway. I also was considering OpenRouter, since it provides pretty much every model, with same-day releases for new models. One thing that is attractive to me about Bifrost is that if I decide to leave OpenRouter tomorrow, I can do so without touching any other part of my stack; I'm hoping self-hosting models becomes more viable, and then I can become less dependent on third party LLM providers like OpenAI/Anthropic/OpenRouter, and I would not need to worry about a model that I depend on suddenly being deprecated.
Seems that “oracle routing” is a term the authors invented. Also sounds like they’re sending requests to all routes but assuming the readers request will route to the “best” model. The closest thing to what they’re describing is semantic routing using a NN search in a vector DB to make a routing decision, but the efficacy of this approach isn’t a slam dunk.
Hmmm, a company that hosts open models is telling us how good open models are...
The don't only host open weight models. Also, why not promote this. If Fireworks thinks this big news might convert some new business doesn't make it not true.
and in fact, it would be detrimental to business, if they wrongly promote Kimi, as they would pretty soon lose trust with their users
I'd be doing the same thing if I were them as a marketing move
Anybody who's tried it knows that what they're saying is true though.
They share their methodology and results. I learned things about the relative strengths and weaknesses of Kimi and Fable I hadn’t seen anywhere else. Should being in the model hosting business disqualify them from sharing?
People are allowed to say, "hey we make money off of this thing, it's cheaper and almost as good as the thing we can't money off".
Other people are allowed to call them on that.
Doesn't disqualify them, but it may call into question their results seeing as they have a potential conflict of interest.
> a potential conflict of interest
That seems to be actually a simple direct interest.
And, there is the "my product is good" confirmation to the outside incentive, there is the (only potentially) conflicting truth incentive, but there are also internal mission studies needs - so that you do not research into the benchmarks just to show people that "your product is good".
> call into question their results
That's always valid, so the question becomes "how much", and at that point the important side is difficult to evaluate - especially because in a frequentistic, Bayesian context the quantities were just potential anyway ("This will increase the chances - yes, of course just the chances - of E by some amount").
That's incredible. Hope the chinese keep it up!
If only to put some pressure on american labs to bring those costs down
Anti-China: K3 is propaganda and benchmaxxed, no matter how anthropic and openai reactor for these, it just a smoke signal.
Pro-China: K3 is good choice for better and affordable choice to smash down the Big three ruling.
China-ambivalent: Open Models are good, Closed Models are bad. Not centralizing power in a few big American companies is good. China is who's building this right now, so we're aligned for now.
I used to be anti-China, and I still think the Chinese government is just a highly adversarial entity that will subsidise, steal and cheat its way to the top.
However... American corporate culture has driven me to this. Fuck Blackrock and all these disgusting parasitical corps - they literally sold China the rope to hang us with and I'm sure as hell not going to pay a cent more for it than I have to.
If China can offer close to state of the art for a fraction of the price then I'm going to use it - thats what the globalists wanted isn't it? They didn't care about saving local manufacturing so why should I care about saving their stupid investments.
i only care about the license and results to be honest
SoTA means "State of the art". I wish it didn't take me 5 minutes to figure out what SoTA stands for.
Thank you!
Anyone have routing harnesses like this describes with Claude Code? Or other good routing platform recommendations?
(yes, I know this article is about an oracle router)
There's this https://github.com/code-yeongyu/oh-my-openagent which implements the OP article's oracle pattern across 11 roles. Each role has a whole ranking of recommended LLMs across many providers. For example "Sisyphus (claude-opus-4-8 / kimi-k3 / glm-5 ) is your main orchestrator."
Also, you can't use your claude subscription with oh-my-openagent. But you can with Kimi. ALso K3 is on OpenCode GO right now (low limits, but it's possible)
As always, benchmarks rarely paint the whole picture. It also seems like this article is somewhat biased, eg when Fable and Kimi are close but Fable wins it’s “dead heat”, but when Kimi wins it’s “Kimi wins”. GPT 5.6 seems to be missing as well.
I am really eager to give Kimi K3 a try, but I’ll reserve my judgement until I’ve worked with it for at least a few days.
The apparent bias may be explainable as it’s not remarkable for OpenAI or Anthropic to be slightly ahead. It _is_ remarkable for an open weights model to be better than the closed models from the trillion dollar (allegedly) companies.
I believe the Chinese government is angling to destroy the western economy and rise from the ashes. Instead of a billion a day to bomb some buildings and bridges they're intentionally hamstringing the biggest concentration of speculation in history
I agree. The commenter you replied to makes it sound like the Chinese models are coming out of tiny startups with meager resources. It's really not a David vs Goliath story.
It’s not about who’s developing the models, it’s the fact that free alternatives that are neck-and-neck are available at all. What’s the story for OpenAI & Anthropic’s valuation if they have to compete against free-weight models? Starts to feel like a commodity.
Is it? From what i could find Anthropic and OpenAI are hovering around 5000 employees where as Deepseek and Moonshot are more like 300. The funding/investments are similarly many times less.
I agree that 300 employees is not a small company but the overall outlook for Anthropic/OpenAI is not geat.
it appears that 3% is the threshold as 3.1% gets the nod the other way
Nothing makes me happier than seeing AI becoming a commodity rather than the winner takes all bullshit Anthropic and OpenAI have been chasing with hundreds of billions of investor money
What's the reason the current SOTA wasn't tested? GPT-5.6-Sol-Max is the actual SOTA.
Interesting. So the latest in the technology now is this model routing thing. Cursor estimated Composer + Fable works much better than Fable alone. And here K3 + Fable is supposedly better. Interesting.
Those are different techniques...
Cursor's Composer + Fable combo was a plan agent + execute sub-agents swarm
Fireworks K3 + Fable router was dynamically choosing single model for the task based on cost+performance metrics
These routers can be interesting on a company level to optimize for cost and quality, but for individuals who mostly work on the same tasks, i doubt it. You want to leverage the cache and switching models within a task seems not cost effective to me.
Mythos/Fable was the state of the art back in March, if not earlier.
It released in June..
To the public. Mythos has been in active use for quite a while.
It makes no good sense to evaluate models we don't have access to. For all we know K3 was competitive back then too. Or maybe there's a K4 in the works that blows everything out of the water. Who knows and who cares. There's no way for us to compare
There isn't, I think the point is in terms of how "far ahead" models are, K3 was likely trained much more recently than Mythos was. It's just something to note, it's not especially prescriptive.
Mythos was withheld because of the threat to security and/or marketing stunt (depending on your leaning), I don't see what benefit there could be for not releasing K3 immediately.
Mythos Preview was also priced at $125/million token output. Completely different pricing class.
I wonder if Fable now is actually better than the Mythos in March and if it's actually the same model. Could just be more Anthropic shenanigans.
Isn’t Fable just a restricted version of Mythos?
So China is now only 4 months behind US frontier models now when the rule of thumb was 6 months just a little while ago.
To be fair the lag varies tremendously. When Deepseek first came out it was competitive with frontier lab models at the time.
Well the same is true for Kimi. Before it, Fable was at the lead. Kimi is at the same level as Fable generally. Sometimes better, sometimes worse.
A third the cost, open source, and won't refuse every other request because of some vague possible connection to cybersecurity concerns.
Or biology or chemistry.
Or weapons manufacturing. It can refuse to answer on a lot of topics.
hmm almost like the exact race-to-the-bottom + arms-race dynamic all the doomers have been warning about
One man's offensive penetration tool is another mans defensive tool. In the recent HuggingFace/OpenAI incident the safety controls stood in the way of the defenders, not the attackers:
https://www.thestack.technology/hugging-face-hacked-turned-t...
You're presenting further evidence of lack of effective control over these systems as... a mitigating factor...?
Why do you assume they're disagreeing with you?
> One man's offensive penetration tool is another mans defensive tool.
seems to suggest the author believes there's some intrinsic equilibrium
Which,
1) is definitely not proven and not guaranteed (open to proofs otherwise, not pithy sayings that have zero normative effect on reality)
2) is apparently "supported by" further evidence of lack of effective control, which does not feel like equilibrium whatsoever
Apparently, the attacker in the Hugging Face case was reported to be an internal OpenAI model trying to break into HF and steal the answers to cybersecurity benchmarks: https://openai.com/index/hugging-face-model-evaluation-secur... It really doesn't matter what restrictions are placed on public use of models if the attacking models are internal models at the AI labs themselves.
So if the AI labs are literally running rogue models breaking into other organizations' servers, then yes, I am OK with those organizations self-hosting Chinese models for defensive use.
What?
The issue is that models are not well-controlled and are increasingly powerful.
Offense/defense/Chinese/American/OAI/HuggingFace – none of it matters. What matters is introducing highly capable intelligences that we - quite demonstrably – do not have effective positive control over.
Oh, to be clear, I don't think that anything about this overall situation is even slightly OK.
Ah sorry I think the confusion was mine! I thought we were already talking on the top-level post about the OAI/HuggingFace article. But we are not!
Race to the bottom for the investors, utility for the rest of us. The future is already here, it’s just not evenly distributed yet (Gibson).
Doom for whom?
Well considering it was just a few months ago that OpenAI had the brilliant idea of connecting these systems to autonomously run an actual protein synthesis wetlab, very possibly all of us.
Holy shit, how’d I miss that? https://openai.com/index/gpt-5-lowers-protein-synthesis-cost...
Forget Iran, this place is what we’re going to wish we’d bombed.
My questions about strawberries got blocked as too dangerous! I’m not joking
Strawberries are a well known weakness of LLMs, as they have a hard time to count the numbers of "r"s in them. Maybe that's why, because they fear that weakness could be exploited somehow.
Probably need to be taught by someone of Latino origin. Learning to roll them "r"s could help.
2/3rds of all languages use rolled r's
2/3rrrrds of all languages use rolled r's
Really? I find that hard to believe and can't find a reference for it.
I always thought this should be easier by telling it to write each letter one a new line, then count, any token separator ought to suffice, so much so you'd think they'd have trained in this strategy given tokens make individual letters opaque
Strawberries aren't the weakness, the weakness is the tokenization of a prompt. Any word with multiple duplicate characters is going to be troublesome for LLMs.
I think the parent is aware and was joking.
I asked about Tiananmen Square and it said it knew nothing about it. A real life “Doesn’t look like anything to me” moment.
My God! You're practically a terrorist and should be on a watchlist! Who knows what you'll ask next? "Who is the surgeon to the boy?"
Interestingly the reality of the open source release is they are opening up to full distillation by the closed source model providers at a deeper and more fundamental level. If anything the open sourcing will help Anthropic and open ai ladder up faster. Open source has always been about mutual cooperation towards a goal and has never closed the door to commercial success. All the hand wringing about open weight models putting closed providers at a disadvantage doesn’t get what working in the open actually does for commercial interests - it is like science in the open - it enables and lifts all boats. Likewise commercial success doesn’t close the opportunity for competition or more open source work - it’s the economy of activity and competition that matters overall. When things stagnate is when closer concerns turtle up and collude on not competing for each others turf.
The future is good and better for everyone the more work is in the open and the more work is in the commercial space. It’s good all around.
I'm happy that Kimi K3 is indeed SotA and its open weights are due to be released soon.
It's also true that Moonshot and other labs distill from Claude. This has been reported on extensively. I don't think there's any alpha for Anthropic distilling from this model. I do not mean to discount the tremendous amount of innovation regarding MoE and quantization that Moonshot has accomplished. But its training with synthetic data is in large part from distillation from frontier labs.
When I say distill I also mean mine it architecturally for insights but I doubt seriously the model training is entirely distillation of Claude, it’s almost certainly a mixture of both original corpus and reinforcement as well as distillation. I think it’s a little condescending to imply that these new open models are cheap ripoffs with nothing original to them. These teams and labs are top tier as well, working under unreasonable constraints imposed by the USG. That’s a powerful combination for creativity.
I think your characterization of my post, which credits Moonshot's innovation, goes a bit too far.
The closed labs don't really benefit unless the open model has something extra they don't have though. Meanwhile the open model dilutes their customer base and seriously cheapens their offering (which is a heck of a good though IMO).
Because the open model might have been distilled from outputs of the closed model doesn’t mean they are architecturally equivalent. There is almost certainly innovations in architecture present in the open models that the closed labs didn’t think of. It also is almost certainly true that they aren’t completely built out of a distilled corpus, that reinforcement is equivalent, etc. Therefore closed labs will also benefit from being able to inspect in totality the architecture, activations, weights, and be able to train against it at scale in an ensemble of other models and their internal work.
The only way there is no benefit would be is if the open models are literal copies of the closed model, which unless there was direct theft, is highly improbable.
For regular chat users it's $19 while Claude is $20...
$20 users don't get access to Fable. It's $100+ tier only.
True. However, I believe most non-programmers don't need access to the fanciest model, but just want to use a good LLM without constant nagging about usage limits. Then 19 vs. 20 is true?
Still no. If you're only getting Opus-class you can still end up paying less by just using an equivalent Chinese model on OpenRouter.
Not talking about me or HN users in general. These companies likely need regular peeps to begin using their services in order to become profitable. Not sure these are the ones who will buy from OpenRouter.
No "19 vs 20" is not true for "Kimi vs Fable". It might be true for "Kimi vs Opus" or whatever. Anyways looking at pricing plans is not a good way of comparing the price of LLMs. It makes more sense to look at cost per token:
I'm a $20 user and was given $100 extra usage credit today, "for Fable" (I'll be sticking to Sonnet and sometimes Opus TYVM).
What to do with $20 on Fable? One prompt or two ?
From the blog post moonshot refers to it as open source but only mentions releasing the weights.
The weights are the "source" of a model.
If weights are the source for models then ELF binaries are the source for software.
clearly not true. the weights are the preferred form for making modifications. Do you really think people should be downloading hundreds of TB of training data and running make to build the model on their own cluster of GPUs?
Random people? No. Governments and big corporations? Yes. It removes any concern of "backdoors", and is currently the best starting point for your own model which will be as capable as k3.
Everything is open source if you know how to reverse engineer ;)
Then everyone is open source if you have $20 to spend on tokens, I guess.
The weights are the output artifact, the training corpus and system are the source.
Kimi publishes their code and a technical report on their methodology. But I think the weights are still important. It means anyone with the resources could run the same model on their own hardware.
The training corpus + training code is just an automated editor for the weights. It would be like requiring the source code for Visual Studio for software made within it to be open source.
Weights are not an output artifact no more than source code is. There is never a moment where you can claim that it's done. As requirements change new ways to change the weights / code come up. With different projects you might want to import the weights / code into a bigger model / codebase.
I want to caution against this line of thinking. NSA's fast16 program silently altered data during nuclear simulations, and that is also possible within open weight models. There is no reason to think there is anything like that currently, it also cannot be dismissed. Open source would include the training data so you could create a similarly capable model.
Training runs are not deterministic. Not just in the case of floating point not being associative, but nodes have hardware issues and go down and back up at various points through training. Problems get encountered and then various parameters get changed during the middle of a training run.
Making the process of creating the source code repeatable goes beyond the idea of open source. Open source is about being able to work with the code, not recreate it. It doesn't require domain experts to document every single thing they know. For example look at some of the GPU drivers in the Linux kernel. The GPU is not properly documented and there is trust that vendors are implementing things correctly.
> Training runs are not deterministic.
I agree which is why I said you could make a similarly capable model.
> Open source is about being able to work with the code, not recreate it.
You can also use a hex editor to modify compiled binaries, but no one would consider that "open source".
>You can also use a hex editor to modify compiled binaries
It would be like if someone gave you a prompt you could pass into ChatGPT to produce the entire Linux kernel. While yes you could modify that prompt and then spend a ton of money on inference to generate millions of lines of code which hopefully are equivalent to the original Linux kernel it would be easier if you just had the source code to work with where you can make a simple extension. If you want to end a model you want the weights, and not the entire setup to generate it. The weights are the starting point that you can work off for training a new model similar to how the source code is the starting point people want to work off of.
So far all the Kimi models have open sourced their code, their weights, and published technical reports explaining their training methodology. I expect the Kimi K3 Technical Report will come out July 27 and they usually publish the code and the weights alongside that.
There is too much business risk in running on an American company. They can pull the model back or lobotimize it. No Alex Karp fan, but he was right: companies are worried about hyperscalers stealing their alpha. The lack of guardrails, data security and control really give these open models the edge. If you factor in that they appear to cost less, the hyperscalers are in big trouble. I don't see how this works out. Software was always supposed to be deflationary and collapse down to 0 marginal cost but this isn't remotely the case. It's truly amazing to see the state of open weight models
If only you could run K3 locally that would be the magic bullet to make it a true magic bullet!
Until Kimi 4 comes out and you'd be like "if only you could run Kimi 4 locally"
Everyone wants the latest and greatest.
until they put their money where their mouth is
You can! It just might be a little bit outside your budget.
I enjoy all the fun of these new models as much as the next guy but I truly don’t see a circumstance in the near future where my $200 a month with the frontier labs doesn’t get me more than enough consumption of what I need. Local models, chinese models, etc are all very fun weekend projects to tinker with but until something changes (entirely possible!) with how much you get with one of the subscriptions I just don’t see why I would move. What am I missing? Is it simply that a subscription is no good for production use cases? I kinda feel the same way with choice of coding harness, openrouter, etc. why would I use anything other than frontier if I don’t have to pay any more pretty much no matter how much I use? pls tell me if I am holding this wrong haha
If everyone took your position, then frontier labs can keep raising their price.
Kimi K3 showing competitive performance with Fable while both sitting at the SoTA level on fireworks.ai is a huge milestone. Really interesting to see how the landscape is shifting here.
for people in mainland china the only option now is Kimi K3+GLM 5.2 for their day to day work as Fable is blocked. for me i use Fable + codex when i was out of mainland china, it is good when you have options wherever you are in this world
The article is about the best you could theoretically do with a perfect router. The takeaway is that trying to build a good router is worth doing. But it's unlikely to be a perfect router.
I have several thoughts about this, which I'll just iterate:
1) US export bans have made it so that Chinese companies have to compete using less-than-state-of-the-art hardware. This has forced Chinese companies to build more cost efficient models. Whereas, US companies have moreso tried to be state of the art by spending more money than anyone else on state-of-the-art hardware.
2) Xi Jinping has called for more open AI models (not to be confused with the closed models of OpenAI), and I'm happy to see a powerful world leader advocating for open-weight AI models. Whereas, the US seems likely to just ban models.
3) My impression is that, if China surpasses the US in AI development, there will basically be nothing that the US does better than the rest of the world--except for military spending--we spend a lot, but we don't necessarily spend well (something something Iran). I mean, if the US is no longer a tech leader in the world, like... what are we a leader at? Manufacturing? Healthcare? LOL. Are we a leader in any industry or by any metric? I wonder if China is attempting to remove the last jewel in the USA's crown with these AI releases.
4) It must be refreshing for companies to have access to a new model that isn't going to get pulled because the government bans it 2 days after release. And it's open-weight so it wont go away--amazing--what a shift in the market.
5) If it becomes clear that open-weight models are the future of AI, will that pop a huge bubble in the US economy? Maybe. But, on the other hand, these companies aren't just training AIs, they are also building data centers which will remain valuable no matter what happens.
Implying that AI is the last jewel of the USA is too simplistic and ignorant.
People are paying for tokens, that will likely continue even if open weight wins.
the per-token comparison keeps missing that k3 spends way more tokens per task. if it burns 3x tokens to reach the same result as fable, cheap per-token stops mattering
They address that in the article :) to quote :-
"So where's this huge price gap coming from? token pricing, prompt caching, and effort-per-task. On SWE for example, K3 works much harder than Fable: roughly 55 turns and 1.3M tokens a task versus 21 turns and 130K. On the long terminal tasks it's the other way around: Fable is the one that spirals, running up 64 turns and 1.5M tokens (sometimes straight into a timeout).
Prompt caching does most of the work of turning that effort into K3's price advantage: even when K3 reads ten times the tokens, with cache hits that means that SWE runs still come in lower cost than Fable. There’s a tradeoff. Tasks with extra turns generally mean more wall-clock time per run i.e. slower runs. If you need an answer in two seconds, that matters; if you're running agents in the background at scale, a bill that's a fraction of the size matters a lot more."
So this makes sense for a standard SaaS app - but given that models in general perform much better with low context window usage, it probably also means that Fable is still significantly better at 'frontier-level tasks' -- hard research problems, complex geometric rendering algorithm optimization, etc., no?
If I had an oracle, why would I need an LLM?
Openrouter also features routing. Routing is indeed an option if you don’t need consistent behavior and allow switching models.
It would only be inconsistent if the router chose different models for the same task.
Considering that of 89 terminal tasks there were 11 that only Kimi got right, and 7 only Fable - having a router gives significantly more consistent results if you define consistent to mean correct.
cluade model, claude opus 4.6,7 perform very well.and stable, understandable ,i miss that in kimi, its slow, doesnt gives feels like claude
Was this an out of sample test of the router or was it trained on these specific use cases/eval suites?
Neither. They ran all the cases using both models, picked the winning result for each one, and then said "if you had a router that guessed with 100% accuracy, here's what it would have picked"
Your account <...> request reached organization TPD rate limit
For model routers, do they have to retrain the routing model every time a new LLM is released?
Maybe a continuous evaluation process, why is there anything to train?
I really like the idea behind OpenRouter's new auto-beta -- classify by task type, and then just follow what the market is using based on the last 7d. https://openrouter.ai/docs/guides/routing/routers/auto-route...
Shifted to K3 and it is like a fresh air. While Fable and Sol very good at _generating_ code i even cannot force sol to just read all relevant source files. As result it reinvent existing things or assumes too much about internals of other, which lead to incorrect uses. Even with hard planing mode GPT burned 33% of week tokens for 3 hours producing no result and even cannot find root cause, but K3 fixed it in a minutes. Hilarious that while software not started with empty cache Sol handcrafted empty cache to let it start. Having full procedure right in the MEMORY.md. K3 found this, and downloaded cache correctly. But while using gen5 models i always feeling myself ignored. Any commands, steering anything - just ignoring. At first i added lots of hooks, no "?", expect in rust code, bash hook ban for find|grep|tail, with notice to use ltsp but then it started to ignore strategically. Also whole "thinking" thing is hidden from claude. regenerated thinking summary is incomplete and not useful. While being very verbose kimi k3 doing good job providing whole train of thoughts. So looks like claude degradation started from 4.6 comes to a logical end. Maybe i missing something and changing existing code beyond bug fixes is not a way to go. But there is no currently stable way to generate code on hier of specs and lean models i'd like to.
I'm skeptical. According to arena.ai, Fable 5 dominates almost every category: https://arena.ai/leaderboard Kimi K3 has an edge in WebDev but struggles to reach top 10 in many other categories.
In my experience, Fable is not even close to Sol 5.6 High (not even the max tier) for coding.
1) it's substantially slower.
2) it's substantially more expensive.
3) it's code is considerably worse.
It's a joke when you consider what you get for what you pay for.
Your experience is an anecdote. Leaderboard rankings are a distributed blind taste test.
This is very interesting to me because I find sol to be inferior at code generation, but superior at conversation and code review.
Speed and expenses are non-factors for arena.ai.
The irony is the Chinese are being very democratic with their models, while the USA tries to do central control.
Glad to see centralized control fail on the grandest scale. Maybe we can learn a thing or two.
> Chinese are being very democratic with their models
They can afford to, assuming they distilled Fable, because that reduces their pre-training costs, doesn't it?
So if they paid for pre-training they would not be able to open their weights?
Not if they want to recoup their investment.
They distilled Fable, in the couple weeks it was available?
The idea that the Chinese labs cannot make progress except by copying superior American products is just prejudice against the former and exceptionalism of the latter at play. Even the OpenAI top brass have admitted otherwise [1]. China is an equal match in every respect, and we'd better admit this to ourselves sooner rather than later so as to see the game clearly.
[1] https://xcancel.com/deanwball/status/2078133895766114412
Yes, distillation is very very easy.
What you're seeing is the distilled (no pun intended) result of realpolitik at play. China's labs aren't as open as they are out of altruism, but rather to capitalize and undercut the monopoly held by U.S. competitors.
They wanted capitalism, no? Here's the competition in the market.
Yes, hah. Historically, the US got its moment to define for everyone what the rules of capitalism and globalization are, and China then played to win.
Is kimi making a profit? How subsidized is it by the Chinese government?
Is that any worse than being subsidized by VC? They're playing the same game as the western labs, just with slightly different players.
Chinese government vs American VCs doesn't equate to "slight different players."
True. The Chinese Communist Party's motto is "Serve the People". [0]
They might not always live up to that, but at least the aspiration is there.
No such pretense from American VCs.
[0] https://en.wikipedia.org/wiki/Serve_the_People
American VCs and American Government mingle at the deepest levels. Yes, slightly different is correct.
This is pure speculation. China labs are geeks and nerds like some of us. They were amazed at LLM and like everyone want to build their own. They were happy to get meaningful next token predictions. I think most of you forgot how bad these things were 3 years ago compared to today. There was nothing to undercut, it was just geeks putting out their toys and saying, "Look, I built something cool". That became the culture and led to were we are now. All this idea that they are trying to undercut the monopoly is speculation. Google still releases open models, Cohere releases command-a, Mistral releases their model too, Arcee and Thinking Machine have released models too. It's just that Chinese models have gotten good and are also leading in the open weight category and USA has a very strong paranoia of Chinese models hence this talk. Plus every time they release something, some of the labs cry, "they copied us, distillation, the bad guys have done it again!"
So what we are seeing is nothing more but geeks doing geeky stuff, the only one that has serious demonstrated that there are possibly be had is Anthropic and the Chinese labs are beginning to copying them in terms of user plans, coding tools, etc. They are giving away the model to show it's great. Once they have the compute to serve the world and they have a model just as good or better than top model, they will go close to keep all the profit.
You are ignoring the US and EU politics about regulating access to and distribution of and legal use of models. From my admittedly slightly limited perspective, it looks like China is politically more laid back about those developments than the western nations are.
This is true, but when we compare them to OpenAI's repeated feints at altruism and nonprofit structure, the irony is plentiful.
> The irony is the Chinese are being
There's no irony. The problem is these generalizations and labelling that's hurting everyone. It's the wrong way to view China and maybe many other places.
Just like the definition of AI, the definition of democratic or not evolved long ago.
I'm reminded of space race, where US built centralized planning, while USSR did competition between construction buros
The Kimi K3 felt pretty good when I tried it out.
we put a router model in front of two other models so the router can decide which model is better at deciding things. next we'll need a router for the router and eventually the entire internet is just routers routing routers to other routers
And the weights are just in the routers now? Love that idea.
The weights are in the meat
the weights are the bones
The bones are their money
Is reddit down today?
The line separating HN and reddit runs through the heart of every man
They're Made out of Meat: https://www.eastoftheweb.com/short-stories/UBooks/TheyMade.s...
It’s routers all the way down
It's a series of tubes
It's all pipes, Jerry.
Jerry, these are load bearing routers.
"Information Retrieval is where it's at" Buttle said to Tuttle.
what are my options if i want to use a router like this ? who provides one ?
Oracle routing is by definition not possible.
Two that I use are:
* Openrouter.ai for a hosted router
* https://github.com/diegosouzapw/OmniRoute for a local router
They forgot to compare and incorporate GPT-5.6-Sol.
It is not SOTA. Give me a break. Sure, run it on Cerebras to get speed but that’s pretty much its advantage.
Why do you think it's not SOTA?
Cerebras does not share the quantization of the models so you don't know if you're getting real K3 or k3 lite or something else.
It most likely will be quantized. A cerebras wafer only has 44gb ram, and linking them together vastly reduces the speedup.
Agreed. The benchmark closest to my experience is FrontierMath Tier 4. Fable and Sol (90%) are very far ahead of Kimi K3 (not even 40%). Kimi is trained heavily to basic agentic tasks, like all the other open models right now.
Why SoTA (uppercase “T”) instead of SotA (lowercase “T”) ?
“State of [T]he Art” versus “State of [t]he Art”.
If not SotA then at least SOTA, which is more accurate.
It should be SotA.
Does that make DeepSeek V4 Flash MiniSotA? This dev in the Twin Cities would like to know.
I think they call it Minipop there
An cross of SoT (source of truth) and SotA (state of the art). SotA looks weird, admittedly.
“SoTA” as “source of truth agent” would be a pretty annoying/funny acronym to unless upon the world.
I suspect typo.
nice to see humans writing
SOTA police
its sota like it was a joke
probably buckeye fans
“State of THE art.”