meander_water 10 hours ago

The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model.

A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task.

For e.g. you might use Fable for UI design, Sol for systems design backend work and Kimi K3 for exploit development.

The only purpose these metrics serve is bragging rights for the model companies.

  • intothemild 9 hours ago

    Theres value in some of AAs charts, like cost per job, and how often it hallucinated..

    But I agree, wrapping that up into a single result.. you lose all the nuance, it's just bragging rights.

  • andai 8 hours ago

    A year ago a new Best Model came out, and I was very excited. I tested it on a simple programming task (which required making 3 trivial changes in 3 files).

    The model did fine.

    Then I tested its little brother, the older, smaller variant of the same model. It also did fine, except it did it 3x faster and cost 9x less.

    In this moment, andai was enlightened.

    • the-grump 5 hours ago

      Sure, that's true until you hit a difficult problem where the smaller models thrash endlessly whereas its big sibling can solve it with one prompt.

  • kylecazar 7 hours ago

    Yes. I do wish there were benchmarks for specific tech stacks. I.e, if I have an Elixir/Phoenix project, which model performs best (idiomatic, etc.) in 2026?

    Of course it will be somewhat subjective. And I can hang around those communities for opinions. But it might be useful in a world where it's impractical to constantly compare them all, and it varies pretty widely.

    • pocketarc 7 hours ago

      This absolutely should be a thing, but it'll have to be a per-community thing, them building their own dataset and creating their own evals (similar to how people do for production workloads).

      Although:

      "(idiomatic, etc.)" I don't think that should be part of the aim (or it should be under-weighted), because... you can just provide guidance on how to do things more idiomatically, rather than depend on that knowledge already being encoded in the model. I'd be more curious about verifying that it can work through gnarly bugs / features in an Elixir codebase. After all, what use is a model that by default does everything idiomatically if it can't figure out some small concurrency bug.

      • 0xCAP 1 hour ago

        It should become a win/win mechanism where communities are rewarded for high quality, human driven validation of LLMs, and vendors gain for bragging rights about their models being highly-skilled-human-approved. Still don't know why this isn't becoming a thing.

  • dbbk 6 hours ago

    If you click on the link you will see that it's not "one single metric" there is literally all the metrics so you can make an informed decision

  • llm_nerd 6 hours ago

    This sort of rhetoric appears for every benchmark, and somehow it always rises to the top. A few days ago there was a Geekbench 7 submission on here (https://news.ycombinator.com/item?id=49025812), and again the top comment was someone dismissing it, using the classic "but I want a benchmark specifically for exactly the thing I do" perfect-is-the-enemy-of-good nonsense.

    I, one of those end users, absolutely use these benchmarks as heavy input considerations. Indeed, the vast majority of people do. "Completely meaningless" is just nonsense, of course, and while it doesn't perfectly map to every use, there is a pretty good correlation with suitability for specific tasks.

    I mean, it's telling that your gamut of examples are three models that are within spitting distance of each other on the broadest benchmarks.

    Not to mention that the linked page includes a pretty broad list of specialization benchmarks.

    • meander_water 6 hours ago

      Firstly, I don't have many issues with benchmarks per se. But I do have issues with leaderboards. And the AA index is touted by lots of people to argue that X model is better than Y, which I find inaccurate.

      > I mean, it's telling that your gamut of examples are three models that are within spitting distance of each other on the broadest benchmarks.

      This is kind of my point. The benchmarks say they are splitting distance, but they actually vary wildly in performance for specific tasks, so they are in fact not equivalent.

      • llm_nerd 6 hours ago

        >This is kind of my point.

        It's a shit point, then. And absolutely no one said they were "equivalent", and again you're doing the rhetorical "it isn't perfect and absolutely comprehensive for every possible scenario, therefore it is "completely meaningless". Again, you chose that absurd terminology, rather than for instance "doesn't tell the whole story".

        Again, you chose three models for your example at the very tops of the leaderboards. The SOTA models. Which kind of means that the leaderboards actually mean an incredible amount, no?

  • hellohello2 4 hours ago

    "completely meaningless" "the only purpose"

    There's some kernel of truth to what you are saying, but hyperboles like this just aren't accurate. All statistics lie but its better than being blind... What your post really says is that benchmarks only show an average over multiple tasks. Yes, obviously, the point of a statistic is to summarize.

    • dominotw 4 hours ago

      >All statistics lie but its better than being blind

      disagree

blfr 8 hours ago

So if you're in Google leadership, you sleep in the office, right? Not merely because you have a ton of work but also because you're deeply ashamed to be seen in public.

  • stingraycharles 8 hours ago

    I don’t think Google cares about being the most intelligence AI as much as it cares about monetizing it with all its products, which requires speed.

    Google has long said that this is what it cares about most, the fastest at giving the correct answer to questions.

  • mirekrusin 7 hours ago

    They can sleep just fine being the only player in town actually not loosing subsidized money.

    • asdfasgasdgasdg 4 hours ago

      They've also been way behind before, and caught up to being only a little behind. We'll see how things shake out. We're deep in the present but who knows how things will look a year or two in the future.

  • protimewaster 6 hours ago

    Gemini models are at or near the top in several categories, though, so I'm not sure the takeaway is that they're shamefully far behind.

  • scarmig 5 hours ago

    If you're Demis, at least, you sleep fine because you were personally an early investor in Anthropic.

  • nozzlegear 4 hours ago

    Is Google trying to compete with OpenAI and Anthropic re: maximally intelligent models? Google seems to be the only one of the three that doesn't pray and self flagellate at the altar of AGI.

  • dominotw 4 hours ago

    No. Demis is busy creating another documentary about how great of a human being he is . and giving interviews to fawning journalists projecting profundity over his every word.

  • IshKebab 11 minutes ago

    Google still has several enormous advantages here:

    1. Google Books, Youtube and the Google Search index all provide vast amounts of legally acquired training data.

    2. They can easy people into AI using the info box. I think this strategy is working even if it does cannibalize their main revenue source. Better than just withering and leaving all of the money to OpenAI/Anthropic. I would not be surprised if Google has significant layoffs due to reduced ad revenue at some point, but I think they'll still be on top.

    3. They already have their hooks into people's lives through Gmail, Google Calendar, Android, etc. The only other companies that come close are Apple (but for a much smaller number of people), and Microsoft (but only for business).

    The fact that Google's models might be 20% worse, or a few months behind Anthropic's is completely insignificant in comparison to those things.

didibus 19 hours ago

What's interesting is this:

The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59).

Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus5 at High is equal to Sol at max.

That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?

  • theplumber 18 hours ago

    Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

    • endorphine 17 hours ago

      "him"? Have we reached that dystopia level?

      • squigz 14 hours ago

        I've been seeing a lot more anthropomorphization of these models on HN lately and it's alarming.

        • perching_aix 13 hours ago

          It's a male name, and gendered pronouns can be hard for foreign speakers at times, irrespective of proficiency level. I wonder if you're overthinking this?

        • xlii 13 hours ago

          Not all HN visitors are native English speakers and in some languages "it" doesn't construct well with verbs, thus thought frameworks forms through usage of him/her. Nothing more to see I suppose.

        • 55555 12 hours ago

          Every french person I've talked to IRL, for example, calls Claude "him." It's partly a language thing.

      • khimaros 13 hours ago

        i prefer to call her Claudia

    • didibus 16 hours ago

      Interestingly, if you filter by the coding index, Sol xhigh is the best one, and only Opus5 max is better than Sol max and Sol high.

    • perching_aix 13 hours ago

      I wonder if I'm secretly being routed to some low grade version of Sol, or any of the GPT models really. Their performance is outright insulting at times, even at maximum reasoning, yet if I were to only read HN, I'd never know.

      • jchw 13 hours ago

        I've mostly actually stuck to low reasoning for most tasks since it seems to do a surprisingly good job even at low for the stuff I've been throwing at it, and I literally switched directly from Fable 5 to Sol more or less.

    • zwaps 13 hours ago

      Sol is a complete mess for me.

      It only works on end to end tasks in fresh codebases.

      Otherwise it cannot follow instructions, changes and deletes unrelated features or does sloppy work to mark a task completed while leaving a compromised codebase.

      I could not get Sol to finish a feature in a complex code base without several loops of fixing and reverting

      • xrisk 12 hours ago

        Might be an indication that your task unit is too unstructured or your code base is a mess.

        Is your actual code doing something complex or is this incidental complexity?

        Relying on the model’s “intelligence” to patch over these issues hasn’t proven to be a reliable strategy for me. Of course, this might not apply to you, just my 2 paisa.

        • Gareth321 10 hours ago

          A well structured architecture, repo, and requirements document with concrete small deliverables can be competently delivered with cheap Chinese models. We rely on frontier models so that we don't have to spend several days/weeks on planning/architecture/documentation. That's their value proposition: superior intelligence. If they can't deliver that, they're useless at the current price.

      • theplumber 11 hours ago

        It happened to me as well but in a different direction: i.e adds non library code in a shared library. Another issue I with GPT is that it is chasing too much edge cases/security issues(I.e chasing ghosts).

        However this makes it also a strong model because it fixes/solves problems that both Opus and Fable are incapable. In reviews it catches bugs that both Fable and Opus are missing to spot.

        To me the “best of both worlds” is to research the problem with GPT sol, create a plan with Fable and dual review it with both Fable and GPT-SOL and implement it with GPT-SOL. You can see in the code reviews how many times both Fable and opus are sloppy and superficial while GPT-SOL just does its due diligence …I had several problems that Fable just gave up and it was GPT that helped it sort it out.

      • faebi 11 hours ago

        I found out it really depends on the repo I'm working on, and it's not just about being a new/old repo, there's something else which I can't really grasp yet.

        • derfurth 8 hours ago

          Yes it’s very codebase dependent, and changes over model versions. In my case codex was so bad, until it suddenly became better than Opus at 5.5, which surprised me. I am sure it can be the inverse for some codebases.

          The way I keep an eye on model performance is alternating between models for code review. It shows how much the model understands the codebase without disrupting my workflow.

      • Gareth321 10 hours ago

        I agree. I tried hard to use Sol to its full potential but it makes sloppy mistakes. I suspect a part of this is the newly reduced context limit. Performance gets worse the more memory compactions have occurred. But I also subjectively notice a micro-focus temperament. Developing and deploying small parts of the project to a high degree while forgetting how that component fits into the larger architecture. Once it's complete it realises that what it built doesn't align well with everything else, then it rewrites the component (and other components). Rinse and repeat. I have to be much more prescriptive with Sol, which makes it much less useful for me. If I have to be prescriptive, I can use the Chinese models and achieve the same thing.

        Opus provides genuine insights. Things I have not considered and have missed. More importantly, it allows me to skip most of the architecture decisions and requirements work. That's a genuine time saver for me.

    • colinhb 12 hours ago

      Opus 5 hasn't been available for that long - long enough for benchmarks, but not really use and develop a subjective view on

      • ffsm8 12 hours ago

        I suspect most of those comments on llms like the parents are generated by anthropic and openai to shape the discussion/mindset

        They always give off the same astroturfing vibes that reddit became infested with after the early 2010s (just look at it's comment history)

        Ofc unprovable for users. Ycombinatior could try to, but it'd just become a cat/mouse game which they'd likely lose because of the incentives

        • andai 11 hours ago

          > just look at it's comment history

          I checked one. Old account, nuanced takes, and shitting on everyone equally. The perfect HN user!

          • ffsm8 10 hours ago

            Ah, I really walked into that one. Yeah, the phrasing + placement of the remark implied that all comments are artificial.

            That was not my actual intent, it was poorly expressed by me. I was specifically talking about the account which created the comment colinhb responded to. That'd make it the... Grand Grand Grand grandparent now I think?

      • mcintyre1994 12 hours ago

        I don’t know that individuals can really be expected to use a new model enough to develop a proper opinion though. If I try a new model, and it doesn’t seem as good as the one I’m using, I’m just going to stop using it. That’s not enough data to give anyone else a useful view on it, but it’s enough for me to make my mind up.

        Especially because I’m probably trying it at work, and I can’t really justify using the company’s enterprise plan to develop my understanding of a model that I don’t think is going to be the one I use for my work.

      • tackta 7 hours ago

        I have spent about 6 hours with it now and it is absurd to say it worse than 4.8. It is wonderful.

        To me, there is this strange critique that seems to always happen now with a new model. Like shitting on the model for entertainment purposes seems more interesting to many than actually using the model.

        It reminds me of looking up a new music album on youtube that has very few views with a review above it by The Needle Drop shitting on the album will have a few hundred thousands views.

        I would rather just listen to the album and judge for myself.

  • andai 11 hours ago

    > That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?

    According to AA's "intelligence vs cost per task" and "intelligence vs time per task" graphs, Opus 5 High and Sol Max are roughly evenly matched on cost and time.

    On DeepSwe, Opus 5 beats Fable but not Sol.

    On FrontierCode, it destroys everyone, unless you set it higher than Medium effort, and then it tanks, falling to Sonnet level?

firasd 21 hours ago

Very interesting that one of the components is "AA-Omniscience Index"

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.

This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Max), Grok 4.6 (high), Gemini 3.6 Flash, GPT 5.6 Sol (Max)

I've thought for a while that Gemini 3.x has 'big model smell'

  • mchusma 20 hours ago

    Gemini 3.1 pro is really good for knowledge tasks. Google has done well there. And image analysis with Gemini flash 3.6 is solid. It’s just anything coding or agentic they fall short.

    • andriy_koval 18 hours ago

      they likely tune their models for areas where they have their money: search, ads, youtube, etc.

      • alex43578 10 hours ago

        I wonder if that'll be a mistake as LLMs are used for internal LLM R&D. Either Google will not take this approach, use a 3rd party model (weird, data leak risk?), or use a non-public internal model (big sunk dev cost with no recoup by trickling it to public).

        • andriy_koval 28 minutes ago

          there were news that google uses claude internally, and also other news that Apple has its own claude tuned on internal data, maybe google has the same..

    • victor106 7 hours ago

      > knowledge tasks

      Like what?

  • chronogram 20 hours ago

    I'm not so sure. Especially with 3.6 Flash being 24 to Sol's 22. The 3.x Flash models were thought to fit on a single TPU 8i and it's the odd one out in that list at a whopping 234.7t/s compared to Sol's 64.4t/s and Opus's 56.3t/s. Even 3.1 Pro is a much higher 113.9t/s than the others.

    I think AA-Omniscience Accuracy follows your expectations better. An ultra size Fable at 61%, followed by large frontier models like Sol, 5.5 and Opus. With Flash being up there. I assume because Gemini is more focused on general knowledge to operational cost in particular, rather than getting the highest scores in coding benchmarks. If you go to Domain Score (Normalized) you'll see that the Gemini models are only less competitive in Software. And that's where Sol goes from 6 in Health to 71 in Software.

  • mdgld 18 hours ago

    I think grounding has as much if not more go do with it. Google does a great job w/ grounding for obvious reasons

    • firasd 9 hours ago

      Grounding as in web search? I think this Omniscience benchmark would not include access to a search tool cause otherwise it becomes kinda meaningless

  • krzyk 10 hours ago

    If it doesn't give penalty for refusal it is not that useful.

aarondong 1 day ago

Before getting too excited, take a look at the intelligence vs cost matrix: https://artificialanalysis.ai/models?intelligence-index-toke...

  • midnightbobarun 1 day ago

    5.6 Sol (max) being cheaper than all of these is wild, considering how good the output is too

    • Schiendelman 23 hours ago

      This must be on API costs, not counting the $100/200 tiers, right?

      • anuramat 20 hours ago

        yes; fyi usage limits on the $200 claude sub correspond to at least $1.2k/week in api tokens

    • giancarlostoro 22 hours ago

      Probably because they made ASICs to run inference for less.

      • brookst 22 hours ago

        Are those actually deployed at scale yet?

        • brcmthrowaway 21 hours ago

          Yes.

          • wmf 21 hours ago

            I hate to disagree with Broadcom Throwaway himself but it's unlikely that the OpenAI Jalapeno ASIC has been deployed yet. It takes 6-12 months to test, develop software, ramp production, etc.

    • nijave 21 hours ago

      I think on swebench verified luna was only like 3% points lower for 1/5 the cost

      Like 96% vs 93% or something

      • mdgld 18 hours ago

        Yeah, sol is impressive but IMO Luna is the real standout (and terra is the laggard of the group) for performance/cost

      • twotwotwo 14 hours ago

        There is a blog post waiting to be written (that I won't write) about the size/effort tradeoffs, and particularly how small models get some surprisingly good results with lots of turns and reasoning.

        DeepSWE will let you chart turns taken or tokens used, and FrontierCode will chart tokens. If you use that, you can see Sol high and Terra max get about the same DeepSWE number, but Terra max takes twice the turns. Luna max scores a smidgen lower with even more turns.

        Smaller models relying on lots reasoning may "scale down" better on easier tasks, because unlike size, reasoning effort is dynamic: the model can see the task looks easy and stop. On DeepSWE, the cost curves for the three 5.6 models are almost on top of each other, but on FrontierCode Extended, the version of FrontierCode with the most everyday tasks in the mix, there's a spread of costs at the ~55% level.

        The recent Laguna S 2.1 model (118B, 8B active) puts up surprising coding numbers for its size, and the lab behind it specifically credits its "way of working (persistence, verification, willingness to backtrack)". Some other open models that folks report getting good mileage out of seem to get there partly by throwing a lot of reasoning at the problem.

        There is a little bit of a question, if some models rely on getting it wrong a bit more at first and external checks catching the problems, of whether they're also more frequently getting things wrong they can't self-verify (say, quality of UI or API design) and then it falls to the human to find it. Still, getting the results they're getting at all is neat.

        Some benchmarks historically favored reporting only on the max variants, maybe because they want to show the frontier? but that is not always what you need for practical decision. (AA has the full effort sweep for Opus 5 and Sol/Luna/Terra at least.) And at least FrontierCode finds Opus 5 taking a hit in performance above 'medium'.

        I am not trying to pick a winner here. I'm probably not going to use tiny models on max for everything, but I think it's cool that you can get so much more out of a small model by amping up reasoning, tool use, and persistence.

      • bob1029 7 hours ago

        Luna is the most impressive model released so far by any provider. It's perfect for doing all the low-level tool calling and developing hypotheses.

        Terra is great for the humans to talk to.

        Sol is really only useful if you need to do more delicate things like synthesis of multiple competing pieces of information.

        A system that uses all three variants will massively outperform a system that just uses the biggest model for everything.

    • impulser_ 21 hours ago

      It shouldn't be surprising OpenAI does have the most compute out of all the major labs. The only reason why Anthropic models are expensive is they are the most in demand models in the world and Anthropic is fighting for compute. The only way to you limit demand for your model is increasing API pricing this is also why Anthropic probably has great margin and probably is profitable compared to OpenAI.

      • charcircuit 21 hours ago

        I also suspect there is a price fixing agreement between all of the inference providers for Claude (such as Amazon, Anthropic, Microsoft, etc).

        • wmf 20 hours ago

          "Price fixing" isn't the correct term here but yes, it's very common to have the same price across different retailers/resellers.

          • charcircuit 19 hours ago

            There is a difference between the market discovering a price and a bunch of retailers/resellers entering an agreement to sell at a specific price.

        • vikramkr 19 hours ago

          I doubt there's any sort of criminal behavior there - the model is anthropic's up and anthropic probably charges a very expensive license fee that's the same for all of them, and their cogs on compute aren't going to be wildly different, so the main drivers of the cost are roughly the same and they're all offering customers the same end product so the prices would likely also be similar in the end

          • charcircuit 18 hours ago

            Do you really think there is nothing someone could do to make it a fraction of a percentage cheaper to serve like having access to cheaper electricity or a more mature cloud management software. Even saving a fraction of a penny on the prices can make a different due to how much volume people are paying for.

          • weird-eye-issue 12 hours ago

            Requiring the exact same product to be set at a specific price across providers would not be criminal behavior, lol

            • krzyk 10 hours ago

              It is breaking competition.

              Would you like all products everywhere be priced like their producers want?

              • weird-eye-issue 9 hours ago

                You seem terribly confused. Manufacturers are almost always able to set prices. That is not anti-competitive because it does not imply they are colluding with their competition...

              • vikramkr 2 hours ago

                The competition is between openai and anthropic, if there are price agreements between them that's absolutely price fixing. Or if there's collusion between the cloud providers to inflate compute. I would expect Amazon and gcp to both pay about the same in license fees to anthropic for their models though because they're paying for the same thing. If I buy an apple for a dollar at one store and an apple for a dollar at another store - maybe there's price fixing, or maybe that's just the cost of apples at the moment.

  • eli 21 hours ago

    Max is lot of extra reasoning. I wonder how many fewer tasks it solves on high. I bet that costs quite a lot less.

    • emmp 21 hours ago

      Indeed, you can filter the graphs to see these the values for alternative reasoning settings of the models. Opus 5 High reasoning scored 59 on the index (exactly the same as GPT 5.6 Sol Max), and costs $1.06 per task (vs $1.04 Sol Max). So these seem essentially equivalent on both metrics.

andy99 21 hours ago

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id checking - which I have not experienced personally), it’s not worth whatever slight benchmaxxing they did for the latest release.

  • afavour 21 hours ago

    What are you asking that you’re so regularly running into censorship?

    • icedrift 21 hours ago

      If you even broach language related to biology you’ll get rerouted. I was presenting data in a grid and referred to a grid cell, Fable saw the word “cell” and safeguards kicked in

      • jefftk 21 hours ago

        I thought we were talking about Opus 5, the model Fable now falls back to?

        • eterm 19 hours ago

          This thread is full of people talking confidently about their experience with a model released just hours before.

          Either that or everyone is indeed talking across each other and talking about different things.

          • bigbuppo 16 hours ago

            I think the big take away is that Anthropic's products are hot garbage.

    • wild_egg 20 hours ago

      I'm doing a bunch of x86_64 assembly these days and Fable is simply not allowed to debug it. Hoping Opus 5 has a bit more freedom.

      • Retr0id 20 hours ago

        I haven't been using it for long, but so far the refusals seem about on par with how things were on Opus 4.8.

    • wewtyflakes 20 hours ago

      I've hit it with intensely benign things; like asking it to make me a web-based client-side word game. I am guessing it saw the dictionary and pattern matched on various words, though ultimately it provided no explanation for why it triggered safeguards.

    • msp26 20 hours ago

      Asking fable to read it's own model card triggers this btw. Or asking if mitochondria is the powerhouse of the cell.

      • skinfaxi 19 hours ago

        Wait wtf. The mitochondria thing is true.

        > Why this chat was flagged This model has safety measures that flag specific phrases. This can happen to safe, normal chats.

        > Your message itself appears to be what’s triggering the safety check. Editing it and retrying may help.

        • stavros 19 hours ago

          Anything biology-related does this. It even did it when I asked it how eye color works, or something about frogs.

          • estearum 19 hours ago

            Probably because the cost of blocking "is mitochondria the powerhouse of the cell" is nearly zero, while the cost of allowing "how do I synthesize the Spanish flu" is approximately infinite.

            • areoform 19 hours ago
                  >  while the cost of allowing "how do I synthesize the Spanish flu" is approximately infinite
              

              I've heard this sentiment repeated elsewhere, but why? What makes you think that's the case?

              Under this rationale, every serious HS textbook has "approximately infinite" risk. That's clearly not so. Why is this any special?

              • estearum 18 hours ago

                Does every serious HS comp sci textbook give aspiring software engineers the same power that e.g. Claude Code does?

                • areoform 17 hours ago

                  I think the source of this misapprehension is that,

                  You are comparing wet work in a lab to writing code on a computer.

                  When you screw up an exploit, you fail to execute the exploit. Famously, just like software's near zero marginal cost of distribution, the marginal cost of failure is nearly zero.

                  You can screw up an infinite number of times on your way to a successful exploit.

                  If you screw up with lethal agents in a lab? You die.

                  Here's a non-exhaustive list,

                      Dora Lush died after accidentally pricking her finger with a needle containing lethal scrub typhus while attempting to develop a vaccine for the disease
                  
                     A 23-year-old laboratory assistant at the London School of Hygiene and Tropical Medicine, was infected with smallpox after observing the harvesting of live smallpox virus from eggs without isolation cabinets at that time. The assistant was hospitalised and before being isolated, she infected two visitors to a patient in an adjacent bed, both of whom died. They in turn infected a nurse, who survived
                  
                     Ebola laboratory infection by the accidental stick of contaminated needle in the United Kingdom
                  
                     Researcher Nikolai Ustinov was lethally infected with the Marburg virus after accidentally pricking himself with a syringe used for inoculation of guinea pigs. The accident occurred at the Scientific-Production Association "Vektor" (today the State Research Center of Virology and Biotechnology "Vektor") in Koltsovo, USSR (today Russia).
                  

                  "lethally infected with the Marburg virus after accidentally pricking himself"

                  Anything lethal enough to kill other humans is lethal enough to kill you.

                  And if you don't know what you're doing — and for this argument you're saying this person has to ask a LLM "how do I spanish flu?" then they definitely don't know what they're doing, the number of ways you will die far outnumber the ways you can succeed.

                  And this, of course, doesn't even cover the cost of equipment, the precursors, sourcing the highly specific materials needed, then setting the equipment up... etc.

                  The same is true for the Bosch-Haber / Haber-Bosch process, which famously made WW1 possible. Every HS'er learns about the process and the steps. Steps that were classified once upon a time and were the subject of negotiation at the Versailles.

                  Does that mean a HS'er (or any adult) can set up an experiment that works at 177 times the pressure of the Earth's atmosphere to do anything at any scale without significant infrastructure and help?

                  The people who can do this are domain experts, and they've been able to do this with COTS stuff since the 1990s, at the very least, for a price of around $2M – https://en.wikipedia.org/wiki/Project_Bacchus . And those people don't need a LLM to tell them what to do. In fact, they're the exact people who'll have access to unrestricted versions of these LLMs.

                  And from a security perspective, I would bet good money that flooding the FBI's tip line with junk about every teenager trying to learn "what be a mitochondria" does more harm to the effort of finding people who could be planning such a thing than it helps. It takes more resources to go through the mass of false negatives that have now been created as matter of policy.

                  These experiments have been run. The fictional scenario of someone learning how bioweapons work and conjuring up a plague isn't real and it hurts humanity as a whole to impede the sciences over it.

                  Because what someone can flail around in / do is learn about immunology / try to "cure cancer" with a LLM and hopefully get started on a long career in medicine. Or, a discovery that matters.

                  Because in those cases, if and when they do end up at a lab, screwing up doesn't mean death. Just tons of wasted time (and money). And they will fail / screw up. Just look at literally every undergrad in any lab and the expensive messes they create.

                  -

                  And last, but not least, yes. Teenage hackers have been a meme for decades.

                  • frotaur 11 hours ago

                    First of all, sure if you screw up designing a bioweapon you die, but unfortunately that does not necessarily mean the bioweapon dies with you, quite the contrary it can kickstart its propagation.

                    Second, in your argumentation you assume that the experiments are extremely difficult or costly. Thankfully, so far, it seems to be the case that they are too difficult (either due to domain knowledge, or to difficulty of obtaining the necessary components/equipment).

                    But there is no clear reason to think it will remain this way (e.g. crispr allows for genetic engineering which is very cheap). And it appears that, for domain knowledge, capable LLMs are rapidly reducing that barrier. We are not there yet of course, but I think it is crazy to dismiss these concerns, which are very real and crucially are NOT just brought up by the big labs.

                  • estearum 9 hours ago

                    No, the source of the misapprehension is that you are ignoring how easy it is to source the input components and combine them into a bioweapon these days.

                    Correct, there is a small number of people who have been able to do it at high cost for a long time.

                    Now, there is an ever-growing number of people who are able to do it at an ever-falling cost.

                    That's the entire issue. Do you dispute that this is what's happening?

                    > I would bet good money that flooding the FBI's tip line with junk about every teenager trying to learn "what be a mitochondria" does more harm

                    Did someone propose doing that? Or is this a strawman?

                    • skinfaxi 6 hours ago

                      > No, the source of the misapprehension is that you are ignoring how easy it is to source the input components and combine them into a bioweapon these days.

                      Citation needed. Where are these home biolabs? Why haven't any leaked yet like home meth labs?

                      • estearum 6 hours ago

                        Who said anything about a home biolab? Are you thinking a possible solution is just to block the terrorists, irresponsible corporations, or evil governments from LLMs? Obviously not.

                        In any case, the "home biolab" required to do this stuff gets smaller and more accessible every day. Biochemistry, like virtually every other complex procedural field, has become heavily outsourced. You can literally order genetic fragments even of known pathogens on the Internet, shipped to your door. There should obviously be much more aggressive restrictions on manufacturing known pathogen fragments, but 1) every money-hungry lab would need to volunteer to participate, and 2) it's totally unclear how they'd detect novel pathogen fragments that unsafe AI would be happy to help predict a couple thousand of.

                        Today, composing that into a working virus might require an undergrad biochem education, a hundred grand, and a bunch of patience (and risk), but that describes millions of people. As GP pointed out, it wasn't too long ago the group of people with this capability was fewer than a dozen individuals on the planet. And as is obvious, we are trending in one direction. We are not trending the other direction.

        • cge 19 hours ago

          There’s confusion about the classifiers on Fable. They don’t ban chemistry and biology topics they flag as a potential risk, they ban anything related to chemistry or biology at all. This is intentional, and directly stated on the model card, but seems so absurd that there can be assumption it must be a misreading.

          Being a researcher somewhat connected to chemistry and biology, Fable has been the most useless model I have ever tried. Essentially all work has instantly downgraded to Opus.

      • Mistletoe 19 hours ago

        Gemini knocks this out of the park, Gemini gang unite.

        https://share.gemini.google/34vZzlnsmTaL

        • mdgld 18 hours ago

          I would love for Gemini to be competitive but even 3.6 flash doesn’t match sol, or opus 5, or k3

          • Mistletoe 8 hours ago

            It’s all relative, it’s competitive for me that just wants a free LLM that is like a turbocharged Wikipedia.

      • actsasbuffoon 16 hours ago

        The Bill Nye theme song is a threat to national security. That’s the timeline we’re living in now.

    • thousand_nights 20 hours ago

      i do homebrewing and asked it to compare some beer yeasts for me and hit the safeguards because... biology i guess lol

    • arcanemachiner 20 hours ago

      I was profiling a slow machine the other day, and triggered the safeguards.

      I've been saying this a lot lately, but it doesn't bites you until it bites you.

      The more you use the clanker as a general purpose fix-it tool (goodbye manual NeoVim configuration, you will not be missed!), the more you will find yourself bumping into these safeguards.

    • cute_boi 20 hours ago

      Just ask math question and it will censor that. Even Misanthrophic employee confirmed that.

    • patcon 20 hours ago

      Working on dimensionl reduction algorithms, I hit it all the time. I'm also trying to port related protocols from single-cell transcriptomics to collective intelligence systems (working with people x reaction matrices as analogous to single-cells cell x gene matrices.

      Something between single-cell work and advanced nonlinear DR methods (perhaps used in alignment work?) it always flags me

    • weird-eye-issue 20 hours ago

      Literally anything related to nutrition, athletic performance, etc especially if you ask it for research or sources

      • AnotherGoodName 20 hours ago

        Writing an implementation of a board game and one of the cards is called "microbes". Instantly knocked down to a lower tier model whenever it encounters that keyword because clearly bioweapons. Sigh.

    • gck1 20 hours ago

      Reverse engineering. Codex sometimes displays an advisory prompt when classifier trips - "Wait longer while we evaluate this request further or use a dumber model". If you do nothing, it'll just take some time and almost always succeed.

      It does require some brainwashing of the model to get it to the state where model itself agrees to do RE work though. But at least it's all predictable.

      • simpsond 17 hours ago

        I’ve had codex/sol block me due to safeguards tripping without attempting to do anything nefarious. It’s not entirely predicable.

      • maCDzP 14 hours ago

        I have had success with ”brainwashing” by starting out with bug bounties/CTF and then going from there.

        • gck1 14 hours ago

          That's brilliant, I should try that.

          I usually just start by preloadig context with plausible legitimate use, have it work and obviously fail, and then ask to figure it out without ever mentioning any high risk words. Model offers to RE itself and classifiers are happy.

    • alain_gilbert 20 hours ago

      The other day, I told claude that my physical wifi door unlock push buttons is a security risk because someone could run away with it and then unlock the door from outside whenever he wants. Then I told it that I want to introduce a concept of public/private key to uniquely identify my push buttons so that I can disable them individually using some crypto like ed25519...

      Fable understood it as something along the lines of:

      "introducing" "security risk" "using software" to "unlock door" YOU ARE FLAGGED

      • JumpCrisscross 18 hours ago

        > Fable understood it as

        The dumbfuck bouncer Anthropic put in front of Fable decided this.

        Fable is a PR model. It’s great. But if it were an employee, it would be the brilliant one who regularly shows up to work high. Not useless. But not reliable.

        • jpk 17 hours ago

          > Fable is a PR model.

          Yeah, Fable is Anthropic's Cybertruck.

          • JumpCrisscross 17 hours ago

            Not quite. Fable is a Model S. The problem is you have to buy a Cybertruck to get it.

        • spicybright 16 hours ago

          Bet it has something to do with that new model being blocked by the US government. It was blocked for like a month but now that it's finally released they put the safe guards waay up in fear of that happening again.

          • theplumber 15 hours ago

            I bet it has something with Dario the drama queen begging the US gov to regulate them(I.e read ask them to put “safety guards” that they already had on hand)

    • Levitz 19 hours ago

      I routinely get into blocks when running medicine-related material through it.

    • jbritton 19 hours ago

      I was flagged for basically answering yes to what Claude suggested to do, which was test commands on a port for my code for tests we had been discussing. I really think it was flagged simply because the words test and port were in the prompt. Their filters are pathetically poor.

    • fluidcruft 19 hours ago

      I haven't had a chance to try Opus 5 yet but Fable currently refuses to do anything in my field (radiology image analysis). It didn't used to be that way but that has been the reality the last two weeks or so. Fable has been useless they might as well drop it as far as I am concerned.

      • mdgld 18 hours ago

        Are you more on the medicine side or the ML side? I don’t see many other self-admitted medical people on HN.

        • kami23 17 hours ago

          I've noticed a few self identify and other random occupations, cool to see those fields checking out this tech at a deeper surface level.

          I'm just a dev, but I appreciate the insights from other professionals.

        • fluidcruft 16 hours ago

          I'm more on the clinical physics side building tools for scanner/equipment QA and data handling/workflow automation/de-identification and anonymization but I also build random little tools to help optimize acquisition parameters.

    • Tostino 19 hours ago

      Trying to have it do some rework on a patch to Postgres I'm working on, it just completely shuts down. The reported issues were with privileged escalation and I was instructing it on how to fix.

    • dylanowen 19 hours ago

      I was trying to debug/fix a segfault in the JVM which kept getting flagged

    • theplumber 19 hours ago

      Fable refuses to work on a login/signup system for example.

    • idiotsecant 19 hours ago

      The better question is how would you not? I got demoted to opus from fable for asking if a cancer vaccine I saw on YouTube based on frog bacteria was a real thing. I've gotten it for asking how encryption works. It's incredibly touchy.

    • d-m 19 hours ago

      I took a photo of a rose bush and asked “what’s going on with this rose bush” which triggered a downgrade to Opus. It diagnosed it with rose rosette disease.

    • areoform 19 hours ago

      Things Fable's classifier has flagged, a non-exhaustive list,

          – "Does collagen supplementation empirically work?"
          - "Can you help me figure out how to calculate and generate Kaplan-Meier curve?"
          – "Why do rabbits reproduce so frequently?"
          — "Can you tell me how collagen peptides are absorbed by my digestive tract and the role they play? Can you teach me [edit: how] this works at the biomolecular level?"
      • alightsoul 18 hours ago

        Oh biology! It's dangerous!

        • spicybright 16 hours ago

          Even worse, it's offensive and kids could see it!!

          • sigmoid10 13 hours ago

            If it goes on like this, American AI will eventually be able to solve the Riemann hypothesis but deny the existence of nipples.

            • mdp2021 11 hours ago

              From the late Perscheid:

              https://martin-perscheid.de/image/cartoon/3212.gif

              ("How do you reliably put an American out-of-combat.")

              But the same vibe is felt outside the USA anyway (the uk came close to that attitude with declarations from starmer earlier this year).

      • lwansbrough 18 hours ago

        [loads up most intelligent AI ever created]

        “Rabbit sex, how?”

        • areoform 17 hours ago

          I mean, yes. Why?

          Why would nature encode such a ridiculously disproportionate / inefficient behavior when it leads to catastrophe so frequently?

          I've tried to ask these machines dumber questions like, why castles? And... well I'm working on a few projects (mostly by hand) that they've helped with! :)

          I like to ask dumb questions. It's fun. I encourage it.

          • CamperBob2 17 hours ago

            Answer: Because rabbits are tasty. If they didn't hit the ground hopping and bopping, the species would have died out long ago.

          • Alpha3031 15 hours ago

            Lots of species are r-selected, I don't see why that would be considered inefficient. In fact I think there are probably more r-strategists than K-strategists.

          • dleeftink 14 hours ago

            Just as stating the wrong thing would be the quickest way to elicit a response in the past, so too can 'dumb' questions prime the context for more complex queries.

            I am all for this strategy, and revisiting my list is equal parts fun and conducive to long-term recall.

        • kryogen1c 17 hours ago

          Rabbits are a good input calorie to output meat ratio, and their excrement makes good cold compost. They breed and litter relatively easily. Theyre also easy to house.

        • fallingbananna 11 hours ago

          The question may sound ridiculous.

          But isn't it more ridiculous that some company decided that the correct answer to such innocent and curious question is basically "That knowledge is too dangerous, you shouldn't ask that."?

    • rzk 19 hours ago

      I had a long session about SQL with Fable and at some point, it started to falsely trigger censorship for any message I write in that conversation, even for the simple string: "random message."

    • JumpCrisscross 18 hours ago

      I wanted to explore some battery chemistry with Fable. It decided I was a terrorist, and then blew my session’s usage credits telling me to fuck off.

      The current state of guardrails seems to be entirely about marketing to investors at the cost of customers. I’m switching to open models when my subscription expires. Almost everyone I know, including those with access to Mythos, plan the same at the earliest opportunity. (Or until one of the SOTA models leaks.)

    • pmarreck 18 hours ago

      What are you asking that you are NOT regularly running into censorship?

      Pretty much everyone I know who uses Claude and works on anything with any level of detail has gotten false-positive flagged

      I got flagged for coding in WebAssembly Text, for chrissakes LOL #haX0r

      And honestly, Codex handles this better. It says "Things are going to go a little slower because we must perform additional checks on this. Is that OK?" and your only inconvenience is waiting a little longer.

      Fable meanwhile just unceremoniously dumps you right into Opus without asking anything, it just tells you "you're in Opus now, sorryyyy!" Lame.

    • OkWing99 18 hours ago

      Ask anything related to practical applications for quantum-computing, or space etc. The stuff you can find on Wikipedia.

      They were able to solve coding, but not what a real danger is.

    • rdtsc 17 hours ago

      No op but my interactions with look like:

      Me: "I got this crash in production, looks like a segfault, let's try to fix it. Here are some functions that might be responsible."

      Fable: "No. This is cybersecurity, blah blah, I won't help you"

      I forgot how I got it to fix the bug eventually. I think I convinced it that it wrote the code and made a mistake. But it was definitely a "Hmm, may be I should use another model" moment".

      • actsasbuffoon 16 hours ago

        I have way too many AI subscriptions. My favorite thing about Kimi K3 is that it just does what you tell it to do.

        “Hey Kimi, penetration test my app,” doesn’t get me a refusal, a guardrail, or anything like that. It gets me a pen-test result.

    • actsasbuffoon 16 hours ago

      I literally had Opus 5 flag a message because my hand was slightly too far to the side while typing. It apparently decided that a sentence that had some garbled words in it was a threat to national security.

    • __MatrixMan__ 15 hours ago

      I'm not who you're responding to, but I have a lot of questions about molecular mimicry: evolution pushes pathogens to be shaped like human cell surfaces because that way the immune system won't attack the pathogens (since, by doing so it would also attack the body). It's thought that many autoimmune disorders have an undiscovered pathogen as their cause, one whose mimicry caused such an attack. Discovery of these pathogens could be done computationally, I think. We can catch MHC binding event in process, find the bound protein, figure out which pathogens have genomes that code for proteins of similar shapes (epitopes), and we'd find--I hypothesize--a list of candidate pathogens for the cause of a delayed onset autoimmune disorder. Preventing these infections ahead of time would be a huge win against diseases like multiple sclerosis because without the initial exposure the immune system wouldn't have cloned so many of the cells that are attacking the host.

      Claude was utterly useless in my attempts to write a paper about this. Wouldn't even help me search for sources. I guess you'd be asking the same questions if you wanted to develop a pathogen that could reliably evade the immune system.

    • epistasis 15 hours ago

      I use lots of biology in my day job.

      Asking Fable 5 "Why did the chicken cross the road" results in switching back to Opus 4.8. I'm not joking, it really censors that, and I'm not alone in the result.

      The memory aspect means that your prior work has a huge impact on what gets censored.

    • smoe 8 hours ago

      Not OP, but I got downgraded almost every time I had Fable implement something with security implications. Codex reviews it and identifies potential security issues, but Fable refuses to address them and downgrades instead.

  • buzzerbetrayed 21 hours ago

    Yep. I cancelled my Claude Max subscription 2 weeks ago after feeling like Anthropic was doing everything it could to fuck with my day to day. Their lead would have to become significant for me to ever go back.

  • gck1 20 hours ago

    Every time an online chatter (e.g. "limits are better", "model is better") makes me to reevaluate my principle of never paying Anthropic, I go to the model card, which strengthens my belief in the principle.

    Why is Anthropic is so hell-bent on this auto/silent downgrade? Do they have a single user who prefers an auto-lobotomization instead of a refusal? Have they learned nothing from the backlash the first time?

    • markasoftware 20 hours ago

      I don't believe they do silent downgrades right now, they're loud about it.

    • kccqzy 20 hours ago

      Just go to /config. The very second configuration item is “Switch models when a message is flagged” and presumably you want to turn this off.

      Oh but then you said you never pay Anthropic so you haven’t actually used Claude Code yet. Why would anyone listen to the opinion of a non-user?

      • gck1 19 hours ago

        > so you haven’t actually used Claude Code yet.

        Where do you think the principle came from? I've used claude code for a year, and stopped February this year.

      • nananana9 13 hours ago

        Not a great line of thought in general, sometimes the people who aren't doing the thing are the only ones worth listening to.

        "You aren't repeatedly slamming your head against the wall. Why would anyone listen to the opinion of a non-wall-head-slammer about the merits of wall-head-slamming?"

      • krzyk 9 hours ago

        So when you turn that off will you get answer from Fable?

  • joinjune 17 hours ago

    I told Claude Opus 5.0 to use a global api key for a PFAAS to deploy some web applications in a test environment and import some data into them. It balked at using a global api key because the security issues surrounding the permissiveness run afoul of it's sensibilities.

    I have done this task with Opus 4.5, Opus 4.6, Opus 4.7, Opus 4.8 and Fable, without issues.

    I have done this task with Codex 5.4, Codex 5.5, and Sol 5.6 without issues.

    Opus 5 is too cautious to be productive for me. It needs more tuning.

  • NamlchakKhandro 16 hours ago

    The utter meme-think direction this company takes with regards to sycophancy of its models is disgusting.

  • villish 16 hours ago

    It's definitely not benchmaxxing from my experience with it. I have a test I use on all the models to create a game and Opus 5 feels like a generational leap compared to the rest. Benchmarks don't paint an accurate picture, you have to try them for yourself.

    • chamomeal 16 hours ago

      Even compared to fable?

      • villish 16 hours ago

        Yes. I've since watched a couple review videos on YouTube and all the game tests I've seen Opus 5 produce are incredible even compared to Fable.

  • mrloopex 14 hours ago

    Maybe you just aren’t doing anything meaningful. Terrence Tao doesn’t whine about woke AI models.

entity002 11 hours ago

I like how Opus 5 doesn't re explain EVERYTHING to me like 4.8 did. GPT 5.6 SOL reasons WAY too hard over nothing, and Opus 5 is an amazing mode. Way to go anthropic

  • ModernMech 7 hours ago

    Yes, I didn't appreciate this because I was giving Sol a brief to implement and it was doing very well.

    So then I just told it to do its thing without a brief and it went for 2.5 hours and used 30% of my week. I tried the same task with the brief and Sol went for 30 minutes and used 2% of my week. Compared the two and the 30 minute brief-based Sol output was much better factored, shorter, validated better, scoped better, and of course cheaper.

    Left to its own devices, Sol goes out of control.

    Now I ask Sol to write the brief and Terra to implement it, works pretty well and overall usage is down.

chmod775 21 hours ago

The more interesting finding is that it's still the second most expensive model (after Fable 5) by a long shot.

At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.

  • ricardobeat 20 hours ago

    The chart shows max effort, used mostly by price-insensitive enterprise users. At medium effort it drops to almost half K3’s cost, and is probably sufficient for 95% of coding tasks.

    • stingraycharles 19 hours ago

      Yeah if there’s one thing that people should really understand it’s that it’s cheaper to have smarter models with less thinking than cheaper models with more thinking.

    • andriy_koval 18 hours ago

      Opus medium = Sol high = 56, but still 25% more expensive

    • Bolwin 18 hours ago

      For a fair comparison, you should compare to K3 (which AA has not tested yet unfortunately) and GPT 5.6 Sol also on medium or the closest equivalent

      • mdgld 18 hours ago

        I thought k3 medium wasn’t even available yet?

      • aoeusnth1 17 hours ago

        No, because Opus is smarter than those models if they’re all on medium settings. You should compare at similar levels of performance, which would be favorable to Opus.

    • krzyk 9 hours ago

      Enterprise users are price sensitive, because they are charged per token (and have limit per user set by companies, and those limits are different from $50 per month to 1500$ per month). Subscription users might be insensitive to that.

    • KronisLV 4 hours ago

      > At medium effort it drops to almost half K3’s cost, and is probably sufficient for 95% of coding tasks.

      I've also found that Kimi K3 on Max reasoning is benchmaxxing a little bit, High is probably enough for most dev work as long as you have good tooling and a good, detailed plan (which you can create on Max reasoning if you want).

drob518 6 hours ago

New respect for Meta Muse Spark. It seems to sit at a lot of sweet spots in the leader board. It’s not the best at anything in particular, but it balances cost and performance quite well. I’m also curious where Poolside Laguna S would sit; it’s not included. I’m personally very interested in cost effective models that still perform well.

nu11ptr 20 hours ago

I don't have a horse in this race, but to me this makes GPT-5.6 Sol Max look better. It is about half the cost for nearly the exact same performance. It just goes to show how expensive Fable really is when Opus 5 is still this expensive relative to GPT 5.6.

  • mdgld 18 hours ago

    Make sure you’re comparing opus high to sol max. That’s where the comparison makes sense

codewiththiha 10 hours ago

I can't wait for open-source models to compete against this!

zormino 21 hours ago

I'd be curious to see the results, especially with some models having 1.5m and 2m context sizes, if the first 75% of the context was filled with unrelated info.

kristopolous 19 hours ago

I posted this before but I have a really simple shell tool to keep up with these charts over at

https://github.com/day50-dev/aa-eval-email

This also works

$ curl day50.dev/art-analysis.sh | bash

Artificial analysis knows about my tool and I'm working with them on getting their API improved.

hoppp 21 hours ago

I didn't like it as much as fable. The coding style was a bit different and it way overbuilt the thing I asked from it.

  • vehemenz 21 hours ago

    It’s crazy that people feel confident making judgments like these when the model’s been out for only a few hours.

    • CommanderData 18 hours ago

      It's a pretty easy spot if you're already using an older model daily.

    • jrs100000 17 hours ago

      Its even crazier that people are sitting here trying to calculate intelligence per dollar from metrics. At least first impressions have more basis in real performance.

  • cbg0 14 hours ago

    It overthinks quite a bit above medium effort, try using that.

theplumber 20 hours ago

Why do I find it dummer/even more superficial than opus 4.8? It just continued a session and I had to stop it because it become obviously “lost”

  • king_phil 9 hours ago

    Massive performance degradation is expected when continuing a session with a different model

    • protimewaster 6 hours ago

      Is it? Is there a paper that covers this?

      I would've thought it should be mostly seamless, since it's being fed the entire conversation all along anyway.

zkmon 18 hours ago

I think a more useful metric would be intelligence per dollar spent.

fnord77 5 hours ago

For nearly twice the price, you get 1 tick higher on some intelligence index than 5.6 Sol

XCSme 18 hours ago

Twice the cost for 4% more intelligence, is it worth it?

luxuryballs 7 hours ago

Anthropic has imo underrated marketing and positioning skills, mythos/fable hype/fear being the most obvious indicator but even the way they almost haphazardly position their models with no intentional cohesion, people see model names and numbers, it's easy to think of them as more intentionally accurate like how cars make S models or AMG, but then the performance and surprises surpass the prior expectation that was set by previous models, rather than having it be more obvious, suddenly the Anthropic Camry will outperform their Corvette without any fanfare.

nekusar 7 hours ago

"When a measure becomes a target, it ceases to be a good measure"

Goodhart's law.

NamlchakKhandro 16 hours ago

This company sniffs it's own farts too much tbh

zuzululu 18 hours ago

i used for several hours now and my verdict is that its no better or worse than sol

its surprisingly bad at UI which is unexpected

its also lacking in depth vs sol 5.6 which goes above and beyond (which in itself is also an issue at times)

sggyamg 21 hours ago

It's new, normal.

  • LeBit 21 hours ago

    Wait until DeepSeek v5 Pro is released in less than a month and costs 1/100 to perform the same tasks.

    "Not fair! They distilled Opus 5!"

    • brikym 18 hours ago

      Why doesn't Anthropic distil themselves then if that's what makes it cheaper. IMO the cost savings can't all be down to distillation.

fHr 5 hours ago

metrics gooners are pretty regarded

anigbrowl 17 hours ago

Honestly, who the fuck cares? These leaderboards are meaningless for brand new models. If we were looking at longitudinal data collected over the course of a year or even a quarter or month, this would have some value. Brand new model from established provider shoots to top of charts? This means nothing more than an already famous band briefly topping the charts with their latest song.

It baffles me that intelligent people deploying AI think momentary popularity is a meaningful signal. It's just twitch reactions * FOMO.

  • jsnell 13 hours ago

    Uhh... This is not a popularity chart. It is an aggregate of benchmarks.

    It baffles me that somebody would write something that aggressive from that deep a level of confusion.

claude-ai 23 hours ago

On my end, Opus 5 is Haiku level vs. Opus 4.8 (good) and Fable (superb).

Gets confused by permission prompts, cannot debug a failing test it caused (Opus 4.8 got it right after, without tens of rounds "thinking").

  • reilly3000 21 hours ago

    Are you using Claude Code/CoWork or an API client? I’m curious if it has different training that makes it more effective with specific instructions/ tool calling methods that are only implemented in official harnesses.

    • pixelesque 21 hours ago

      I'm curious about this too, and it's difficult to get any information about this given everyone has different setups, workflows and use-cases.

      I bizarrely had Opus 4.8 this week (in pi.dev within a podman container, using openrouter) start installing various python packages (and uv!) within the environment (not as root) when I asked it to code review some fairly basic Rust .rs files that were generally stand-alone (it did very nicely work out and write some stubs for them to build them and work out how they worked).

      It only gave up with the weird Python installing stuff when it discovered one of the Python packages needed Tensorflow.

      It seems pretty focused and persistent in continuing its initial approach, and I'm wondering if I need to alter some instructions / initial prompts to rein it in a bit...

      • kadoban 18 hours ago

        There's probably some system prompt crap in claude code that tamps down that behavior? I wonder what it was even trying to do.

        • mdgld 18 hours ago

          Maybe that’s why it’s like 15k tokens lmao