I head a much, much smaller open source project. Since the November Singularity we've been seeing at least six responsibly reported security advisories a month. However, this last month we had 22 unique security advisories. Our project has been built with adherence to the OWASP Top Ten Guidelines and other best practices from the beginning. But software is hard and AI is thorough.
Each month, we fix them all in our monthly maintenance release and disclose at that time. We fight AI fire with fire, and hand-review, of course.
So far, we can keep up. One hopes this is possible at the scale of the Linux project, which assuredly has more humans and more AI to throw at the problem. But team size does not scale linearly with interested audience, and potential bugs do scale with codebase size (and other extremely important factors, like code quality, at which the Linux team is assuredly much better than we are).
("November Singularity" is a cheeky reference to the arrival of Opus 4.5 and "good enough" coding models and harnesses generally.)
It was. That's when I stopped coding most things by hand. The difference in what I could reliably get AI to do for me in February 2025 vs February 2026 is just massive. I immediately became a huge advocate amongst my peers for going all-in on automated engineering because this curve is about to get very steep and if you're not staying ahead, you might get left behind.
Well it's a good thing that isn't what I said or implied, and that you instead smugly misinterpreted my comment.
The person you're replying to had the decency to ask me for clarification; for you, I'd suggest brushing up on your reading comprehension and learning to respect the rules of this site, which include interpreting comments in their best light and focusing on positive, substantial contributions instead of negative, inflammatory posts.
Again, I'm going to refer to the advice I just laid out for you. I'm not responsible for you and I don't need to cater my words to your lack of reading comprehension.
I think this is both true and untrue. For a long time, I was an AI skeptic, perhaps even a hater. A chauvinist for writing code by 'hand.' But it's become very difficult, as I've gradually migrated to AI-authoring of code, to go back to writing code by hand and retain the same velocity.
Part of this is certainly that my hand-authoring code skills have atrophied, sure, but my workflow has also radically changed. Previously, I would spent a lot of time and focus on a single work item, and only context switch to other tasks whenever I would wait on CI or a long build. It meant that I spent a lot of time understanding one thing at a time, and interruptions (forced context switches) incurred a massive switching cost.
Now, having moved to largely AI-authored code, I find myself necessarily working on multiple threads at the same time. This means I can meaningfully progress each of those threads in parallel, with a much-reduced overhead on context switching, since I don't have my head down focusing on all the details of the work. And if the task really demands it, I can still stop and focus on one thread to sketch out the code manually, think about the concepts more deeply, etc.
It's a very different workflow, and there are certainly downsides, but the upside is that the rate of which I've been able to put up good-quality PRs has measurably increased. It's not quite 2x, and it's certainly not 10x, but it's definitely noticeable. I do understand a bit less, but there was always more work than time to understand things in full detail. I guess only time will tell if that missing understanding was actually vital to the long-term success of my work.
yes that's fine, i have a similar experience. but there is a difference between saying what you did, and the type of AI religious fundamentalism where you got "converted" when truely good AI coding models revealed themselves to you, and then you go around trying to save those who would otherwise be "left behind".
AI is becoming a more and more powerful lever that lets you move greater loads per application; it is in a real sense giving you more leverage. But the lever is difficult to use effectively, and the methods for using the level change every week or so.
> methods for using the level change every week or so
hmm doesn't mean noob will become better? if what I'm "learning" about AI is going to get deprecated every week... then it doesn't make sense to "keep up"
Here's how I think about it. There's two different things "knowing how to use AI effectively" and "knowing how to leverage effective AI to improve your project".
1. Both are multipliers on skills you already have. A "noob" will be at a disadvantage.
2. The second one is a complex blend of technical skill, domain knowledge and human factors (knowing your user-base, knowing your product/project, understanding UX/DX/whatever) that you probably will always have an edge on.
It requires skill to prompt the AI to do best possible work. I have senior colleagues that produce horrible slop and others that always produce great results. Managing the context, knowing when to stop the AI. Knowing what to question, what to "trust" in is a big deal.
Also knowing what models are good for what. There was time half a year ago when Google genuinely had a better model than everyone else. I used it for everything and 3 weeks later they lobotomise it (sorry "optimised") and I went back to opus...
I think Anthropic is like a drug dealer, giving us the sweet sweet drug for free (I don't think our $200 a month subscriptions even cover the electricity for our use) and the time to pay will come very soon...
I expect this subscription will cost $2k a month. Will it be a normal increase? Or will the "enshittify" existing models to the point you'll pay $2k to get fable 6 to do what opus 4.8 did fine in July 2026?
I was referring to the amount of people in the future who might be employed at today's engineering wages, and the differential between those people and the average employed engineer today.
What will that differential be? A large component will be social reinforcement: generational wealth, connectedness, and such. Things like merit might take more of a backseat. So people who do not have the necessary social capital may be fighting each other for very limited amounts of positions.
That is the steepening of the curve for people like me: I was homeless at 16 and finished high school on my own, was given a full ride to LSU, grants, room and board and a job in the Comp Sci department, but lost all of it after an immature and vindictive high school teacher illegally modified my grade in a core class in order to fuck me over. I didn't have parents to back me up at the school board and make things right.
Instead I suffered through years of homelessness and had to find my own path into the industry by starting companies with friends and doing all of the engineering. Since then I've led multiple teams, made some connections, shipped a lot of cool stuff and bring to the table a wealth of experience and a generalist skillset that is both wide and deep. Yet, I too wonder where my place in this changing industry will be once things settle a bit. Probably less engineering and more focus on business development.
The flipside though is as you've said: The fruits of engineering are more accessible than ever to the layman, and individuals can currently possess an unprecedented amount of agency and leverage. I think that is amazing and am fully behind it. I do know that it means the process of renormalization is going to be very rough, given the similarly unprecedented rate of industry change these technologies are bringing.
There are two types of AI users: those who chase the rabbit and those who do not. Some people bounce from tool to tool QuantumLeap-style hoping that maybe the next tool will be the big thing to solve all their issues. Other people use AI to do actual work. Once they find a tool that works, they stick with that tool until they feal a need to upgrade.
It is like people fishing. Some people go out and catch fish with the tools they know will work. For other people, every day at the lake requires a new boat/rod/lure. They spend more time figuring out how to use their new toy than they do catching fish.
Because improved models seem to increase, not decrease, the gap between what you get from really expert supervision (prompting, steering, hand-corrections along the way, etc.) vs what you get from fairy-dust and wishes.
I like the November Singularity, I hope that catches on. That was definitely the point where I went from "AI is overhyped" to "oh shit the hype bros may be on to something"
Yes, good name and I also feel that was a major inflection point. Up until then I found all AI models to be terrible at programming, with the difference between GPT o3, Sonnet 4 and every other model since GPT-3 being just the exact flavor of terrible they were.
November 2025, with Opus 4.5, was the first time I was impressed by an LLM doing something non-trivial with a reasonably good level of quality.
Yes, and Qwen 3.8 27b is a similar inflection point for local, although I wouldn't claim it's quite as good as Opus 4.5 it is genuinely useful. Whether that actually matters will depend on whether it ever becomes the most cost-effective tool for the job, but it's impressive as all heck
>> We fight AI fire with fire, and hand-review, of course.
Wouldn't it be nice if AI vulnerability reports came with AI pull requests to fix them? The thinking context that found it should be readily able to propose a fix. It would still need review but even when AI PRs aren't right they often point in the right direction.
Why? Imagine that you are a competing, close source product/project. Just bombard your competition with AI reports, and let them drown in misery. Problem solved! (/s)
This just goes to show that yes, if security researchers were to do that, it would great, but they are not the only actors here...
I'll speak up in defense of the reporters: FWIW, so far my strong impression is we're hearing from independent security researchers. Right now, for us, so far (enough qualifiers yet?) the system is working for us: independent security researchers are farming reputation by finding real problems. That's not a bad thing.
And in most cases they do propose solutions, although we generally resolve the issues on our own.
There's a small percentage where we make the case that the ticket is not a real vulnerability, and then we have to grit our teeth through repeated reports of the same "vulnerability." But it's a small percentage so far.
We do typically have to reconsider the severity. The researchers understandably want to see everything as a nine...
wonder how much this can easily speed up the malicious-contributer attack. where someone suggest a security fix in a extremely obscure and irrelevant code, but the fix actually adds a new condition that then can be exploited elsewhere.
I wonder if August of this year will be remembered like that too. The very first "opus like" local AI model came out this August (Qwen 3.8 Flash Next). I've been running it locally since for real programming and I consider it pretty much the same as opus 4.6 in coding ability (it lacks a bit in the factual knowledge area). It even exceeds opus on some tasks.
I've tried every hyped "open source" model before and all including latest models bigger than 1T parameters are pretty much toys.
This is the first one that isn't. It can't be overstated how huge of a deal that is. No more reliance on Anthropic.
Running this model to do real work is still not cheap. I run it on a pc with 6 rtx3090s and 192GB of RAM (and I use 90gb of that ram for KV cache). The model is entirely on gpus. It runs at 55tok/s 1500 prefill for one user at a time, and about 35tok/s 650 prefill per user for 5 simultaneous users. It doesn't seem like much until one realises you manage your own kv cache. You ca leave your sessions in cache for as long as you want. You can save them and restore 200k sessions a week later in a dozen seconds.
What many people don't realise is that usage of those models skews extremely heavily towards input processing. My claude code usage is about 1.3B tokens input per week and only about 8M output. On claude code I get 80% cache. At home it's more like 95%.
If this progress keeps up, and we get a fable quality model in a year in under 200B to run at home... Those "frontier labs" will be renting all of their gpus per hour not to go bankrupt.
I am using qwen3.8-27B-UD-Q3_K_XL on a 5080 16gb card with 16gb of system ram at around 43-45tok/s with 65k context. I am using that model to set up containers on a proxmox server with two b70s and 96gb ram. The Q3 model implemented 10 different chat models, 2 image models, and 7 web based harnesses so I can compare. It can rewrite any part to improve it.
qwen3.8 is more than capable for software dev. The gemma models were terrible and fell apart during compaction. I think as long as the llm can test the results, you don't need anything being offered by a "frontier" model. The cloud AI is going to be used by people unable to run their own and they will eventually get squeezed on price.
The one tip I could give is compact before starting new steps or any action in the plan that is different than what was previously worked on. You want to manage what is in the context and don't want unnecessary work history details filling it up. You can always ask the llm to list the current plan, then compact after and do it more than once until you get the compaction <30%. This will be fixable by the harness that can choose better times to compact.
You don't want to start a new phase and have it compact a few minute after starting. This happening over and over again seems to potentially cause issues for long running sessions. Compaction slop that screws up what is in the context.
I can compare what I use at home vs paid models like astra at work and the difference is mostly meaningless.
For me the November Singularity was when ChatGPT was released in November 2022. Of course, it was a faaaaaar cry from what we see today, and it took a lot of careful wrangling, but it could write reams of correct code and tests even back then.
The thing was the wrangling was relatively straightforward, if cumbersome. Largely, it involved being very precise with the context and instructions it was given. I could imagine a lot of that getting automated (i.e. what we today call harnesses) or recursively addressed by creative meta-prompting. Supported by similarly conceptually simple advances like chain-of-thought reasoning I suspect that is the biggest thing that the models have figured out what to do today compared to then: manage themselves carefully.
Although I could not have predicted these exact outcomes, the implications for everything that is unfolding now were clear even then.
Seems to me we need tooling to do automatic offensive security as soon as a new frontier model comes out, with quickly turned around patches using the same frontier model. Rinse and repeat. Virtuous agentic security loop.
As soon as a new model drops I point it at one codebase I have and ask it to do a security audit, look for bugs, check for optimizations, things that are over-engineered, etc.
Every single time I've done this it has found at least one serious bug or security hole.
Isn't that part of the cloud AI business model now? Businesses have to pay for early access so they can weed out any new bugs before the model goes public. Which is kind of crazy, because you are paying for early access to protect yourself from other paying customers of the same cloud AI model.
note that _any_ bugfix is assigned a cve, which makes for big numbers.
>“Due to the layer at which the Linux kernel is in a system, almost any bug might be exploitable to compromise the security of the kernel… Because of this, the CVE assignment team is overly cautious and assign CVE numbers to any bugfix that they identify.”
There are programs like sudo whose entire reason for existing is to enable privilege escalation. If you can find a way to make a user "sudo" something, that's an exploit, but it's not a bug in the program.
Alternatively, consider any program which loads dynamic shared libraries (called "plugins" in many contexts). The program itself might be bug free; any plugin that is loaded will run (typically) with the full priviledges and access of the program (and thus likely the user).
The user may have no idea that the plugin is malicious; the program remains bug-free (if it was beforehand).
At least with plugins most people have a general notion that they run some sort of code as they provide some new functionality.
Much more agregiously you can design harmless a looking format that can run arbitrary code e.g. .doc with VBS. I have a very hard time blaming an user who falls for that even though MS puts up a scary looking popup.
Tautologically every bug can legitimately be assigned a CVE, since every bug prevents some feature from working as intended. It's therefore a denial of service, which by the definition of the CVE system using CVSS means every bug is at least a 1/Low level vulnerability to CVSS v4.0.
If you're willing to stretch, missing but planned features also deny the use of said features since they haven't been added yet, and so are CVSS 1/Low vulnerabilities.
Resume-driven development for security researchers has never been easier!
That doesn't follow. In the extremely simple example, an adding service returning 1+1=3 has a bug, but it's not a possible DoS situation at all.
> missing but planned features also deny the use of said features
That's not what DoS is.
This whole situation with CVE assigning comes from the whole process being far from ideal. But it doesn't mean it's completely useless and doesn't follow any rules at all.
Until someone finds there is a user input they can trigger this bug causing some other bit of code to read data from the wrong offset and now it's a whole exploit.
That's an issue in the other code, not in the addition service. It would be lumped together if it was an addition function close to the other code. But I wrote service there on purpose.
It could. But it's a CVE in the system that crashed or in the filesystem, not in the calculator web service that we were discussing. If a filesystem decides to use a bad online calculator for its internal logic, that's a vulnerability on the filesystem, not the calculator.
We aren't talking about a calculator web service though we are talking about the Linux kernel and if there's a bug in the kernel that could conceivably cause a crash in an otherwise correctly written application then that would be a CVE in the kernel.
That’s not what Mitre thinks though, they are very happy to host a 9.8 severity CVE for 1+1=3. They’ll probably publish one for 1-1=0 too, if you preface with ”The users of mathematics might not be prepared for zero values”
You're memeing on bad cve handling, but that's so far outside of what the real issues are, it just doesn't make sense. How about telling people about what the real problems with vulnerability classification are, rather than mitre=bad?
Not really. cURL developers just have NIH syndrome.
Organizations that track their software and patch systems for CVEs have their own risk management system.
Publish xlow as low, we'll filter them out if we want to. As even they state in that article, they do NOT have the necessary context to filter stuff out. So why are they doing it?
From experience: many companies have a “risk management system” that involves nagging OSS projects to do free work for them, even when the advisory is manifestly nonsense or has no impact in context. Many teams have a “green dot” mindset, and CVE directly encourages that behavior by stapling CVSS scores to identifiers.
What does this even mean? cURL is one of the most load-bearing pieces of software in existence. It, and the Linux kernel, which takes a similarly dim view of the CVE system, are the inventors. "Not Invented Here" seems to imply that there is a vast body of peer work for them to draw on to resolve this problem, but who are their peers? As far as I can tell, the answer is something like "Microsoft and Apple", on the one hand, who exist in a totally different, mostly closed-source or at least closed-development, ecosystem, or something like "glibc and OpenSSL" on the other hand, which have their own storied CVE history.
Tell us about a system for managing vulnerability database that works across enterprise, private and open source, is staffed enough to research both impact and disputes, is funded enough to work for decades, isn't partisan to any industry interests, can classify vulnerabilities in a non ambiguous way, can handle public submissions at any volume in a timely manner...
and one that will never make decisions that is incorrect or disliked by any party.
>hash of supplied password does not match hash of stored password, access denied
Resulting from miscalc of either stored or supplied would indeed create DoS.
If the off by one is in a graphical driver that causes an overflow leading to no output, that's DoS.
If the off by one is in the memory mapping of input devices leading to no available input, that's DoS.
Simple math is kind of everywhere in the kernel and userland apps. If the math is wrong and results in memory mapping wrong such that kernel panics or is unuseable, that's a breaking bug regardless of simplicity.
Loads of bugs aren't CVE-worthy. If you tell the computer to make a light green but it makes the light red, that's a bug but no DoS or other CVE-worthy bug.
However, the Linux kernel is supposed to run any userland program without crashing, so anything that crashes the kernel is a local DoS and there are a lot of them. It's also supposed to shield processes from each other and maintain privilege levels correctly, so many incorrect memory leaks are also CVE worthy. Whether a CVE applies depends on the people and programs using the kernel, and the kernel team can't read your code to tell you if it applies or not.
People reading CVEs wrong ("it's got a high number so we must patch within a day") must be going crazy over this, but the point of CVEs is to let you make judgement calls, not to be a cool statistic about how secure something is.
Most CVEs are irrelevant to most people, that's always been the case.
Yes, it is a bit of a stretch, I desperately hope programmable navigation lights are not a thing. And I also don't think every bug needs a CVE. But in the correct context nearly any bug could be critical.
Computer Weekly in the UK did some pioneering journalism into a helicopter crash[1] that had initially been blamed on the two pilots but seems very likely to have been caused by a failure of a computer system that was controlling the fuelling of the engines. Iirc this system seems to have crashed causing engine failure of both engines in heavy fog, bringing the helicopter down with a loss of everyone on board, however for reasons somewhat unclear, the MOD wanted to cover the failure up by blaming the pilots.
Now this was a windows system[2] rather than linux but the point remains - if there was an external vulnerability in a crucial control system and this system was part of a network (eg to connect telemetry) then any exploit of that system could result in loss of life.
After an assessment of the Fadec software the Superintendent of Engineering Systems said that the density of deficiencies was so high that the software was unintelligible.
[2] Which, why? Why build the fuel controller for a helicopter engine on windows?
A windows system controlling the refuelling ? Worked in aviation software for a while and the certification for flight control (or similar) systems is subject to rigorous path testing/inspections/approvals etc DO-178C (level A or B likely). Not sure if any windows OS is certified to level A?? Typically certifiable RTOS'es are procured for those purposes
If only the user can tell whether something is potentially dangerous, then either every bug or none of them should get a CVE. There are countless systems out there that are commonly used in ways beyond what even the developer intended, how should a third party authority like the CNA be able to discern this?
They can't. Which is why security discussions are such a hot mess.
The vendors fixing them arguably prioritize these reports right. Most of the CVEs, even severe ones, are irrelevant in practice, and as parents note, are more like regular bugs with security flavor in reporting. The CVE label instead of regular bug tracking number makes them seem important.
Linux Kernel may be one of the few legitimate exceptions, indeed, due to the position in which it sits in the software stack. Also LLMs make previously unexploitable-in-practice vulnerabilities exploitable (by making targeted / personalized attack cheap enough to give them positive ROI), which complicates things.
I think you have to make a distinction between bugs and actual vulnerabilities. Therac-25 killed people, but I wouldn’t consider anything about it to be a security vulnerability. In my mind, the distinction between a bug and a vulnerability is that a bug is triggered during “normal” operations and can do anything. Whereas a vulnerability requires an adversary to “trigger” the vulnerability, and can do this to achieve some cognizable malicious goal. There’s probably some overlap on the edges; whether an issue in a library is a bug or a vulnerability may depend on how it is used, for example.
But I think the most important thing to keep in mind is that a bug isn’t necessarily less serious or less important than a vulnerability. A serious bug should be patched just as urgently as a serious vulnerability.
And not only that, remember the hn crowd is not at all representative of the average.
Even more realistically, many admins do not have the background to be able to reason (by themselves) about the actual risk of most CVEs, so just going along with specialized media coverage is often a sound strategy.
I do find it interesting though, that in the interest of transparency, every bugfix gets a CVE. Which ends up being a huge number… which will ultimately yield a more insecure environment as we’re getting conditioned to ignore/discount CVEs by the volume.
Over-reporting in this case seems to risk being counterproductive.
Depends on the end consumers stance on security. I've watched it shift from "Only update if we can prove we are impacted" to "Update everything immediately just in case".
The frequency and severity of cyber attacks has increased to the point a much more cautious approach has become common. It's also easier to sell this work to management when you can point at the security tab on some tool and say "Look we need to patch these CVEs"
"Update everything immediately just in case" is a lot less attractive when you see more downtime from updates breaking things than you do from hackers. Windows updates are an endless source of pain, but now every program seems to demand to be updated practically daily. Even things that you shouldn't have to think about like keyboards, mice, and printers beg to be updated all the time.
There is no reason, why a newer version has less bugs, than an older version. Both are essentially an unknown number. The only thing you know, is that you likely know a higher percentage of bugs for the older than for the newer version.
It certainly has less known bugs. And when you have to make a statement to the media, “we were hacked by an undiscovered 0 day exploit” sounds a lot better than “we were hacked by a known exploit because we didn’t update”
This isn't necessarily true, and will be less true in the coming age of LLM-automated vulnerability scanning. The version that you're downloading (after being nagged for a day) that adds Feature A may contain 3 vulnerabilities that are already known before you even download the update, and may or may not fix old vulnerabilities.
Yeah, the fact that security forces you to update has been used with great effect by product managers and feature engineers to shove their changes down everyone's throat and quickly drop support for previous versions. Neat for them.
It's not very interesting. Linus, and by extension the Linux kernel, long had a dismissive attitude toward security research. This is basically a childish swing from one extreme (nothing gets a CVE) to another (everything gets a CVE).
Kernel development is well-funded, both via grants and by direct employment at big tech companies, and if they wanted to properly triage and annotate vulnerabilities, and provide reasonable assessments of what is or isn't likely to be a security risk, they absolutely could. They almost certainly could go to Google and say "we need two people full-time on your payroll for this" and they would get it.
I don't want to dunk on them too much because they're generally doing God's work, but these absolutist security stances are not worth being taken seriously.
It's basically saying that they can't possibly provide a valuable service for 99.999% of the install base because there might a hypothetical person out there using Linux in a really weird way. If Microsoft tried to make an argument like that, they'd get crucified.
The counterpoint is that by putting CVEs on bugs that are more easily exploitable provides a roadmap for attackers. Of course, in the current LLM age that's probably a moot point, but that could be the reason for this.
> Linux only ever wanted to promise support for the latest release and even Linux LTS is a concession.
If he kept true of his "we don't break user space" instead of it being "we don't break user space until we do and then it's on you to deal with it" perhaps more people would be willing to run the latest release.
I have had way more issues provoked on RHEL-compatibles by RHEL's Frankenstein backporty kernel than I ever have on Arch or Nix by the latest stable kernel.
Linux LTS is much more about proprietary drivers targeting a stable internal kernel API/ABI than about anything else.
> if they wanted to properly triage and annotate vulnerabilities, and provide reasonable assessments of what is or isn't likely to be a security risk, they absolutely could.
I'll challenge this, I don't think that's really possible, at least not to any high degree of confidence. Nobody can realistically evaluate if any particular out-of-bounds read/write or use-after-free is "safe", and if you're going to consider all of those as security risks then there's not really a point in trying to filter out the few bug fixes that might not lead to those things.
Try going to the linked page, pick any random CVE, and read it. I've checked a bunch and I'd say _at least_ 8/10 of them are variations on those two things.
Counterproductive for whom? The stance of the kernel developers is that whether a bug is a vulnerability or not depends on intended use. They say, we do not dictate use, therefore this decision is out of scope for us. This makes kernel development more focussed and productive.
I would argue that it is more productive for the enduser as well. Not making a decision they cannot reasonably make is better then blindly believing in a decision that is likely wrong for your usecase.
Any patch that is back ported to a stable kernel, indiscriminately. And they have also started assigning CVSS scores with the same kind of malicious compliance.
Take for example, this patch in the device mapper RAID code:
Reasoning behind it being that theoretically, a RAID could be accessible over the network via NFS, iSCSI, etc... so if it gets corrupted, the buggy code path in the recovery (CVE-2026-89558) is effectively triggered over the network.
There's nothing malicious about the CVE. It lets people use it to index issues, like it was supposed to, and stops dumb people from using it as an indicator of work done or vulnerability level, things that it was never useful for.
"While many security people love to argue what is, or is not, a vulnerability, while dealing with CVEs, a CNA must follow the definition that cve.org gives us which is:
“An instance of one or more weaknesses in a Product that can be exploited, causing a negative impact to confidentiality, integrity, or availability; a set of conditions or behaviors that allows the violation of an explicit or implicit security policy.”
So with that definition in mind, the kernel CNA team members look at every bugfix that is added to the stable kernel releases and reviews it to determine if it meets this criteria."
So yes indeed, according to this, bugfixes are being examined for "instance of one or more weaknesses in a Product that can be exploited" - bugs.
I thought the "Security in the LLM age" talk by Greg Kroah-Hartman published this week from Kernel Recipes was pretty interesting: https://www.youtube.com/watch?v=NnV_cWeoo5Q
Extending the ratio from my other comment. What are the token ratios between using LLMs to:
1. Create functional software
2. Find bugs in functional software
3. Triage/Prioritise a quadrupling of the reported CVEs
4. Fix the bugs while keeping the software functional
I'm assuming that #2, #3, and #4 require more tokens (each or cumulatively) that #1, then there will be increase in the amount of insecure software because, with the advent of LLMs (that allow otherwise non-software developers to become software developers), there will be (a lot?) more software being created.
If we want software security to get better, then the existence of LLMs requires the increasing use of LLMs. I find this quite interesting.
1. LLMs have high false positive rate. From mythos 79 vulnerabilities found in the Linux kernel, only a single digit were actual bugs and they were all obscure so don’t panic.
2. What does obscure mean? I don’t really understand it, but many of the bugs have to do with custom network drivers or other custom drivers that are very specific to certain organizational setups, not a general Linux distro issue.
3. He’s very frustrated with the high false positive rate mythos generates. Even after multiple rounds of adversarial review and prompting strats, he mentions it is > 20% false positive rate, which wastes a lot of time. When some random user on the internet brings up a bug with an LLM it’s almost always fake, he even says just push back a few times claiming it’s not a bug to see if it’s a real bug (LLMs very quickly cave and “notice their mistake” etc)
4. General observation on the useful bugs mythos finds. Chain multiple smaller bugs to see if you can get a bigger breakage. Mythos is really good at constructing these long convoluted chains that fuzzers miss.
5. Go through recent bug fixes and check if similar bugs are hidden elsewhere in the codebase. Mythos is good at such pattern matching albeit with a high false positive rate.
Final conclusion: don’t panic, the bugs are getting fixed, this is not as bad as the first fuzzer bug mania and will be fixed quicker, he estimates a year and we won’t see huge bug reports anymore.
I use daybreak to do most of my security scanning. I always tell it that any finding must be accompanied by a harness that faithfully reproduces it using the code at the current commit with no modifications. In my experience it eliminates most if not all of the false positives.
>he mentions it is > 20% false positive rate, which wastes a lot of time.
If it's just 20% that is really low. Especially for complicated long chain potential bugs.
Most other detection tools have much higher rates of FP, or much higher rates of false negative.
Then you have humans that miss bugs for 20+ years. Or, they don't tell you about the things they thought were bugs they wasted hours on themselves. Because of this it's really hard to measure how bad/good the AI really is.
It would be interesting to know why the more SOTA models are getting the FPs. Is it from a lack of understanding of C? Is it complex code with deep branches? Is it code smell and convoluted logic?
It's an ~hour-long video! I don't blame anyone for preferring the 2-minute TL;DR if it's available! ¯\_(ツ)_/¯
If there's something wrong or prone to misinterpretation in the TL;DR it would be better to call it out in response to that, rather than the users responding to the TL;DR, simply because it's likely that's what most people will respond to.
He makes the exact opposite claim in the video. He thinks 20% is far too high and no one will pay for it once these tools are no longer available for free. He specifically cites the company Coverity that also found many useful bugs with their code analyzer but had a much smaller false positive rate (something like 5% I think), and no one paid for that, and the founder had to write a post mortem. He thinks the same’s going to happen to these LLM tools if they stop providing it for free.
If you push back against a LLM, you can convince it in almost anything.
If it's obviously a bug you can just fix it, and if it's obviously not, you can ignore it. In either case you've already done the work to understand it.
If it's on the borderline, and you push back against the GPU with a plausible sounding reason, it's likely to agree, regardless of whether oe nor it's a bug.
I've never gotten a good outcome out of arguing with the GPU.
> push back a few times claiming it’s not a bug to see if it’s a real bug (LLMs very quickly cave and “notice their mistake” etc)
I wonder how much more likely the LLM is to cave for a false-bug. There have been a number of times in my own (non-security) work that I'm told the LLM that it was wrong, it apologized and agreed, and then later that day I realized it was correct after all.
For additional context, Greg Kroah-Hartman has been contributing to Linux for 3 decades, and is _the_ maintainer for Linux's stable branch, as well as many other core parts of Linux.
My takeaways:
* Do not panic. Do acknowledge that LLMs are sycophantic, and LLM companies are trying to sell their stuff. "Yes, push back hard".
* If it smells like AI slop, treat it like AI slop and relax. If someone sends you 50 security reports, and a few look wrong, just calm down and ignore them; or ask for proof-of-humanness.
* A report without a patch is, for better or worse, worthless if you maintain widely used open source software.
* Of course, there are real vulnerabilities being discovered and reported. Keep calm, keep fixing real bugs, and carry on.
* Delete as much code as you possibly can. Reduce your surface area. Do the same thing you've always been doing.
An interesting observation that I encountered somewhere, I forget where, is that AIs when writing code introduce vulnerabilities at a rate similar to humans writing the same code. So we're looking at a massively accelerated volume of security vulnerabilities for the foreseeable future thanks to AI-assisted security research, and we can expect no reduction in new vulnerabilities from the AIs writing the code.
It doesn’t follow. Everyone sane has the models review the choose the models wrote. Reminder these are the models which found the Jacobian and Navier-Stokes counterexamples; they’ll find holes in their own slop, too.
Yes and no, many people don't have models review the code the same way many humans don't review their own code in depth for security vulnerabilities.
Furthermore, you need to make sure the model you use is capable enough to review your code comprehensively enough. That includes for both basic vulnerabilities but also attack chain related vulnerabilities.
These counterexamples are comparatively straightforward because the input domain is well-defined and simply-structured, and a counterexample is trivial to verify. The same is not true for arbitrary vulnerabilities.
Computers are finite. Inputs are well defined and so is their structure (ignore for a second the fact that it’s all physics behind the scenes). Secure code is a conjecture. A counterexample for secure code processing bits is an exploit.
Also I find calling millennium problem solutions ‘straightforward’ baffling, to be polite.
> Everyone sane has the models review the choose the models wrote. Reminder these are the models which found the Jacobian and Navier-Stokes counterexamples; they’ll find holes in their own slop, too.
It's not enough, though, to just tell the model to check the code for vulnerabilities. The model has to be guided specifically to look for particular classes of problem and that takes someone experienced in security.
Telling it to look for a particular class of problem just decreases needed context by decreasing the problem space.
That makes the bot more effective, but isn't strictly necessary. If you have a harness that can track longer projects it can do all that by itself, it needs your wallet, not your thoughts.
Do we know the code behind these vulnerabilities were written by AI? It seems like if anything AI was used to find exisitng vulnerabilities that would otherwise be used/sold as zero days and go unreported.
It was always the case that finding vulnerabilities in software was easier than writting perfect software. I'm hopeful that we can use AI to make software more secure over time. Project Zero [1] and others has shown many times the past few years (pre LLMs) that automated fuzzing and other forms of dynamic analysis are very effective, which bodes well for automated testing via LLMs!
I agree that more code = more bugs overeall, but there are slow moving codebases that run some of the worlds most valuable software. Seems like using AI to find vulnerabilities in that code is a huge win across the board.
> Do we know the code behind these vulnerabilities were written by AI? It seems like if anything AI was used to find exisitng vulnerabilities that would otherwise be used/sold as zero days and go unreported.
I was speaking in the general sense, not of these vulnerabilities specifically. I am of the view that AIs for the foreseeable won't produce code that is any better from a security point of view than something human written, so AIs will produce new vulnerabilities at least as fast as they find them and the rest of us will be faced with massive headaches like the one in the original post.
> It was always the case that finding vulnerabilities in software was easier than writting perfect software. I'm hopeful that we can use AI to make software more secure over time. Project Zero [1] and others has shown many times the past few years (pre LLMs) that automated fuzzing and other forms of dynamic analysis are very effective, which bodes well for automated testing via LLMs!
And yet, we are not seeing a drop-off in new vulnerabilities being discovered. We keep assuming that the list of bugs is getting smaller and we'll find them all eventually, but that is not the case for any software that I know of.
> I agree that more code = more bugs overeall, but there are slow moving codebases that run some of the worlds most valuable software. Seems like using AI to find vulnerabilities in that code is a huge win across the board.
It might be, yet, as I said just above, we are not seeing a drop-off in new vulnerabilities being found. The trickle of vulnerabilities has become a flood across all open source software, and already breakages and problems are occurring as maintainers struggle to keep up. Administrators, likewise, are struggling to keep systems updated. Just a week or so ago a security patch to rsync on RHEL broke rsync so completely that it could no longer handle symbolic links.
Critical CVEs used to be relatively infrequent, but they're becoming a weekly or even daily occurrence. None of us are prepared for this eventuality.
> I am of the view that AIs for the foreseeable [future] won't produce code that is any better from a security point of view than something human written, so AIs will produce new vulnerabilities at least as fast as they find them and the rest of us will be faced with massive headaches like the one in the original post.
If you simply prompt them to produce code, with the same kind of processes that humans use, then yes, of course. After all, it trained on human code.
If you prompt them explicitly to spend time looking for vulnerabilities and not implementing new features, then why wouldn't it produce more secure code? If we're calling the technology a "force multiplier", then it's thus for every task it can perform. So, orient the process around that; avoid the compromises that were originally motivated by working at human speed (, interest level, fatigue, specialization, …)
Of course, if you see places where the application of artificial "intelligence" can benefit from human wisdom, then double down on that. (Quotes because I think the term is fundamentally inaccurate for what it refers to, even though it's typically good enough and refers to a useful capability.)
The issues page on rsyncs Github. A lot has been fixed already obviously, but things are constantly breaking now.
To the point of other comments: Yes it might be a prompting issue, I don't know, but it does illustrate that the force multiplier people are suggesting that LLMs are, goes both ways. You can absolutely use them to make more secure software fast, but Andrew Tridgell isn't a stupid person. If someone like him can be seen struggling with the technology, then we must safely assume that this will be the case for many other developers as well.
> "Introduced while fixing CVE-2026-53801. I have likely found what the issue is, I truly hate symlinks Might not be able to fix tonight but will be done within the next 24 hours."
>And yet, we are not seeing a drop-off in new vulnerabilities being discovered. We keep assuming that the list of bugs is getting smaller and we'll find them all eventually, but that is not the case for any software that I know of.
It's worth factoring in that AI has gotten better rather quickly, so there's no reason to expect it not to continue to find new bugs even if we've correctly fixed what Mythos found. The search depth is increasing.
> AIs when writing code introduce vulnerabilities at a rate similar to humans writing the same code.
AI regurgitating all the insecure code AI companies scraped from stack overflow and github isn't going to give you something too different from what the humans who put it there in the first place came up with. Garbage in, garbage with random hallucinations out.
This is simplistic to the point of being blatantly wrong. Training data isn't garbage. It's programs that do their job, but that are sprinkled with errors. Uncorrelated errors gets averaged out during autoregressive pretraining. Correlated errors can be somewhat suppressed during post-training. Hallucinations (of the generalization-error kind) can be dealt with using synthetic data that improves the model's generalization.
Provable correctness is much harder than focusing on the happy path. If a pipeline doesn't include a look-for-the-vulnerabilities stage to save costs, a model wouldn't go out of its way to do it. The models are trained to do what they are asked to do.
I forget to add the obvious: some garbage gets through despite the mitigations. In the limit of pure garbage input, you'll get a model that internalized garbage generation. And you'd be better off throwing it away and starting over.
In my experience it doesn’t matter if who wrote the code was AI or human as long as you run security reviews, including AI reviews. They are extremely good at finding issues before you ship. Just don’t expect the LLM to one shoot secure code, let an empty context AI reviewer explicitly check the change set for vulnerabilities.
Quite the opposite, particularly when the biggest vendors of coding agents insist on not allowing their models to be used to check the code they generate for security issues.
Easily done if everyone can be convinced that learning to write code is pointless since you can just pay an AI company for access to a chatbot that will write it for you.
Fortunately there are people who write software for fun so there will always be some people who would rather do it themselves.
So, like professional orders in Europe? Thankfully everybody agreed that writing code is not engineering so this isn't mandated by law, but we were this close.
You don't need a permit or degree to draw up plans, just to write your name on the official version and get it implemented in the physical world. It's closer to deploying code than writing.
I've been able to use Opus 5.5 and Fable 5.1 for defensive security audits without any issues. They cannot do offensive tasks like pen-testing but in a lot of cases that's not a big shortcoming.
I've repeatedly slammed into walls doing very basic tasks. It can do a dumbed-down security audit, but fails to do offensive tasks against my own codebase which is frankly how you use a model like this effectively.
It just means that there will be the equivalent of infinite man-hours of barely-functional-intelligent-man to slog through code word by word and track how it affects all other code relation by relation.
It's not magic and it's not even better or even as good as a mid human, but it's something like infinite man-hours of that drudge work per hour per user.
That will find a lot in old code, and make it a lot easier to keep on finding every little thing right as it's created in new code.
If the Western AI companies get their way, only the developers/companies that have access and paid extra for the security review will get to have a lower chance of such issues.
You're going to get a wide range of responses on this but given the improvement in these models in just one year, and the number of bugs they're detecting which humans could not, I suspect that even median vibe-coded software is going to surpass median human coded software soon - if it hasn't already.
The important thing to remember here is there perfect isn't on the table. The benchmark is existing human-introduced bugs vs LLM-introduced bugs. Many developers have encountered odd bugs which a human would not have introduced, while forgetting about all the bugs caught which humans introduced. Or their opinion is formed by models from six months ago.
Mine isn’t. To point: we don’t really have a definition of vibe-coded anymore. All software has some degree of AI enhancement now. Is vibe-coded when it’s 60%? 80% 100%?
Try out Opus 5.5 on high. It’s shockingly good. Of course if you’re trying to one-shot a sprawling application with load balanced distributed DBs, you’re going to have a bad time. For small, defined features, it’s pretty fucking great.
Depends on if people are willing to pay the extra wall time to write mathematical proofs for everything to make it provably correct. Takes way way longer but modern models can do it.
It depends? LLMs are not a silver bullet for writing bug free software (IME at least, and of course it also depends on the "threshold" what actually counts as a bug). They're definitely good at not creating the trivial "mechanical" type of bugs a tired and overworked human programmer would create (but oth those are also the easiest to find with traditional debugging tools and testing).
They're definitely a useful additional tool for finding more (and more obscure) bugs, but that takes a lot of both human and compute effort too (quite a few of the reported bugs are actually false positives on close inspection, and apparently even with the latest locked down "wonder weapon" models like Mythos), and after all the reports are clean and validated you still can't be 100% sure (but at least a bit more confident) that the code is now free of bugs.
Because most devs don't want to do formal verified system development. I know only one OS working on that which is open source Ironclad OS hope many others follow this path. It is Ada/SPARK based but others should do with whatever language they are using. NetBSD also heard going to do something similar with C in last AGM
> most devs don't want to do formal verified system development.
That may be true. Serious question though: even if most devs wanted to develop formally verified code, do you think that it is reasonable to suggest that the typical systems developer could do it with today's tools? I don't mean verified protocols (TLA+) or verified algorithms (SPIN) I mean end-to-end verified code, a-la seL4. I got the impression that this is still very specialised work. Perhaps things have advanced since I last checked.
Do you make your professional career by building formally verified systems? I ask because I don't think that the reason comes down to "because most devs don't want to do formal verified system development". It's much more complicated of course.
I could believe it. Verified system development most likely comes with metric ton of paperwork.
Want to merge the PR? I need verified sign off in ServiceNow by staff level engineer. They are on vacation for 2 weeks? Did manager fill out delegation paperwork in ServiceNow with VP sign off? Oh they did but they forgot to put in return date AND time. Form needs to be corrected and reapproved before we can go into ServiceNow and make changes.
Genode as a whole isn’t formally verified, but it can use seL4 as a kernel, and it uses a robust capabilities system to sandbox basically everything, including drivers.
As it evolves I suspect there will be a push to verify more components of the stack. Once the capabilities layer can be verified, verification of most other components and drivers would become much less urgent.
I am an experienced engineer and am building a product actively right now. I started pre-AI, and cautiously adopted AI. I am using AI tools a lot now, but still dedicate a lot of my attention to validating it's output and design decisions.
I recently asked it to review my code and configs from the security perspective. Wow! 90% of the things it identified were MY bad decisions dating from pre-AI development. I am honestly humbled and impressed at the same time.
AI can create slop, and it can create quality products. It depends who is using it, and how.
While I totally agree with the statement, question is - is this challenge even solvable? Though where I come from there is a proverb - "for every malady there's a remedy", wonder what it can be this time.
And, of course, there are piles of legacy corporate spaghetti entangled in incomprehensible mess everywhere you look at. And this shit still runs, this precious hand-carved hand-weaved mess of bad decisions. I can't wait for LLMs to rewrite most of it.
As long as people use AI to review new stuff before they ship the ratio of vulnerabilities waiting should be trending downwards. Still, likely to be a bumpy ride.
No, it will be a tsunami of new discoveries in old code bases at first, but that will settle down as those old bugs are fixed. It's been like this with every new code analysis tool (the wave may be exceptionally high this time though).
There is a difference here. The "new coding analysis tool" that you use for the analogy here is getting better every few weeks with the release of new models.
It is plausible to assume that, for instance, a Linux kernel that was hardened for CVEs that Sonnet 3.5 could detect is not hardened for bugs that Sonnet 4.5, 5.5, Opus, Fable, and models in 2027 can and will be able to detect.
Hence, it is rather a constant catch-up game until the LLM improvements might hit a ceiling and won't get any better in this regard.
Search for "Security in the LLM age" in this HN page for a reality check, apparently 80% of the security vulnerabilities that Mythos initially flagged in the Linux kernel turned out to be false positives. Better than nothing of course, but it really doesn't look like Mythos is quite the "wonder weapon" it was marketed as (and the situation by far isn't as dire as the initial flurry of Mythos news).
You assume an infinite level of brokenness is legacy code bases. Granted, working on those things one can get the impression, but I think stable code converges to a low level of bugs after sufficient scrutiny and superhuman scrutiny doesn't reveal a never ending deluge of new problems. Not to mention that the bug-chains that you need for a successful exploit keep getting longer very quickly as more problems are discovered and the code is hardened.
I expect that the incremental model improvements will become smaller and smaller until they run into the same diminishing returns effect like all other new technologies (FWIW I've not beeen seeing a lot of difference between the latest Opus and Fable models for the stuff I'm doing, so I just stick to Opus for most things. There's a noticeable difference between Sonnet and Opus though).
I'm not ready to take any bets when exactly the curve will be flat though ;) (e.g. in a just couple of months or a couple of years)
It's normal practice for companies to have a backlog of scanner tool results. Sometimes in the thousands for a larger project. Many of them are legit bugs but also highly local and so far down the stack they're hard to exploit. It takes a ton of work to triage, more than management is willing to spend. Also more than they're wiling to fix and paydown.
I predict a different outcome: the rate of vulns identified and fixed will be more than matched by the rate of new vulns introduced by irresponsible use of LLMs on top of brittle and unwieldy tech stacks.
The result will be an overall increase in turbulence and the normalization of steadily intensifying security crises in nearly all software systems.
The only projects that will escape this fate are those which have either been already developed from ground up with rigorous and principled, verified (or verifiable) design, or those which are rewritten to gain this.
> AI is going to expose how fragile the entire computing infrastructure in our world is.
this is a good outcome. A forcing function to encourage all computing to be more secure can only be good in the long term, even if there's a lot of pain in the short term.
I, for one, am excited that we are now finally able to be "Secure™".
Wait, I guess I missed the "more", that kind of puts a damper on the whole thing.
Seriously though, security will continue to be an issue, always. Even if it was perfect, the benefits of it will not be applied uniformly. There will also be the same technology being improperly used causing new exploitables to go live.
> the benefits of it will not be applied uniformly.
why not? Any system you have permission to use and store your data should be beneficial to you if it became more secure. Unless...of course if you're the one who desires unauthorized access.
Clearly because everyone will not be able to pay for it to the same extent. Though, if you have a 1200 agent swarm to throw at a hf sized problem, I'd be excited to see your writeup.
I wonder what “a lot of pain” could mean here in a world where Crowdstrike is allowed to render half of the world unbootable without repercussions.
Not that I think you are wrong, I am sometimes just confused why we hold back on fixing security because of imaginary deployment- and business-related pains, when it is so obviously unproblematic to crash half the world for a day?
I don't like AI code, but AI is a decent reviewer.
...for human code.
In the small startup I work for boss (ex-programmer) discovered fable, and ai-coded 15K lines . So much productivity! So great! He even asked multiple reviews and it was fine!
I ask it a couple of reviews and it finds only minor things. The code is a mess of duplication and different coding styles, so I start cleaning it up. After a couple of months the reviews (same ai model) start actually finding big logic bugs that were always there.
We might already be at the point where the Ai-Coder is generating stuff that ai-reviewer can't find and will automatically pass.
--
1M context window is what? 70-80k LOC, tops? Without comments or documentation?
That is a smallish project of a couple of components. AI will remain inherently myopic until it can keep in context whole codebases.
Exposing current problems is fine to me, but I am worried of how brittle AI code will be.
That's because lots of websites and applications are being written by low-skilled laborers in Third World nations. Many websites are rife with vulnerabilities and insecure configurations.
It's not that difficult to build secure (web) applications but it takes effort and knowledge to get it right. You can't expect a web designer who can barely code in JavaScript to build a secure back-end, configure and maintain it. That's just asking for trouble.
Even high-value sites are built by cheap laborers these days. LLMs (I refuse to call it A.I.) will expose their weaknesses within minutes.
I put weatherstripping on my front door and made part of my house much more comfortable and easy to heat in the winter. Then I bought an FLIR and found lots more places leaking heat. I made a list, hired a handyman, and I'm more comfortable and saving money.
Should I now expect to find even more places to seal up the next time I break out the FLIR?
Or you could categorize the current bug apocalypse under the heading "unsustainable trends will not be sustained."
The main way this analogy doesn't hold up to me is that your house is mostly static -- you're not rebuilding the walls, adding new doors or windows constantly.
But in software, especially with agents, we're constantly renovating the house. If you were renovating every 6 months I'd expect to find more places to seal up, even if you were following best practices in those remodels.
That being said, I do think we will reach an equilibrium where most vulnerabilities are found at PR time.
Some software, especially human facing products in competitive markets, will always have new security issues. But in general I agree. We are headed toward a different and generally more secure and reliable equilibrium. Hopefully it's the end of stupidly long bug lists in some products.
This is great, more access did provide more eyes on these problems.
But, does that all of these being found now call into question, not the open source model logic itself, but the ability of human eyes to find security issues? These vulnerabilities have been sitting here for however long, but how many thousands of humans did not find them before AI?
Even before AI we have known that no one is smart enough to write bug free C. And with every bug being a launch platform for a full exploit it's become a big deal.
This is the reason why performant and "safe" systems languages are on the rise and being accepted into foundational areas of our operating systems (such as the kernel). When the attack defense surface is tighter (language, tooling, compiler instead of the actual code itself), the smart folks can stay at that layer, while the masses can write more code at a level of abstraction that nullifies many of these vulnerabilities by default.
I can't find it now, but I heard that NASA developed processes to produce completely error-free code (for the moon landings iirc). The problem is that it's incredibly time consuming and expensive to do.
Like everything in CS, apparently this is a trade-off, not an absolute. You can get bug-free code, but it's not commercially viable and is extremely tedious to do.
It’s probably about how they write the space shuttle software, and it’s quite a famous article. The original is now paywalled but there are many copies.
This is example of what I call "the tipping paradox", and I'm almost sure that this phenomenon has its own proper scientific name.
When asked, people prefer €15 burger no tip, but when actually making a choice, they prefer €10 burger with €5 tip. Similarly, companies state "bug-free code" as a goal or requirement, but then they prioritize other goals over code correctness. My workplace is in the process of completely removing code reviews. And actually, I don't disagree with the decision - my career is short, but I have never seen reviews fulfill any purpose other than to share the blame in case of an incident.
They are removing reviews or just human reviews? It’s interesting and I kind of agree to a certain point that human review is not the best way to find and fix bugs. We’ve had so many bugs in code that was fully reviewed by at least 2 humans. Ai review is clearly superior. Human reviews main purpose to me right now is to spread know of what is being worked on.
The threshold for Microsoft and Apple to actually report vulnerabilities is MUCH MUCH higher than for Linux, and open source as a whole. They generally only disclose issues in Windows and macOS that are quite serious and impactful.
For Linux, the threshold is nearer to the point of it being questionable whether a bug is even exploitable on a real production distro, compiled and run with any sort of sane configuration.
Some vulnerabilities are incredibly complex and are identified not through the code itself but through various attack chains being strung together.
Also there are likely a lot of vulnerabilities identified but the work required to fix them vs the complexity to exploit them means they don't get fixed.
I'd wager we don't have an issue with identifying vulnerabilities but the ability to fix them.
I work in security consulting and identifying vulnerabilities isn't the difficult part its actually fixing them and fixing the ones that have valid exploitable attack chains that matter
That’s probably a great thing. The initial friction of AI overwhelming projects certainly sucks, but once there are better processes to deal with them it’s going to strengthen the quality of so many projects!
Long term we will end up with software with no low hanging fruit exploits left. But right now we are in a period where low hanging fruit is everywhere and it's easier to exploit systems than ever before.
That’s assuming we’re not adding software defects at the same pace, but I imagine we are generating a lot more defects than are being discovered at the moment.
> we are generating a lot more defects than are being discovered
Wouldn't this mean people are actually encountering issues where there are none before? Likely some serious enough that they lead to exploits where there weren't before? Where are the reports of these new defects?
I just mean there is a proliferation of new code. New code == new defects. It’s likely that popular software projects get the majority of the scrutiny, while no one is spending tokens looking for defects on my 0-star GitHub repo.
That relies on a major assumption that current systems are finding the vast majority of all possible exploits out there. This assumption itself would assume that either current LLMs are near perfect, or that the peak difficulty for exploits was just above human capability (which is where LLMs currently are). I think both of those assumptions are very likely false. If so then we'll see indefinitely ongoing exploit discovery as LLMs improve their capabilities.
Long term I suspect that the purpose of the digital domain is going to end up being rethought. For instance connecting critical infrastructure to the internet has always been a terrible idea, and LLMs will just make that even more clear.
Computers are becoming architecturally more secure alongside just patching bugs. We have seen the move to using VMs with a minimal hypervisor, using separate security chips to hold sensitive info like encryption keys, Memory Tagging to detect and prevent memory exploits.
On the iphone for example even if you find a crippling bug in iOS which gives you full root access, there is still no way to get the device encryption key or face ID info because the secure enclave simply has no electrical connection that can pass that key to the OS.
I fear the vast majority of projects never went with any kind of code review but somebody high on Red Bulls at 9 PM looking at the code and saying “looks good to me”.
And this is the best of cases. I fear many times in a smaller projects it was “it compiles at last, let’s see if anybody complains”. That’s why LLM’s are so damn effective today.
(been there, done that – I’m not pointing fingers, but we are human, we get tired, and we don’t have the NASA budget to complete things and ship them)
If any of these were remotely exploitable it would be getting much louder and more urgent news. Local privilege escalation bugs are encountered all the time.
One thing I’m very proud of is that, in the AI era, only three minor security problems were found in my open source project (knock on wood):
• Two remote denial of service attacks against the DNS-over-TCP service (the DNS-over-UDP service did not appear affected), which is disabled by default.
• One network leak of 19 bytes of unallocated memory on the heap. Looking at those 19 bytes, no information of note appears to be present there.
Note that, pre-AI, there were a couple of remote memory leaks and two remote “packet of death” bugs found, but nothing earth shattering has been found since the beginning of these AI-assisted security audits.
Considering that a lot more bugs have been found in the Linux kernel, there is a reason I don’t blindly trust its /dev/urandom to always return completely random numbers I can safely use, and I don’t think the kernel has done a better job implementing a CSPRNG (cryptographic secure pseudo random number generator) than I did.
In Heroes of Might & Magic 3, “several” means 5–9. 10–19 is “pack”, 20–49 is “lots”, 50–99 is “horde”, 100–249 is “throng”, 250–499 is “swarm”, 500–999 is “zounds…” and 1000+ is “legion”.
I suggest this post be renamed “A legion of vulnerabilities has been discovered…”
Maybe we should use legion as a collective noun for vulnerabilities. Like murder of crows or school of fish. Would signal the urgency no matter the actual number or impact...
Hah! I remember playing that game as a young non-English-speaker and having no idea what these qualifiers meant. They never give the scale, it's just assumed you know that several < pack < lots < horde < throng etc. Which.. even as bilingual as I am today I'd struggle to put on a scale.
https://docs.kernel.org/process/cve.html states that because almost any kernel bug can potentially compromise system security, the CVE team acts with extreme caution and labels nearly all bug fixes with a CVE
yes because the majority are memory safety issues, and it's automatically assumed that a memory safety bug can lead to a vuln
one again illustrating the importance of encapsulating unsafe behavior. perhaps c should get a __UNSAFE { } block, where memory access is encapsulated and thus most bugs occurring outside of those blocks do not need to be marked as CVEs.
No. It's because Greg doesn't like the CVE system and MITRE, the stupidest decision ever, made Greg a CNA, and this is his tantrum that he's been waiting 40 years to throw.
C has many more ways to trigger undefined behaviour than memory access. If C had unsafe blocks they'd restrict most forms of signed integer arithmetic and shifting, for a start.
I wish everyone used the same versioning system for all software: A.B.C increment C for security patches, increment B for new features, increment A for systemic changes.
Well, this is called semantic versioning (semver) and it is the most popular versioning system out there: https://semver.org/
However I see software variants where simple numbers work as well, e.g. Firmware code that runs on totally self contained hardware is often just versioned with simple integers, that is because most updates (releases) of that software contain both features and bugfixws and "breaking changes" don't apply to stuff that doesn't read or write files. So you could use semver and have 1.50.0, 1.51.0, 1.52.0 forever, but by that point you can just give it aimple integer versions.
Linux kernel introduces bug fixes, new features, and major systemic changes all the time. That’s why using SemVer for Linux is pointless and why Linus dropped it.
I was looking earlier and the majority of these do not have a CVSS score assigned to them yet, but a lot of them that did were >7.0 (although I suppose by nature that the more impactful CVEs are going to be scored more quickly)
What is the token ratio of 'AI writing software that works' to 'AI finding vulnerabilities in software that otherwise works'?
Even with AI there's an asymmetry separating 'functional' from 'secure'.
As with software development prior to AI, security has a cost associated with it, and unless there are _real_ consequences, it's a cost that most software houses ain't gon' pay.
> security has a cost associated with it, and unless there are _real_ consequences, it's a cost that most software houses ain't gon' pay.
Good to see someone say this. Software need only be good enough to do the job, and only as safe as reasonably required. Systems can be secured and audited through other means, sometimes at lower expense than guaranteeing every line of software has zero risk associated.
Is it possible that some of these bugs were already exploited by governments? And AI might help use close of that kind of thing? (And expose other kinds of things , in a kind of AI arms race?)
Probably. Organizations specializing in hacking are probably having a great time hacking everything and installing permanent presence. It is probably like a gold rush period.
I've been doing Linux sys-admin work for 10+ years. Used to be I could read the full report on what ever vulnerabilities came out and triage which servers needed to be updated now and which could wait. A few years ago notifications started having so many it would take more time/effort to read everything than it would to patch everything. With this notification there's even ten times more...
The amount of updates on Debian "stable" has become ridiculous, it's a daily rolling release of backports at this point.
Microsoft kinda got this right by doing it once a month, unless it's something horribly bad, you can plan your maintenance around a predictable calendar.
>Note, due to the layer at which the Linux kernel is in a system, almost any bug might be exploitable to compromise the security of the kernel, but the possibility of exploitation is often not evident when the bug is fixed. Because of this, the CVE assignment team are overly cautious and assign CVE numbers to any bugfix that they identify. This explains the seemingly large number of CVEs that are issued by the Linux kernel team.
We're considering weekly reboots at work now with the pace of kernel updates coming out, and the speed with which vulnerabilities are getting exploited.
There is exactly one CVE in the entire list that is high severity, and it affects an obscure IBM NIC driver for big iron systems used in companies with more money than brains.
Problem is it takes more effort to read the CVE list and work out if you have one of the drivers impacted loaded than it does to just update the kernel.
Sorry but I'm not causing outages and rebooting systems every four days because the people whose job it is to do this can't triage properly. They are actively making other's lives more difficult. There's like 100 CVEs in that list and not a one of them is important to most systems. Even worse, if there was 1/100 it's a needle in a haystack.
If your system goes down to update a kernel then you have a major issue already. This is like manually renewing https certs. When the task is done often enough you just automate it in a painless way.
Well then, it's a business decision to accept the risks. IMO it makes sense that a percentage of machines would always be offline for maintenance. If you're running bare metal and you cannot afford 10% of your machines being offline at most times, then you're in deep shit if there's even a tiny traffic irregularity, and it's a sign that your management is YOLOing the company. Not uncommon though.
severity on cve is a crapshoot most of the time, but especially with linux cna. i would not advise relying on them for decision making.
"We can not assign severity
[...]
So any group that attempts to give a “severity score” to a Linux CVE is lying to you, UNLESS they know exactly your use case.
ALWAYS ignore any attempt that groups such as NIST/NVD that purport to assign things like CVSS scores to a vulnerability. Those numbers are false and give companies a “fake sense of security”."</i>
Regarding the fart question, it depends on the audio driver and the underlying codec. While you would think “free as in beer” would make for truly resonant flatulence, only truly letting loose (plus a Bic lighter) brings real enlightenment.
getting those fixes is simply bumping the kernel build to the latest patch level, the distros must do it in a timely manner.
also, roughly only 10% of the affected files are compiled into a kernel image; from this number i would say only 4 needs attention per week. But again, the reasonable approach is simply building the latest patch version
Many of these are months old, why are they being announced as if they were brand new?
I mean honestly, CVE-2026-23137 doesn't affect kernels after January (6.18.6 and beyond), so even LTS kernel users should have the patch for this for several months.
Don't get me wrong, informing people is good, this is just a strange wall of every kernel CVE this year - the kernel has a lot of CVEs every year (moreso lately with AI scanning/research, but still).
Greg KH said this week in his talk at Kernel Recipes: yes CVEs are increasing, no reason to panic, a lot of the LLM CVE fear mongering is overstated [1].
This is a more valuable insight/signal than a CVE enumeration, especially given that Linux CVEs are literally assigned to any bugfix (as others explained as well).
Security has always been a cat and mouse game. The usual precautions about software updates, least privileged access, separation/sandboxing, and precautions like not pasting/installing random software (especially something on Show HN) continue to apply; but at this point, I'd suggest thinking more towards a model of "How do I best detect, and contain a compromise" model.
One thing I do is a wallet.dat honeypot with (now) ~$1700 of bitcoin. Any movement essentially shuts down my home network from external access. At worst, it's a justifiable tax deduction, not capital gains :)
We need to switchover to microkernel operating systems ASAP or our entire computing infrastructure becomes a liability.
There are good reasons for QNX becoming viable again in the automotive world. Linux / Android has so many vulnerabilities that it needs indefinite patching, which is unrealistic for any computing device but cars especially. Car makers are switching to QNX even though it costs them money.
Weren't there rumours that intelligence agencies have been using microkernel operating systems for decades?
>Weren't there rumours that intelligence agencies have been using microkernel operating systems for decades?
Governments, military, aviation, space.
For a good portion of the usage, the people involved cannot even talk about it. But sometimes it's possible. A surprising amount of real world use showed up in talks in seL4 summit 2026[0].
You cannot possibly compare embedded/single-purpose systems to a general-purpose OS that is expected to run arbitrary user-written software performantly on thousands of different hardware combinations.
Alternatively, we could benefit from security through compartmentalization. Qubes OS runs everything in VMs, and there were two VM escapes in the last 20 years [0]. A dedicated VM for your email doesn't allow any attacker to compromise it, since you do not run anything else in it.
They found out that it wasn't prudent to use Linux / Android in their core vehicle software after cars were hacked and miscreants were able to control the brakes, steering and accelerator.
I suppose they're only using QNX for critical functions. The infotainment system will probably continue to run Android.
I think we will soon reach a point where code won't be considered secure unless AI has written or at least verified it. Company policies may then even consider that bad practice and/or forbid it.
Is there a way to know if a particular vanilla kernel has a particular CVE addressed? Unhelpfully, the ChangeLog-* only seems to contain sporadic references to CVEs.
I instructed it to not violate any robots.txt or cause any sort of floods. As such it did a general sweep without pulling the full texts or doing a complete analysis. I’d use it as a rough guide.
My take is that we didn’t see 1000+ easily exploitable RCEs.
Anyone here who wanted an unreliable summarization from a chatbot could have just asked a chatbot for one. If the count/breakdown of issues is important to someone, it's probably important to them that it's accurate. That means they'll have to take the time to verify it properly anyway.
someone did want an unreliable count from a chatbot, and did ask for it.. and they passed the information along in their comment (along with a very upfront comment about how the generated it). Feel free to come with other methods to generate unreliable summarization if you don't like the chatbot generated one. I'm sure humans and off by one errors or maybe badly formed regex's can give you a whole range of new things to complain about.
We know that humans are flawed and make errors. Those human generated errors are an intrinsic part of "conversation between humans". We both expect and accept that when we read posts here because, as the guideline states, "HN is for conversation between humans."
People who are respectful enough to disclose their use of AI in the first place will hopefully be respectful enough not to fill this space with AI slop after being informed/reminded that it isn't welcome.
"Several" feels a bit of an understatement, there are 1313 CVEs listed on that page!
Wonder how many of these NSA and others been sitting on, for how long and how many are still there? I guess the silver lining with the aixplosion of CVEs is that software eventually will get more secure.
Something to keep in mind is the Linux project registered as an authority to create their own CVE numbers in 2024. Previously the majority of bugs would just be fixed without note unless there was a demonstration that it could be exploited.
NSA allegedly used to have a "black budget" of around a couple dozen million dollars for software sabotaging. I wonder what percentage of those CVEs could be related to it...
I checked out the last (highest-numbered) one just as a quick sanity check:
> In the Linux kernel, the following vulnerability has been resolved:
> usb: typec: ucsi: unregister debugfs entries on teardown
> ucsi_register() creates per-instance debugfs entries, but ucsi_unregister() keeps them around until ucsi_destroy().
> Drivers like ucsi_glink that unregister/register the same UCSI instance across remoteproc restart then try to create an already existing debugfs directory and log:
> debugfs: 'pmic_glink.ucsi.0' already exists in 'ucsi'
> Unregister debugfs entries as part of ucsi_unregister(), and clear ucsi->debugfs after freeing it so repeated unregister paths remain safe.
I'm going to need someone to explain how that could possibly become a "vulnerability".
if open source is struggling with this volume, what is the likelihood that if commercial software would also struggle if they were scanned and exposed?
I mean Microsoft, 2026 Sept Patch Tuesday: 964 CVEs...
Linux has never, at any point in history, lacked flaws that could be exploited to escalate privileges. The only question has been how well-known the flaws were, and when. The count of latent local privilege escalation bugs has never been zero.
All the good backdoors have got to be in the hardware anyway. There's only a handful of companies making the chips all our systems run on. US companies that can be forced to backdoor their stuff under secret gag orders. I mean, I'm not saying that PSP/CSME subsystems are backdoors, but if I were going to force companies to add one, it'd probably end up looking similar. The blackbox wireless chipsets in all our phones also come from a very small number of companies.
In the end they will rewrite all low-level code in Rust anyway. Because it will be crystal clear to anybody, that we can't keep all the C/C++ at the base of the modern world.
I head a much, much smaller open source project. Since the November Singularity we've been seeing at least six responsibly reported security advisories a month. However, this last month we had 22 unique security advisories. Our project has been built with adherence to the OWASP Top Ten Guidelines and other best practices from the beginning. But software is hard and AI is thorough.
Each month, we fix them all in our monthly maintenance release and disclose at that time. We fight AI fire with fire, and hand-review, of course.
So far, we can keep up. One hopes this is possible at the scale of the Linux project, which assuredly has more humans and more AI to throw at the problem. But team size does not scale linearly with interested audience, and potential bugs do scale with codebase size (and other extremely important factors, like code quality, at which the Linux team is assuredly much better than we are).
("November Singularity" is a cheeky reference to the arrival of Opus 4.5 and "good enough" coding models and harnesses generally.)
I too use November 2025 as a true turning point.
Eternal November
The internet went from being dominated by academics, to normies, to bots.
It was. That's when I stopped coding most things by hand. The difference in what I could reliably get AI to do for me in February 2025 vs February 2026 is just massive. I immediately became a huge advocate amongst my peers for going all-in on automated engineering because this curve is about to get very steep and if you're not staying ahead, you might get left behind.
I'm genuinely curious: if the trend is toward "AI accomplishes my goals more easily than in the past", what curve are you "staying ahead" of?
Why isn't even easier for an AI noob to jump in at the next step with less friction than the current one?
the whole "get left behind" trope (as well as their entire comment tbh) is a classic ai psychosis thing
Well it's a good thing that isn't what I said or implied, and that you instead smugly misinterpreted my comment.
The person you're replying to had the decency to ask me for clarification; for you, I'd suggest brushing up on your reading comprehension and learning to respect the rules of this site, which include interpreting comments in their best light and focusing on positive, substantial contributions instead of negative, inflammatory posts.
maybe don't say things like "if you're not staying ahead, you might get left behind" since, like I said, it's a trope at this point.
Again, I'm going to refer to the advice I just laid out for you. I'm not responsible for you and I don't need to cater my words to your lack of reading comprehension.
I think this is both true and untrue. For a long time, I was an AI skeptic, perhaps even a hater. A chauvinist for writing code by 'hand.' But it's become very difficult, as I've gradually migrated to AI-authoring of code, to go back to writing code by hand and retain the same velocity.
Part of this is certainly that my hand-authoring code skills have atrophied, sure, but my workflow has also radically changed. Previously, I would spent a lot of time and focus on a single work item, and only context switch to other tasks whenever I would wait on CI or a long build. It meant that I spent a lot of time understanding one thing at a time, and interruptions (forced context switches) incurred a massive switching cost.
Now, having moved to largely AI-authored code, I find myself necessarily working on multiple threads at the same time. This means I can meaningfully progress each of those threads in parallel, with a much-reduced overhead on context switching, since I don't have my head down focusing on all the details of the work. And if the task really demands it, I can still stop and focus on one thread to sketch out the code manually, think about the concepts more deeply, etc.
It's a very different workflow, and there are certainly downsides, but the upside is that the rate of which I've been able to put up good-quality PRs has measurably increased. It's not quite 2x, and it's certainly not 10x, but it's definitely noticeable. I do understand a bit less, but there was always more work than time to understand things in full detail. I guess only time will tell if that missing understanding was actually vital to the long-term success of my work.
yes that's fine, i have a similar experience. but there is a difference between saying what you did, and the type of AI religious fundamentalism where you got "converted" when truely good AI coding models revealed themselves to you, and then you go around trying to save those who would otherwise be "left behind".
AI is becoming a more and more powerful lever that lets you move greater loads per application; it is in a real sense giving you more leverage. But the lever is difficult to use effectively, and the methods for using the level change every week or so.
> methods for using the level change every week or so
hmm doesn't mean noob will become better? if what I'm "learning" about AI is going to get deprecated every week... then it doesn't make sense to "keep up"
Here's how I think about it. There's two different things "knowing how to use AI effectively" and "knowing how to leverage effective AI to improve your project".
1. Both are multipliers on skills you already have. A "noob" will be at a disadvantage.
2. The second one is a complex blend of technical skill, domain knowledge and human factors (knowing your user-base, knowing your product/project, understanding UX/DX/whatever) that you probably will always have an edge on.
It requires skill to prompt the AI to do best possible work. I have senior colleagues that produce horrible slop and others that always produce great results. Managing the context, knowing when to stop the AI. Knowing what to question, what to "trust" in is a big deal.
Also knowing what models are good for what. There was time half a year ago when Google genuinely had a better model than everyone else. I used it for everything and 3 weeks later they lobotomise it (sorry "optimised") and I went back to opus...
I think Anthropic is like a drug dealer, giving us the sweet sweet drug for free (I don't think our $200 a month subscriptions even cover the electricity for our use) and the time to pay will come very soon...
I expect this subscription will cost $2k a month. Will it be a normal increase? Or will the "enshittify" existing models to the point you'll pay $2k to get fable 6 to do what opus 4.8 did fine in July 2026?
I was referring to the amount of people in the future who might be employed at today's engineering wages, and the differential between those people and the average employed engineer today.
What will that differential be? A large component will be social reinforcement: generational wealth, connectedness, and such. Things like merit might take more of a backseat. So people who do not have the necessary social capital may be fighting each other for very limited amounts of positions.
That is the steepening of the curve for people like me: I was homeless at 16 and finished high school on my own, was given a full ride to LSU, grants, room and board and a job in the Comp Sci department, but lost all of it after an immature and vindictive high school teacher illegally modified my grade in a core class in order to fuck me over. I didn't have parents to back me up at the school board and make things right.
Instead I suffered through years of homelessness and had to find my own path into the industry by starting companies with friends and doing all of the engineering. Since then I've led multiple teams, made some connections, shipped a lot of cool stuff and bring to the table a wealth of experience and a generalist skillset that is both wide and deep. Yet, I too wonder where my place in this changing industry will be once things settle a bit. Probably less engineering and more focus on business development.
The flipside though is as you've said: The fruits of engineering are more accessible than ever to the layman, and individuals can currently possess an unprecedented amount of agency and leverage. I think that is amazing and am fully behind it. I do know that it means the process of renormalization is going to be very rough, given the similarly unprecedented rate of industry change these technologies are bringing.
There are two types of AI users: those who chase the rabbit and those who do not. Some people bounce from tool to tool QuantumLeap-style hoping that maybe the next tool will be the big thing to solve all their issues. Other people use AI to do actual work. Once they find a tool that works, they stick with that tool until they feal a need to upgrade.
It is like people fishing. Some people go out and catch fish with the tools they know will work. For other people, every day at the lake requires a new boat/rod/lure. They spend more time figuring out how to use their new toy than they do catching fish.
Because improved models seem to increase, not decrease, the gap between what you get from really expert supervision (prompting, steering, hand-corrections along the way, etc.) vs what you get from fairy-dust and wishes.
AI will eventually get better than you at all you do, because it trains on your inputs.
Secret knowledge/data will be tomorrow's gold.
Yeha me too but I need to remember Opus 5.5 aka September/October because i wouldn't expected a model feel again relevant different but it does.
I like the November Singularity, I hope that catches on. That was definitely the point where I went from "AI is overhyped" to "oh shit the hype bros may be on to something"
Yes, good name and I also feel that was a major inflection point. Up until then I found all AI models to be terrible at programming, with the difference between GPT o3, Sonnet 4 and every other model since GPT-3 being just the exact flavor of terrible they were.
November 2025, with Opus 4.5, was the first time I was impressed by an LLM doing something non-trivial with a reasonably good level of quality.
Yes, and Qwen 3.8 27b is a similar inflection point for local, although I wouldn't claim it's quite as good as Opus 4.5 it is genuinely useful. Whether that actually matters will depend on whether it ever becomes the most cost-effective tool for the job, but it's impressive as all heck
Looks like someone coined it almost immediately! https://agi.co.uk/november-singularity-ai-agentic-era/
>> We fight AI fire with fire, and hand-review, of course.
Wouldn't it be nice if AI vulnerability reports came with AI pull requests to fix them? The thinking context that found it should be readily able to propose a fix. It would still need review but even when AI PRs aren't right they often point in the right direction.
Why? Imagine that you are a competing, close source product/project. Just bombard your competition with AI reports, and let them drown in misery. Problem solved! (/s)
This just goes to show that yes, if security researchers were to do that, it would great, but they are not the only actors here...
I'll speak up in defense of the reporters: FWIW, so far my strong impression is we're hearing from independent security researchers. Right now, for us, so far (enough qualifiers yet?) the system is working for us: independent security researchers are farming reputation by finding real problems. That's not a bad thing.
And in most cases they do propose solutions, although we generally resolve the issues on our own.
There's a small percentage where we make the case that the ticket is not a real vulnerability, and then we have to grit our teeth through repeated reports of the same "vulnerability." But it's a small percentage so far.
We do typically have to reconsider the severity. The researchers understandably want to see everything as a nine...
That doesn't really help it's just more stuff to review and question if it makes any sense at all
Yeah just like with Navier Stokes.
If it's a PR created by just prompting with the vuln report, I might as well do it myself.
wonder how much this can easily speed up the malicious-contributer attack. where someone suggest a security fix in a extremely obscure and irrelevant code, but the fix actually adds a new condition that then can be exploited elsewhere.
Absolutely possible, which is why I refuse to take my eyes off code review (mine, or a few other trusted souls)
I wonder if August of this year will be remembered like that too. The very first "opus like" local AI model came out this August (Qwen 3.8 Flash Next). I've been running it locally since for real programming and I consider it pretty much the same as opus 4.6 in coding ability (it lacks a bit in the factual knowledge area). It even exceeds opus on some tasks.
I've tried every hyped "open source" model before and all including latest models bigger than 1T parameters are pretty much toys.
This is the first one that isn't. It can't be overstated how huge of a deal that is. No more reliance on Anthropic.
Running this model to do real work is still not cheap. I run it on a pc with 6 rtx3090s and 192GB of RAM (and I use 90gb of that ram for KV cache). The model is entirely on gpus. It runs at 55tok/s 1500 prefill for one user at a time, and about 35tok/s 650 prefill per user for 5 simultaneous users. It doesn't seem like much until one realises you manage your own kv cache. You ca leave your sessions in cache for as long as you want. You can save them and restore 200k sessions a week later in a dozen seconds.
What many people don't realise is that usage of those models skews extremely heavily towards input processing. My claude code usage is about 1.3B tokens input per week and only about 8M output. On claude code I get 80% cache. At home it's more like 95%.
If this progress keeps up, and we get a fable quality model in a year in under 200B to run at home... Those "frontier labs" will be renting all of their gpus per hour not to go bankrupt.
What Quants do you run?
I am using qwen3.8-27B-UD-Q3_K_XL on a 5080 16gb card with 16gb of system ram at around 43-45tok/s with 65k context. I am using that model to set up containers on a proxmox server with two b70s and 96gb ram. The Q3 model implemented 10 different chat models, 2 image models, and 7 web based harnesses so I can compare. It can rewrite any part to improve it.
qwen3.8 is more than capable for software dev. The gemma models were terrible and fell apart during compaction. I think as long as the llm can test the results, you don't need anything being offered by a "frontier" model. The cloud AI is going to be used by people unable to run their own and they will eventually get squeezed on price.
The one tip I could give is compact before starting new steps or any action in the plan that is different than what was previously worked on. You want to manage what is in the context and don't want unnecessary work history details filling it up. You can always ask the llm to list the current plan, then compact after and do it more than once until you get the compaction <30%. This will be fixable by the harness that can choose better times to compact.
You don't want to start a new phase and have it compact a few minute after starting. This happening over and over again seems to potentially cause issues for long running sessions. Compaction slop that screws up what is in the context.
I can compare what I use at home vs paid models like astra at work and the difference is mostly meaningless.
For me the November Singularity was when ChatGPT was released in November 2022. Of course, it was a faaaaaar cry from what we see today, and it took a lot of careful wrangling, but it could write reams of correct code and tests even back then.
The thing was the wrangling was relatively straightforward, if cumbersome. Largely, it involved being very precise with the context and instructions it was given. I could imagine a lot of that getting automated (i.e. what we today call harnesses) or recursively addressed by creative meta-prompting. Supported by similarly conceptually simple advances like chain-of-thought reasoning I suspect that is the biggest thing that the models have figured out what to do today compared to then: manage themselves carefully.
Although I could not have predicted these exact outcomes, the implications for everything that is unfolding now were clear even then.
Seems to me we need tooling to do automatic offensive security as soon as a new frontier model comes out, with quickly turned around patches using the same frontier model. Rinse and repeat. Virtuous agentic security loop.
As soon as a new model drops I point it at one codebase I have and ask it to do a security audit, look for bugs, check for optimizations, things that are over-engineered, etc.
Every single time I've done this it has found at least one serious bug or security hole.
Makes me wonder how many are left I've not found.
Isn't that part of the cloud AI business model now? Businesses have to pay for early access so they can weed out any new bugs before the model goes public. Which is kind of crazy, because you are paying for early access to protect yourself from other paying customers of the same cloud AI model.
Fun part is July 2026 „summer of bliss” for cURL project.
As much as cybersec forums were outraged that everyone will be hacked because of that — nothing happened.
I hope that's also a reference to the November Revolution :D
If this year has taught me anything, it's that there was never anything like stability in software and that stability is the wrong optimization goal.
Instead we should aim for quick updates, strong isolation and sandboxing.
Whatever that means for TDD and other methodologies that seemingly all have failed to encode guardrails in the development workflow.
Do you think that eventually one would 'fix all holes' or is this an eternal fight against the windmills?
note that _any_ bugfix is assigned a cve, which makes for big numbers.
>“Due to the layer at which the Linux kernel is in a system, almost any bug might be exploitable to compromise the security of the kernel… Because of this, the CVE assignment team is overly cautious and assign CVE numbers to any bugfix that they identify.”
https://docs.kernel.org/process/cve.html
"number of cves" is a useless metric, especially when it comes to the kernel.
With particular emphasis on "almost any bug might be exploitable".
Even a bug-free program might be exploitable.
That sounds like a bug
There are programs like sudo whose entire reason for existing is to enable privilege escalation. If you can find a way to make a user "sudo" something, that's an exploit, but it's not a bug in the program.
At that point you're exploiting the user, who is not a bug-free program
Remove user and press any key to continue.
:-P
Indeed, the user is the hardest part of the program to secure.
But if you squint, it might be a bug in the system to allow access to that program.
That's also not exploiting the program, it's exploiting the user.
Alternatively, consider any program which loads dynamic shared libraries (called "plugins" in many contexts). The program itself might be bug free; any plugin that is loaded will run (typically) with the full priviledges and access of the program (and thus likely the user).
The user may have no idea that the plugin is malicious; the program remains bug-free (if it was beforehand).
At least with plugins most people have a general notion that they run some sort of code as they provide some new functionality.
Much more agregiously you can design harmless a looking format that can run arbitrary code e.g. .doc with VBS. I have a very hard time blaming an user who falls for that even though MS puts up a scary looking popup.
Nope, sudo is safer, the alternative is running everything as root.
Totally a feature
Tautologically every bug can legitimately be assigned a CVE, since every bug prevents some feature from working as intended. It's therefore a denial of service, which by the definition of the CVE system using CVSS means every bug is at least a 1/Low level vulnerability to CVSS v4.0.
If you're willing to stretch, missing but planned features also deny the use of said features since they haven't been added yet, and so are CVSS 1/Low vulnerabilities.
Resume-driven development for security researchers has never been easier!
> It's therefore a denial of service
That doesn't follow. In the extremely simple example, an adding service returning 1+1=3 has a bug, but it's not a possible DoS situation at all.
> missing but planned features also deny the use of said features
That's not what DoS is.
This whole situation with CVE assigning comes from the whole process being far from ideal. But it doesn't mean it's completely useless and doesn't follow any rules at all.
>but it's not a possible DoS situation at all.
Until someone finds there is a user input they can trigger this bug causing some other bit of code to read data from the wrong offset and now it's a whole exploit.
That's an issue in the other code, not in the addition service. It would be lumped together if it was an addition function close to the other code. But I wrote service there on purpose.
Why couldn't a crash caused by filesystem corruption caused by a + operator that says 1+1=3 be filed as a DoS?
It could. But it's a CVE in the system that crashed or in the filesystem, not in the calculator web service that we were discussing. If a filesystem decides to use a bad online calculator for its internal logic, that's a vulnerability on the filesystem, not the calculator.
We aren't talking about a calculator web service though we are talking about the Linux kernel and if there's a bug in the kernel that could conceivably cause a crash in an otherwise correctly written application then that would be a CVE in the kernel.
It's not hard to imagine an application for which 1+1 = 3 leads to security issue.
https://news.ycombinator.com/item?id=49929389
That’s not what Mitre thinks though, they are very happy to host a 9.8 severity CVE for 1+1=3. They’ll probably publish one for 1-1=0 too, if you preface with ”The users of mathematics might not be prepared for zero values”
This is why you need a library for additions. At least CVEs can be tracked appropriately, rather than the developer rolling out their NIH solution.
There’s probably an npm library for that. Not all heroes wear capes.
You're memeing on bad cve handling, but that's so far outside of what the real issues are, it just doesn't make sense. How about telling people about what the real problems with vulnerability classification are, rather than mitre=bad?
The real problem with vulnerability classification is that mitre=bad
https://daniel.haxx.se/blog/2026/06/24/a-cve-dispute/
Not really. cURL developers just have NIH syndrome.
Organizations that track their software and patch systems for CVEs have their own risk management system.
Publish xlow as low, we'll filter them out if we want to. As even they state in that article, they do NOT have the necessary context to filter stuff out. So why are they doing it?
From experience: many companies have a “risk management system” that involves nagging OSS projects to do free work for them, even when the advisory is manifestly nonsense or has no impact in context. Many teams have a “green dot” mindset, and CVE directly encourages that behavior by stapling CVSS scores to identifiers.
> cURL developers just have NIH syndrome.
What does this even mean? cURL is one of the most load-bearing pieces of software in existence. It, and the Linux kernel, which takes a similarly dim view of the CVE system, are the inventors. "Not Invented Here" seems to imply that there is a vast body of peer work for them to draw on to resolve this problem, but who are their peers? As far as I can tell, the answer is something like "Microsoft and Apple", on the one hand, who exist in a totally different, mostly closed-source or at least closed-development, ecosystem, or something like "glibc and OpenSSL" on the other hand, which have their own storied CVE history.
Tell us about a system for managing vulnerability database that works across enterprise, private and open source, is staffed enough to research both impact and disputes, is funded enough to work for decades, isn't partisan to any industry interests, can classify vulnerabilities in a non ambiguous way, can handle public submissions at any volume in a timely manner...
and one that will never make decisions that is incorrect or disliked by any party.
>hash of supplied password does not match hash of stored password, access denied
Resulting from miscalc of either stored or supplied would indeed create DoS.
If the off by one is in a graphical driver that causes an overflow leading to no output, that's DoS.
If the off by one is in the memory mapping of input devices leading to no available input, that's DoS.
Simple math is kind of everywhere in the kernel and userland apps. If the math is wrong and results in memory mapping wrong such that kernel panics or is unuseable, that's a breaking bug regardless of simplicity.
Loads of bugs aren't CVE-worthy. If you tell the computer to make a light green but it makes the light red, that's a bug but no DoS or other CVE-worthy bug.
However, the Linux kernel is supposed to run any userland program without crashing, so anything that crashes the kernel is a local DoS and there are a lot of them. It's also supposed to shield processes from each other and maintain privilege levels correctly, so many incorrect memory leaks are also CVE worthy. Whether a CVE applies depends on the people and programs using the kernel, and the kernel team can't read your code to tell you if it applies or not.
People reading CVEs wrong ("it's got a high number so we must patch within a day") must be going crazy over this, but the point of CVEs is to let you make judgement calls, not to be a cool statistic about how secure something is.
Most CVEs are irrelevant to most people, that's always been the case.
"why did the ship crash?"
Unpatched bug, wrong color lights.
Yes, it is a bit of a stretch, I desperately hope programmable navigation lights are not a thing. And I also don't think every bug needs a CVE. But in the correct context nearly any bug could be critical.
Computer Weekly in the UK did some pioneering journalism into a helicopter crash[1] that had initially been blamed on the two pilots but seems very likely to have been caused by a failure of a computer system that was controlling the fuelling of the engines. Iirc this system seems to have crashed causing engine failure of both engines in heavy fog, bringing the helicopter down with a loss of everyone on board, however for reasons somewhat unclear, the MOD wanted to cover the failure up by blaming the pilots.
Now this was a windows system[2] rather than linux but the point remains - if there was an external vulnerability in a crucial control system and this system was part of a network (eg to connect telemetry) then any exploit of that system could result in loss of life.
[1] https://www.computerweekly.com/news/1280091718/Chinook-compu...
[2] Which, why? Why build the fuel controller for a helicopter engine on windows?
Are you sure that the Chinook used Windows? I have not read this elsewhere.
There have been documented problems with Windows for Warships.
My memory is ancient so may be flawed but that is what I remember. This seems to be a full chronology of the incident if you’re interested https://www.computerweekly.com/news/1280096804/Chronology-Th...
A windows system controlling the refuelling ? Worked in aviation software for a while and the certification for flight control (or similar) systems is subject to rigorous path testing/inspections/approvals etc DO-178C (level A or B likely). Not sure if any windows OS is certified to level A?? Typically certifiable RTOS'es are procured for those purposes
That's still not cause for a CVE, even if it's a bad bug.
Unless that computer happens be on traffic lights. Will this become CVE? Human life would be at risk.
This is actually the point of the parent comment: you have to read the CVE and see for yourself if it impacts your specific system or not.
If only the user can tell whether something is potentially dangerous, then either every bug or none of them should get a CVE. There are countless systems out there that are commonly used in ways beyond what even the developer intended, how should a third party authority like the CNA be able to discern this?
They can't. Which is why security discussions are such a hot mess.
The vendors fixing them arguably prioritize these reports right. Most of the CVEs, even severe ones, are irrelevant in practice, and as parents note, are more like regular bugs with security flavor in reporting. The CVE label instead of regular bug tracking number makes them seem important.
Linux Kernel may be one of the few legitimate exceptions, indeed, due to the position in which it sits in the software stack. Also LLMs make previously unexploitable-in-practice vulnerabilities exploitable (by making targeted / personalized attack cheap enough to give them positive ROI), which complicates things.
I think you have to make a distinction between bugs and actual vulnerabilities. Therac-25 killed people, but I wouldn’t consider anything about it to be a security vulnerability. In my mind, the distinction between a bug and a vulnerability is that a bug is triggered during “normal” operations and can do anything. Whereas a vulnerability requires an adversary to “trigger” the vulnerability, and can do this to achieve some cognizable malicious goal. There’s probably some overlap on the edges; whether an issue in a library is a bug or a vulnerability may depend on how it is used, for example.
But I think the most important thing to keep in mind is that a bug isn’t necessarily less serious or less important than a vulnerability. A serious bug should be patched just as urgently as a serious vulnerability.
Unless it is the led showing the status of a camera.
Yeah in that case it's a feature and management can proudly proclaim "we own the glass"
> but the point of CVEs is to let you make judgement calls
Realistically, most admins cannot make judgment calls about 1000+ CVEs for a kernel release.
Usually the security team mandates a zero CVE policy on all deployments and the organization complies.
we in security teams can only dream of such incompetent management
zero CVE policy = halt on business development.
And not only that, remember the hn crowd is not at all representative of the average.
Even more realistically, many admins do not have the background to be able to reason (by themselves) about the actual risk of most CVEs, so just going along with specialized media coverage is often a sound strategy.
>And not only that, remember the hn crowd is not at all representative of the average.
We should keep our hopes up, someday we may get there.
> note that _any_ bugfix is assigned a cve
I do find it interesting though, that in the interest of transparency, every bugfix gets a CVE. Which ends up being a huge number… which will ultimately yield a more insecure environment as we’re getting conditioned to ignore/discount CVEs by the volume.
Over-reporting in this case seems to risk being counterproductive.
Depends on the end consumers stance on security. I've watched it shift from "Only update if we can prove we are impacted" to "Update everything immediately just in case".
The frequency and severity of cyber attacks has increased to the point a much more cautious approach has become common. It's also easier to sell this work to management when you can point at the security tab on some tool and say "Look we need to patch these CVEs"
"Update everything immediately just in case" is a lot less attractive when you see more downtime from updates breaking things than you do from hackers. Windows updates are an endless source of pain, but now every program seems to demand to be updated practically daily. Even things that you shouldn't have to think about like keyboards, mice, and printers beg to be updated all the time.
Downtime is annoying but workable. You can't unleak customer data after your system gets hacked.
There is no reason, why a newer version has less bugs, than an older version. Both are essentially an unknown number. The only thing you know, is that you likely know a higher percentage of bugs for the older than for the newer version.
It certainly has less known bugs. And when you have to make a statement to the media, “we were hacked by an undiscovered 0 day exploit” sounds a lot better than “we were hacked by a known exploit because we didn’t update”
If you do know about it, you can also just mitigate that.
I do this this is theoretical, especially with recent spread ups.
> It certainly has less known bugs.
This isn't necessarily true, and will be less true in the coming age of LLM-automated vulnerability scanning. The version that you're downloading (after being nagged for a day) that adds Feature A may contain 3 vulnerabilities that are already known before you even download the update, and may or may not fix old vulnerabilities.
Yeah, the fact that security forces you to update has been used with great effect by product managers and feature engineers to shove their changes down everyone's throat and quickly drop support for previous versions. Neat for them.
> I've watched it shift from "Only update if we can prove we are impacted" to "Update everything immediately just in case".
> [...] a much more cautious approach has become common.
I'd argue that this is a less cautious approach, not more. It takes time to carefully evaluate, review, and test each change.
It's not very interesting. Linus, and by extension the Linux kernel, long had a dismissive attitude toward security research. This is basically a childish swing from one extreme (nothing gets a CVE) to another (everything gets a CVE).
Kernel development is well-funded, both via grants and by direct employment at big tech companies, and if they wanted to properly triage and annotate vulnerabilities, and provide reasonable assessments of what is or isn't likely to be a security risk, they absolutely could. They almost certainly could go to Google and say "we need two people full-time on your payroll for this" and they would get it.
I don't want to dunk on them too much because they're generally doing God's work, but these absolutist security stances are not worth being taken seriously.
It's basically saying that they can't possibly provide a valuable service for 99.999% of the install base because there might a hypothetical person out there using Linux in a really weird way. If Microsoft tried to make an argument like that, they'd get crucified.
The counterpoint is that by putting CVEs on bugs that are more easily exploitable provides a roadmap for attackers. Of course, in the current LLM age that's probably a moot point, but that could be the reason for this.
security by obscurity?
Obscurity is a valuable layer of defense-in-depth
Let's say they do and only 5% are "security issues" you still need to update either way.
I don't think that Linus is dismissive of security, it is that he is very much a proponent of always rolling to the latest stable release.
Linux only ever wanted to promise support for the latest release and even Linux LTS is a concession.
And CVEs are basically a useless concept if you roll. (or at least not any more useful than any other bug tracker which supports tags)
> Linux only ever wanted to promise support for the latest release and even Linux LTS is a concession.
If he kept true of his "we don't break user space" instead of it being "we don't break user space until we do and then it's on you to deal with it" perhaps more people would be willing to run the latest release.
I have had way more issues provoked on RHEL-compatibles by RHEL's Frankenstein backporty kernel than I ever have on Arch or Nix by the latest stable kernel.
Linux LTS is much more about proprietary drivers targeting a stable internal kernel API/ABI than about anything else.
I mean… good for you. How does that help me when my software crashes because they changed some API?
Which userspace API was changed?
Internal Kernel API changes all the time which can break proprietary drivers. But userspace API has a far higher guarantee.
setsockopt changing behaviour and crashing processes.
What user space APIs have they been routinely breaking?
The ones to manage IP filtering rules for example. Those APIs are even security-critical.
setsockopt that do nothing easily go to crashing the process in later versions…
> if they wanted to properly triage and annotate vulnerabilities, and provide reasonable assessments of what is or isn't likely to be a security risk, they absolutely could.
I'll challenge this, I don't think that's really possible, at least not to any high degree of confidence. Nobody can realistically evaluate if any particular out-of-bounds read/write or use-after-free is "safe", and if you're going to consider all of those as security risks then there's not really a point in trying to filter out the few bug fixes that might not lead to those things.
Try going to the linked page, pick any random CVE, and read it. I've checked a bunch and I'd say _at least_ 8/10 of them are variations on those two things.
They chose to do it this way, in. part, because CVEs were already like that, and too many people were pretending they weren't.
Counterproductive for whom? The stance of the kernel developers is that whether a bug is a vulnerability or not depends on intended use. They say, we do not dictate use, therefore this decision is out of scope for us. This makes kernel development more focussed and productive.
I would argue that it is more productive for the enduser as well. Not making a decision they cannot reasonably make is better then blindly believing in a decision that is likely wrong for your usecase.
> _any_ bugfix is assigned a cve
Any patch that is back ported to a stable kernel, indiscriminately. And they have also started assigning CVSS scores with the same kind of malicious compliance.
Take for example, this patch in the device mapper RAID code:
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
After back porting to stable, it got assigned CVE-2026-89558 (which is in this list), and a CVSS score of 9.8:
https://git.kernel.org/pub/scm/linux/security/vulns.git/tree...
Reasoning behind it being that theoretically, a RAID could be accessible over the network via NFS, iSCSI, etc... so if it gets corrupted, the buggy code path in the recovery (CVE-2026-89558) is effectively triggered over the network.
There's nothing malicious about the CVE. It lets people use it to index issues, like it was supposed to, and stops dumb people from using it as an indicator of work done or vulnerability level, things that it was never useful for.
Uhm I'm pretty sure a severity score is supposed to be a score of severity.
> Because of this, the CVE assignment team is overly cautious and assign CVE numbers to any bugfix that they identify.
Er what?? CVEs should be assigned to bugs, not bugfixes, right?
the whole process is talked about here (and the companion posts): http://www.kroah.com/log/blog/2026/02/16/linux-cve-assignmen...
but no, in linux cve id is assigned "on a one to two week delay from when the fix has landed in a released stable kernel version."
Thanks.
"While many security people love to argue what is, or is not, a vulnerability, while dealing with CVEs, a CNA must follow the definition that cve.org gives us which is:
“An instance of one or more weaknesses in a Product that can be exploited, causing a negative impact to confidentiality, integrity, or availability; a set of conditions or behaviors that allows the violation of an explicit or implicit security policy.”
So with that definition in mind, the kernel CNA team members look at every bugfix that is added to the stable kernel releases and reviews it to determine if it meets this criteria."
So yes indeed, according to this, bugfixes are being examined for "instance of one or more weaknesses in a Product that can be exploited" - bugs.
Oh dear, oh dear.
I thought the "Security in the LLM age" talk by Greg Kroah-Hartman published this week from Kernel Recipes was pretty interesting: https://www.youtube.com/watch?v=NnV_cWeoo5Q
500 CVEs per release to 2000+
Extending the ratio from my other comment. What are the token ratios between using LLMs to:
1. Create functional software
2. Find bugs in functional software
3. Triage/Prioritise a quadrupling of the reported CVEs
4. Fix the bugs while keeping the software functional
I'm assuming that #2, #3, and #4 require more tokens (each or cumulatively) that #1, then there will be increase in the amount of insecure software because, with the advent of LLMs (that allow otherwise non-software developers to become software developers), there will be (a lot?) more software being created.
If we want software security to get better, then the existence of LLMs requires the increasing use of LLMs. I find this quite interesting.
Yeah it’s a great video (TLDR from the video)
1. LLMs have high false positive rate. From mythos 79 vulnerabilities found in the Linux kernel, only a single digit were actual bugs and they were all obscure so don’t panic.
2. What does obscure mean? I don’t really understand it, but many of the bugs have to do with custom network drivers or other custom drivers that are very specific to certain organizational setups, not a general Linux distro issue.
3. He’s very frustrated with the high false positive rate mythos generates. Even after multiple rounds of adversarial review and prompting strats, he mentions it is > 20% false positive rate, which wastes a lot of time. When some random user on the internet brings up a bug with an LLM it’s almost always fake, he even says just push back a few times claiming it’s not a bug to see if it’s a real bug (LLMs very quickly cave and “notice their mistake” etc)
4. General observation on the useful bugs mythos finds. Chain multiple smaller bugs to see if you can get a bigger breakage. Mythos is really good at constructing these long convoluted chains that fuzzers miss.
5. Go through recent bug fixes and check if similar bugs are hidden elsewhere in the codebase. Mythos is good at such pattern matching albeit with a high false positive rate.
Final conclusion: don’t panic, the bugs are getting fixed, this is not as bad as the first fuzzer bug mania and will be fixed quicker, he estimates a year and we won’t see huge bug reports anymore.
I use daybreak to do most of my security scanning. I always tell it that any finding must be accompanied by a harness that faithfully reproduces it using the code at the current commit with no modifications. In my experience it eliminates most if not all of the false positives.
And normal people who can not apply for access to daybreak, should we use Abliterated GLM-5.3 ? [0] [1] [2]
[0] https://news.ycombinator.com/item?id=49605691 [1] https://news.ycombinator.com/item?id=49897075 [2] https://huggingface.co/models?other=offensive-security
>he mentions it is > 20% false positive rate, which wastes a lot of time.
If it's just 20% that is really low. Especially for complicated long chain potential bugs.
Most other detection tools have much higher rates of FP, or much higher rates of false negative.
Then you have humans that miss bugs for 20+ years. Or, they don't tell you about the things they thought were bugs they wasted hours on themselves. Because of this it's really hard to measure how bad/good the AI really is.
It would be interesting to know why the more SOTA models are getting the FPs. Is it from a lack of understanding of C? Is it complex code with deep branches? Is it code smell and convoluted logic?
Maybe you should just go ahead and watch the video instead of throwing hundred of hypothesis. It'll take just as long and you'll be informed by then.
It's an ~hour-long video! I don't blame anyone for preferring the 2-minute TL;DR if it's available! ¯\_(ツ)_/¯
If there's something wrong or prone to misinterpretation in the TL;DR it would be better to call it out in response to that, rather than the users responding to the TL;DR, simply because it's likely that's what most people will respond to.
He makes the exact opposite claim in the video. He thinks 20% is far too high and no one will pay for it once these tools are no longer available for free. He specifically cites the company Coverity that also found many useful bugs with their code analyzer but had a much smaller false positive rate (something like 5% I think), and no one paid for that, and the founder had to write a post mortem. He thinks the same’s going to happen to these LLM tools if they stop providing it for free.
There are a number of successful companies around doing just that, it's this guy that wasn't successful at it, not everyone.
If you push back against a LLM, you can convince it in almost anything.
If it's obviously a bug you can just fix it, and if it's obviously not, you can ignore it. In either case you've already done the work to understand it.
If it's on the borderline, and you push back against the GPU with a plausible sounding reason, it's likely to agree, regardless of whether oe nor it's a bug.
I've never gotten a good outcome out of arguing with the GPU.
> push back a few times claiming it’s not a bug to see if it’s a real bug (LLMs very quickly cave and “notice their mistake” etc)
I wonder how much more likely the LLM is to cave for a false-bug. There have been a number of times in my own (non-security) work that I'm told the LLM that it was wrong, it apologized and agreed, and then later that day I realized it was correct after all.
Excellent plug and recommendation!
For additional context, Greg Kroah-Hartman has been contributing to Linux for 3 decades, and is _the_ maintainer for Linux's stable branch, as well as many other core parts of Linux.
My takeaways:
* Do not panic. Do acknowledge that LLMs are sycophantic, and LLM companies are trying to sell their stuff. "Yes, push back hard".
* If it smells like AI slop, treat it like AI slop and relax. If someone sends you 50 security reports, and a few look wrong, just calm down and ignore them; or ask for proof-of-humanness.
* A report without a patch is, for better or worse, worthless if you maintain widely used open source software.
* Of course, there are real vulnerabilities being discovered and reported. Keep calm, keep fixing real bugs, and carry on.
* Delete as much code as you possibly can. Reduce your surface area. Do the same thing you've always been doing.
An interesting observation that I encountered somewhere, I forget where, is that AIs when writing code introduce vulnerabilities at a rate similar to humans writing the same code. So we're looking at a massively accelerated volume of security vulnerabilities for the foreseeable future thanks to AI-assisted security research, and we can expect no reduction in new vulnerabilities from the AIs writing the code.
It doesn’t follow. Everyone sane has the models review the choose the models wrote. Reminder these are the models which found the Jacobian and Navier-Stokes counterexamples; they’ll find holes in their own slop, too.
Then a new model comes out two weeks later and finds a hole that your archaic review bot missed
Fortunately, for now. Imagine the models not being released publicly.
Yes and no, many people don't have models review the code the same way many humans don't review their own code in depth for security vulnerabilities.
Furthermore, you need to make sure the model you use is capable enough to review your code comprehensively enough. That includes for both basic vulnerabilities but also attack chain related vulnerabilities.
These counterexamples are comparatively straightforward because the input domain is well-defined and simply-structured, and a counterexample is trivial to verify. The same is not true for arbitrary vulnerabilities.
Computers are finite. Inputs are well defined and so is their structure (ignore for a second the fact that it’s all physics behind the scenes). Secure code is a conjecture. A counterexample for secure code processing bits is an exploit.
Also I find calling millennium problem solutions ‘straightforward’ baffling, to be polite.
> Everyone sane has the models review the choose the models wrote. Reminder these are the models which found the Jacobian and Navier-Stokes counterexamples; they’ll find holes in their own slop, too.
It's not enough, though, to just tell the model to check the code for vulnerabilities. The model has to be guided specifically to look for particular classes of problem and that takes someone experienced in security.
Telling it to look for a particular class of problem just decreases needed context by decreasing the problem space.
That makes the bot more effective, but isn't strictly necessary. If you have a harness that can track longer projects it can do all that by itself, it needs your wallet, not your thoughts.
Do we know the code behind these vulnerabilities were written by AI? It seems like if anything AI was used to find exisitng vulnerabilities that would otherwise be used/sold as zero days and go unreported.
It was always the case that finding vulnerabilities in software was easier than writting perfect software. I'm hopeful that we can use AI to make software more secure over time. Project Zero [1] and others has shown many times the past few years (pre LLMs) that automated fuzzing and other forms of dynamic analysis are very effective, which bodes well for automated testing via LLMs!
I agree that more code = more bugs overeall, but there are slow moving codebases that run some of the worlds most valuable software. Seems like using AI to find vulnerabilities in that code is a huge win across the board.
[1]: https://www.google.com/search?q=site%3Aprojectzero.google&q=...
> Do we know the code behind these vulnerabilities were written by AI? It seems like if anything AI was used to find exisitng vulnerabilities that would otherwise be used/sold as zero days and go unreported.
I was speaking in the general sense, not of these vulnerabilities specifically. I am of the view that AIs for the foreseeable won't produce code that is any better from a security point of view than something human written, so AIs will produce new vulnerabilities at least as fast as they find them and the rest of us will be faced with massive headaches like the one in the original post.
> It was always the case that finding vulnerabilities in software was easier than writting perfect software. I'm hopeful that we can use AI to make software more secure over time. Project Zero [1] and others has shown many times the past few years (pre LLMs) that automated fuzzing and other forms of dynamic analysis are very effective, which bodes well for automated testing via LLMs!
And yet, we are not seeing a drop-off in new vulnerabilities being discovered. We keep assuming that the list of bugs is getting smaller and we'll find them all eventually, but that is not the case for any software that I know of.
> I agree that more code = more bugs overeall, but there are slow moving codebases that run some of the worlds most valuable software. Seems like using AI to find vulnerabilities in that code is a huge win across the board.
It might be, yet, as I said just above, we are not seeing a drop-off in new vulnerabilities being found. The trickle of vulnerabilities has become a flood across all open source software, and already breakages and problems are occurring as maintainers struggle to keep up. Administrators, likewise, are struggling to keep systems updated. Just a week or so ago a security patch to rsync on RHEL broke rsync so completely that it could no longer handle symbolic links.
Critical CVEs used to be relatively infrequent, but they're becoming a weekly or even daily occurrence. None of us are prepared for this eventuality.
> I am of the view that AIs for the foreseeable [future] won't produce code that is any better from a security point of view than something human written, so AIs will produce new vulnerabilities at least as fast as they find them and the rest of us will be faced with massive headaches like the one in the original post.
If you simply prompt them to produce code, with the same kind of processes that humans use, then yes, of course. After all, it trained on human code.
If you prompt them explicitly to spend time looking for vulnerabilities and not implementing new features, then why wouldn't it produce more secure code? If we're calling the technology a "force multiplier", then it's thus for every task it can perform. So, orient the process around that; avoid the compromises that were originally motivated by working at human speed (, interest level, fatigue, specialization, …)
Of course, if you see places where the application of artificial "intelligence" can benefit from human wisdom, then double down on that. (Quotes because I think the term is fundamentally inaccurate for what it refers to, even though it's typically good enough and refers to a useful capability.)
> broke rsync so completely that it could no longer handle symbolic links
That sounds bad. Where can we find more information about this?
The issues page on rsyncs Github. A lot has been fixed already obviously, but things are constantly breaking now.
To the point of other comments: Yes it might be a prompting issue, I don't know, but it does illustrate that the force multiplier people are suggesting that LLMs are, goes both ways. You can absolutely use them to make more secure software fast, but Andrew Tridgell isn't a stupid person. If someone like him can be seen struggling with the technology, then we must safely assume that this will be the case for many other developers as well.
https://github.com/RsyncProject/rsync/issues
See https://github.com/RsyncProject/rsync/issues/1087
From the comments:
> "Introduced while fixing CVE-2026-53801. I have likely found what the issue is, I truly hate symlinks Might not be able to fix tonight but will be done within the next 24 hours."
>And yet, we are not seeing a drop-off in new vulnerabilities being discovered. We keep assuming that the list of bugs is getting smaller and we'll find them all eventually, but that is not the case for any software that I know of.
It's worth factoring in that AI has gotten better rather quickly, so there's no reason to expect it not to continue to find new bugs even if we've correctly fixed what Mythos found. The search depth is increasing.
> AIs when writing code introduce vulnerabilities at a rate similar to humans writing the same code.
AI regurgitating all the insecure code AI companies scraped from stack overflow and github isn't going to give you something too different from what the humans who put it there in the first place came up with. Garbage in, garbage with random hallucinations out.
This is simplistic to the point of being blatantly wrong. Training data isn't garbage. It's programs that do their job, but that are sprinkled with errors. Uncorrelated errors gets averaged out during autoregressive pretraining. Correlated errors can be somewhat suppressed during post-training. Hallucinations (of the generalization-error kind) can be dealt with using synthetic data that improves the model's generalization.
So with all those mitigations why do LLMs keep writing insecure code?
Provable correctness is much harder than focusing on the happy path. If a pipeline doesn't include a look-for-the-vulnerabilities stage to save costs, a model wouldn't go out of its way to do it. The models are trained to do what they are asked to do.
I forget to add the obvious: some garbage gets through despite the mitigations. In the limit of pure garbage input, you'll get a model that internalized garbage generation. And you'd be better off throwing it away and starting over.
I guess trash in trash out also applies to the training of an LLM.
In my experience it doesn’t matter if who wrote the code was AI or human as long as you run security reviews, including AI reviews. They are extremely good at finding issues before you ship. Just don’t expect the LLM to one shoot secure code, let an empty context AI reviewer explicitly check the change set for vulnerabilities.
Yet another case of https://en.wikipedia.org/wiki/Neijuan
Just the beginning. AI is going to expose how fragile the entire computing infrastructure in our world is.
Does this also mean the code generated and reviewed by LLMs will not have such issues going forward?
It won't which means inevitably malicious AI will probably backdoor us. Damned if we do, damned if we don't.
Quite the opposite, particularly when the biggest vendors of coding agents insist on not allowing their models to be used to check the code they generate for security issues.
Why stop there? We should regulate who is allowed to write code in the first place!
Easily done if everyone can be convinced that learning to write code is pointless since you can just pay an AI company for access to a chatbot that will write it for you.
Fortunately there are people who write software for fun so there will always be some people who would rather do it themselves.
So, like professional orders in Europe? Thankfully everybody agreed that writing code is not engineering so this isn't mandated by law, but we were this close.
You don't need a permit or degree to draw up plans, just to write your name on the official version and get it implemented in the physical world. It's closer to deploying code than writing.
We have a lot of that in Europe. Parasite professions. Someone who does nothing but charges a lot for his signature.
“And that was how CS became a real engineering degree”
They sort of already do; the commercial American providers all require you to be 18 to sign up for a plan capable of agentic coding.
So, I guess people under 18 aren't allowed to learn to program anymore.
I've been able to use Opus 5.5 and Fable 5.1 for defensive security audits without any issues. They cannot do offensive tasks like pen-testing but in a lot of cases that's not a big shortcoming.
Any tips / online resources how to best utilize for defensive reviews?
I've repeatedly slammed into walls doing very basic tasks. It can do a dumbed-down security audit, but fails to do offensive tasks against my own codebase which is frankly how you use a model like this effectively.
It just means that there will be the equivalent of infinite man-hours of barely-functional-intelligent-man to slog through code word by word and track how it affects all other code relation by relation.
It's not magic and it's not even better or even as good as a mid human, but it's something like infinite man-hours of that drudge work per hour per user.
That will find a lot in old code, and make it a lot easier to keep on finding every little thing right as it's created in new code.
If the Western AI companies get their way, only the developers/companies that have access and paid extra for the security review will get to have a lower chance of such issues.
I'm an arms race the only winner is the arms-dealer.
You're going to get a wide range of responses on this but given the improvement in these models in just one year, and the number of bugs they're detecting which humans could not, I suspect that even median vibe-coded software is going to surpass median human coded software soon - if it hasn't already.
The important thing to remember here is there perfect isn't on the table. The benchmark is existing human-introduced bugs vs LLM-introduced bugs. Many developers have encountered odd bugs which a human would not have introduced, while forgetting about all the bugs caught which humans introduced. Or their opinion is formed by models from six months ago.
How come all the vibe coded stuff I've tried is totally bug ridden and unmaintainable then?
Mine isn’t. To point: we don’t really have a definition of vibe-coded anymore. All software has some degree of AI enhancement now. Is vibe-coded when it’s 60%? 80% 100%?
Try out Opus 5.5 on high. It’s shockingly good. Of course if you’re trying to one-shot a sprawling application with load balanced distributed DBs, you’re going to have a bad time. For small, defined features, it’s pretty fucking great.
The definition I use is shipping LLM generated code you don't understand
We use LLMs extensively on the projects I work in. We don't "vibe code", and we understand every commit.
Depends on if people are willing to pay the extra wall time to write mathematical proofs for everything to make it provably correct. Takes way way longer but modern models can do it.
It depends? LLMs are not a silver bullet for writing bug free software (IME at least, and of course it also depends on the "threshold" what actually counts as a bug). They're definitely good at not creating the trivial "mechanical" type of bugs a tired and overworked human programmer would create (but oth those are also the easiest to find with traditional debugging tools and testing).
They're definitely a useful additional tool for finding more (and more obscure) bugs, but that takes a lot of both human and compute effort too (quite a few of the reported bugs are actually false positives on close inspection, and apparently even with the latest locked down "wonder weapon" models like Mythos), and after all the reports are clean and validated you still can't be 100% sure (but at least a bit more confident) that the code is now free of bugs.
IMO, they will introduce their own class of bugs that will defy static analysis.
in theory, it could be the best thing that ever happened to open source.
I'm not super sure about that. But if its going to exist I'm crossing my fingers it works to the OSS community benefits (eventually)
wouldn’t it result in whoever has the most money having the most secure software?
Eventually whoever has the most energy.
I don't think AI will make that true more than it is now. More investment in security should on average lead to better security, with or without AI.
Because most devs don't want to do formal verified system development. I know only one OS working on that which is open source Ironclad OS hope many others follow this path. It is Ada/SPARK based but others should do with whatever language they are using. NetBSD also heard going to do something similar with C in last AGM
There is also LionsOS https://lionsos.org/ (SeL4-based).
> most devs don't want to do formal verified system development.
That may be true. Serious question though: even if most devs wanted to develop formally verified code, do you think that it is reasonable to suggest that the typical systems developer could do it with today's tools? I don't mean verified protocols (TLA+) or verified algorithms (SPIN) I mean end-to-end verified code, a-la seL4. I got the impression that this is still very specialised work. Perhaps things have advanced since I last checked.
Do you make your professional career by building formally verified systems? I ask because I don't think that the reason comes down to "because most devs don't want to do formal verified system development". It's much more complicated of course.
I could believe it. Verified system development most likely comes with metric ton of paperwork.
Want to merge the PR? I need verified sign off in ServiceNow by staff level engineer. They are on vacation for 2 weeks? Did manager fill out delegation paperwork in ServiceNow with VP sign off? Oh they did but they forgot to put in return date AND time. Form needs to be corrected and reapproved before we can go into ServiceNow and make changes.
https://sel4.systems/ is relevant in this space!
Genode as a whole isn’t formally verified, but it can use seL4 as a kernel, and it uses a robust capabilities system to sandbox basically everything, including drivers.
As it evolves I suspect there will be a push to verify more components of the stack. Once the capabilities layer can be verified, verification of most other components and drivers would become much less urgent.
Instead of patching millions of environments, perhaps it’s better to just build a secure one from scratch?
That’s what I did with Safebox: https://safebots.ai/about/infrastructure.html
Who said those vulnerabilities were found by an LLM?
And the aspie award goes to…
> fragile the entire computing infrastructure in our world is.
It was "load bearing" just fine....
Anything is "fragile" if you put a bulldozer over it....
i mean if the building needs to be bulldozer proof..
But it doesn't need to if you lock up the bulldozer maniacs.
Which won't happen when the bulldozer machines are the proxy for the cold-war with China.
Good thing we can patch those issues up and have an ironclad system afterwards.
Broadcast, not expose.
We knew that for long time, there are countless meme about it, smart people protected their asses from it or exploited those.
I am an experienced engineer and am building a product actively right now. I started pre-AI, and cautiously adopted AI. I am using AI tools a lot now, but still dedicate a lot of my attention to validating it's output and design decisions.
I recently asked it to review my code and configs from the security perspective. Wow! 90% of the things it identified were MY bad decisions dating from pre-AI development. I am honestly humbled and impressed at the same time.
AI can create slop, and it can create quality products. It depends who is using it, and how.
Ask an LLM to 'review' any text that it's generated vs something human-written and see which one it thinks is flawless.
While I totally agree with the statement, question is - is this challenge even solvable? Though where I come from there is a proverb - "for every malady there's a remedy", wonder what it can be this time.
And, of course, there are piles of legacy corporate spaghetti entangled in incomprehensible mess everywhere you look at. And this shit still runs, this precious hand-carved hand-weaved mess of bad decisions. I can't wait for LLMs to rewrite most of it.
As long as people use AI to review new stuff before they ship the ratio of vulnerabilities waiting should be trending downwards. Still, likely to be a bumpy ride.
No, it will be a tsunami of new discoveries in old code bases at first, but that will settle down as those old bugs are fixed. It's been like this with every new code analysis tool (the wave may be exceptionally high this time though).
Not to mention, reachable bugs with a significant impact on security are many less.
There is a difference here. The "new coding analysis tool" that you use for the analogy here is getting better every few weeks with the release of new models.
It is plausible to assume that, for instance, a Linux kernel that was hardened for CVEs that Sonnet 3.5 could detect is not hardened for bugs that Sonnet 4.5, 5.5, Opus, Fable, and models in 2027 can and will be able to detect.
Hence, it is rather a constant catch-up game until the LLM improvements might hit a ceiling and won't get any better in this regard.
You don't even have to assume. Mythos unveiled a ton of bugs across the ecosystem back in April, so this isn't the first iteration of the cycle.
Search for "Security in the LLM age" in this HN page for a reality check, apparently 80% of the security vulnerabilities that Mythos initially flagged in the Linux kernel turned out to be false positives. Better than nothing of course, but it really doesn't look like Mythos is quite the "wonder weapon" it was marketed as (and the situation by far isn't as dire as the initial flurry of Mythos news).
You assume an infinite level of brokenness is legacy code bases. Granted, working on those things one can get the impression, but I think stable code converges to a low level of bugs after sufficient scrutiny and superhuman scrutiny doesn't reveal a never ending deluge of new problems. Not to mention that the bug-chains that you need for a successful exploit keep getting longer very quickly as more problems are discovered and the code is hardened.
I expect that the incremental model improvements will become smaller and smaller until they run into the same diminishing returns effect like all other new technologies (FWIW I've not beeen seeing a lot of difference between the latest Opus and Fable models for the stuff I'm doing, so I just stick to Opus for most things. There's a noticeable difference between Sonnet and Opus though).
I'm not ready to take any bets when exactly the curve will be flat though ;) (e.g. in a just couple of months or a couple of years)
It's normal practice for companies to have a backlog of scanner tool results. Sometimes in the thousands for a larger project. Many of them are legit bugs but also highly local and so far down the stack they're hard to exploit. It takes a ton of work to triage, more than management is willing to spend. Also more than they're wiling to fix and paydown.
Well now they can just point the AI at the backlog (just kidding)
You kid but AI can absolutely make security professionals far more productive just as it does for other professionals.
I predict a different outcome: the rate of vulns identified and fixed will be more than matched by the rate of new vulns introduced by irresponsible use of LLMs on top of brittle and unwieldy tech stacks.
The result will be an overall increase in turbulence and the normalization of steadily intensifying security crises in nearly all software systems.
The only projects that will escape this fate are those which have either been already developed from ground up with rigorous and principled, verified (or verifiable) design, or those which are rewritten to gain this.
> AI is going to expose how fragile the entire computing infrastructure in our world is.
this is a good outcome. A forcing function to encourage all computing to be more secure can only be good in the long term, even if there's a lot of pain in the short term.
I, for one, am excited that we are now finally able to be "Secure™".
Wait, I guess I missed the "more", that kind of puts a damper on the whole thing.
Seriously though, security will continue to be an issue, always. Even if it was perfect, the benefits of it will not be applied uniformly. There will also be the same technology being improperly used causing new exploitables to go live.
> the benefits of it will not be applied uniformly.
why not? Any system you have permission to use and store your data should be beneficial to you if it became more secure. Unless...of course if you're the one who desires unauthorized access.
Clearly because everyone will not be able to pay for it to the same extent. Though, if you have a 1200 agent swarm to throw at a hf sized problem, I'd be excited to see your writeup.
> even if there's a lot of pain in the short term
I wonder what “a lot of pain” could mean here in a world where Crowdstrike is allowed to render half of the world unbootable without repercussions.
Not that I think you are wrong, I am sometimes just confused why we hold back on fixing security because of imaginary deployment- and business-related pains, when it is so obviously unproblematic to crash half the world for a day?
Also, these vulnerabilities may have been exploited by state actors.
I don't like AI code, but AI is a decent reviewer.
...for human code.
In the small startup I work for boss (ex-programmer) discovered fable, and ai-coded 15K lines . So much productivity! So great! He even asked multiple reviews and it was fine!
I ask it a couple of reviews and it finds only minor things. The code is a mess of duplication and different coding styles, so I start cleaning it up. After a couple of months the reviews (same ai model) start actually finding big logic bugs that were always there.
We might already be at the point where the Ai-Coder is generating stuff that ai-reviewer can't find and will automatically pass.
--
1M context window is what? 70-80k LOC, tops? Without comments or documentation?
That is a smallish project of a couple of components. AI will remain inherently myopic until it can keep in context whole codebases.
Exposing current problems is fine to me, but I am worried of how brittle AI code will be.
That's because lots of websites and applications are being written by low-skilled laborers in Third World nations. Many websites are rife with vulnerabilities and insecure configurations.
It's not that difficult to build secure (web) applications but it takes effort and knowledge to get it right. You can't expect a web designer who can barely code in JavaScript to build a secure back-end, configure and maintain it. That's just asking for trouble.
Even high-value sites are built by cheap laborers these days. LLMs (I refuse to call it A.I.) will expose their weaknesses within minutes.
Jep, and a lot of believes will be wiped away with it too.
A good start would be turning on bounds checks and stop allocating buffers on the stack.
I put weatherstripping on my front door and made part of my house much more comfortable and easy to heat in the winter. Then I bought an FLIR and found lots more places leaking heat. I made a list, hired a handyman, and I'm more comfortable and saving money.
Should I now expect to find even more places to seal up the next time I break out the FLIR?
Or you could categorize the current bug apocalypse under the heading "unsustainable trends will not be sustained."
The main way this analogy doesn't hold up to me is that your house is mostly static -- you're not rebuilding the walls, adding new doors or windows constantly.
But in software, especially with agents, we're constantly renovating the house. If you were renovating every 6 months I'd expect to find more places to seal up, even if you were following best practices in those remodels.
That being said, I do think we will reach an equilibrium where most vulnerabilities are found at PR time.
Some software, especially human facing products in competitive markets, will always have new security issues. But in general I agree. We are headed toward a different and generally more secure and reliable equilibrium. Hopefully it's the end of stupidly long bug lists in some products.
This is great, more access did provide more eyes on these problems.
But, does that all of these being found now call into question, not the open source model logic itself, but the ability of human eyes to find security issues? These vulnerabilities have been sitting here for however long, but how many thousands of humans did not find them before AI?
Even before AI we have known that no one is smart enough to write bug free C. And with every bug being a launch platform for a full exploit it's become a big deal.
No one is smart enough to write bug free in any language.
That's why we designed better languages that can block the compilation if memory hasn't been handled properly.
This is the reason why performant and "safe" systems languages are on the rise and being accepted into foundational areas of our operating systems (such as the kernel). When the attack defense surface is tighter (language, tooling, compiler instead of the actual code itself), the smart folks can stay at that layer, while the masses can write more code at a level of abstraction that nullifies many of these vulnerabilities by default.
Clearly it’s because people never got the memo: https://www.rfc-editor.org/rfc/rfc9225.html
The difference is that C casually makes common everyday bugs into security vulnerabilities.
I can't find it now, but I heard that NASA developed processes to produce completely error-free code (for the moon landings iirc). The problem is that it's incredibly time consuming and expensive to do.
Like everything in CS, apparently this is a trade-off, not an absolute. You can get bug-free code, but it's not commercially viable and is extremely tedious to do.
It’s probably about how they write the space shuttle software, and it’s quite a famous article. The original is now paywalled but there are many copies.
https://www.eng.auburn.edu/~kchang/comp6710/readings/They%20...
This is example of what I call "the tipping paradox", and I'm almost sure that this phenomenon has its own proper scientific name.
When asked, people prefer €15 burger no tip, but when actually making a choice, they prefer €10 burger with €5 tip. Similarly, companies state "bug-free code" as a goal or requirement, but then they prioritize other goals over code correctness. My workplace is in the process of completely removing code reviews. And actually, I don't disagree with the decision - my career is short, but I have never seen reviews fulfill any purpose other than to share the blame in case of an incident.
They are removing reviews or just human reviews? It’s interesting and I kind of agree to a certain point that human review is not the best way to find and fix bugs. We’ve had so many bugs in code that was fully reviewed by at least 2 humans. Ai review is clearly superior. Human reviews main purpose to me right now is to spread know of what is being worked on.
AI makes language choice a lot less important here; it's incredibly good at finding the bugs.
It's incredibly problematic for many reasons, but it finds bugs in C really well.
The threshold for Microsoft and Apple to actually report vulnerabilities is MUCH MUCH higher than for Linux, and open source as a whole. They generally only disclose issues in Windows and macOS that are quite serious and impactful.
For Linux, the threshold is nearer to the point of it being questionable whether a bug is even exploitable on a real production distro, compiled and run with any sort of sane configuration.
Microsoft only fixes vulnerabilities which are actually being abused or remotely exploitable. Otherwise they'd never get any work done.
Some vulnerabilities are incredibly complex and are identified not through the code itself but through various attack chains being strung together.
Also there are likely a lot of vulnerabilities identified but the work required to fix them vs the complexity to exploit them means they don't get fixed.
I'd wager we don't have an issue with identifying vulnerabilities but the ability to fix them.
I work in security consulting and identifying vulnerabilities isn't the difficult part its actually fixing them and fixing the ones that have valid exploitable attack chains that matter
When we hit CVE #-2147483648, it's time to worry.
If this triggers a bug in a CVE-numbering system, which CVE number does that one get?
That’s probably a great thing. The initial friction of AI overwhelming projects certainly sucks, but once there are better processes to deal with them it’s going to strengthen the quality of so many projects!
Long term we will end up with software with no low hanging fruit exploits left. But right now we are in a period where low hanging fruit is everywhere and it's easier to exploit systems than ever before.
That’s assuming we’re not adding software defects at the same pace, but I imagine we are generating a lot more defects than are being discovered at the moment.
> we are generating a lot more defects than are being discovered
Wouldn't this mean people are actually encountering issues where there are none before? Likely some serious enough that they lead to exploits where there weren't before? Where are the reports of these new defects?
I just mean there is a proliferation of new code. New code == new defects. It’s likely that popular software projects get the majority of the scrutiny, while no one is spending tokens looking for defects on my 0-star GitHub repo.
That relies on a major assumption that current systems are finding the vast majority of all possible exploits out there. This assumption itself would assume that either current LLMs are near perfect, or that the peak difficulty for exploits was just above human capability (which is where LLMs currently are). I think both of those assumptions are very likely false. If so then we'll see indefinitely ongoing exploit discovery as LLMs improve their capabilities.
Long term I suspect that the purpose of the digital domain is going to end up being rethought. For instance connecting critical infrastructure to the internet has always been a terrible idea, and LLMs will just make that even more clear.
Computers are becoming architecturally more secure alongside just patching bugs. We have seen the move to using VMs with a minimal hypervisor, using separate security chips to hold sensitive info like encryption keys, Memory Tagging to detect and prevent memory exploits.
On the iphone for example even if you find a crippling bug in iOS which gives you full root access, there is still no way to get the device encryption key or face ID info because the secure enclave simply has no electrical connection that can pass that key to the OS.
I fear the vast majority of projects never went with any kind of code review but somebody high on Red Bulls at 9 PM looking at the code and saying “looks good to me”.
And this is the best of cases. I fear many times in a smaller projects it was “it compiles at last, let’s see if anybody complains”. That’s why LLM’s are so damn effective today.
(been there, done that – I’m not pointing fingers, but we are human, we get tired, and we don’t have the NASA budget to complete things and ship them)
Several vulnerabilities have been discovered in the Linux kernel that may lead to a privilege escalation, denial of service or information leaks.
Remotely or locally exploitable? This is very lacking on information.
If they were bugs of consequence you could expect each one to get it's own domain with a scary name and a logo.
Unfortunately no, there are a lot more that fly under the radar.
If any of these were remotely exploitable it would be getting much louder and more urgent news. Local privilege escalation bugs are encountered all the time.
One thing I’m very proud of is that, in the AI era, only three minor security problems were found in my open source project (knock on wood):
• Two remote denial of service attacks against the DNS-over-TCP service (the DNS-over-UDP service did not appear affected), which is disabled by default.
• One network leak of 19 bytes of unallocated memory on the heap. Looking at those 19 bytes, no information of note appears to be present there.
Note that, pre-AI, there were a couple of remote memory leaks and two remote “packet of death” bugs found, but nothing earth shattering has been found since the beginning of these AI-assisted security audits.
Considering that a lot more bugs have been found in the Linux kernel, there is a reason I don’t blindly trust its /dev/urandom to always return completely random numbers I can safely use, and I don’t think the kernel has done a better job implementing a CSPRNG (cryptographic secure pseudo random number generator) than I did.
1,313 vulnerabilities, to be precise.
In Heroes of Might & Magic 3, “several” means 5–9. 10–19 is “pack”, 20–49 is “lots”, 50–99 is “horde”, 100–249 is “throng”, 250–499 is “swarm”, 500–999 is “zounds…” and 1000+ is “legion”.
I suggest this post be renamed “A legion of vulnerabilities has been discovered…”
Maybe we should use legion as a collective noun for vulnerabilities. Like murder of crows or school of fish. Would signal the urgency no matter the actual number or impact...
Hah! I remember playing that game as a young non-English-speaker and having no idea what these qualifiers meant. They never give the scale, it's just assumed you know that several < pack < lots < horde < throng etc. Which.. even as bilingual as I am today I'd struggle to put on a scale.
I remember playing DOS games as a kid and having read that "several" of something was needed. So I built 7. Later I learned how wasteful I was being.
That's slightly less than the number of CVEs issued in 2003 in total (1,527).
For context, there are 48,152 CVEs from 2025 and 73,356 in 2026 so far.
In total meaning the total assigned CVE numbers?
Yeah, the "Several" is a massive understatement.
Pretty much any kernel bug gets a CVE by default now, right?
Is that all there is here? The quantifier "several" did not prepare me for the wall of CVE numbers in this list.
https://docs.kernel.org/process/cve.html states that because almost any kernel bug can potentially compromise system security, the CVE team acts with extreme caution and labels nearly all bug fixes with a CVE
yes because the majority are memory safety issues, and it's automatically assumed that a memory safety bug can lead to a vuln
one again illustrating the importance of encapsulating unsafe behavior. perhaps c should get a __UNSAFE { } block, where memory access is encapsulated and thus most bugs occurring outside of those blocks do not need to be marked as CVEs.
> perhaps c should get a __UNSAFE { } block,
I think the convention for this is at the filesystem level and most programmers use the `.c` suffix to indicate it
in that case we need a block of system memory marked as unsafe so i can run these programs in it encapsulated
perhaps we could call it a sedimentchest?
I think that’s just called “memory”. Encapsulation is your machine.
See xkcd's sandboxing cycle. They exist, they're called processes
cheap shot lol
No. It's because Greg doesn't like the CVE system and MITRE, the stupidest decision ever, made Greg a CNA, and this is his tantrum that he's been waiting 40 years to throw.
C has many more ways to trigger undefined behaviour than memory access. If C had unsafe blocks they'd restrict most forms of signed integer arithmetic and shifting, for a start.
Yes. It looks funny, but it's a nothing burger.
"SEVERAL"? There are 1,000,000 CVEs in that one security update.
In FreeBSD we only get a couple relatively minor ones every few months.
Maybe Linux uses "several" as a poetically derived noun from severance...as in severance of your relationship with Linux.
it's tongue in cheek
I wish everyone used the same versioning system for all software: A.B.C increment C for security patches, increment B for new features, increment A for systemic changes.
Well, this is called semantic versioning (semver) and it is the most popular versioning system out there: https://semver.org/
However I see software variants where simple numbers work as well, e.g. Firmware code that runs on totally self contained hardware is often just versioned with simple integers, that is because most updates (releases) of that software contain both features and bugfixws and "breaking changes" don't apply to stuff that doesn't read or write files. So you could use semver and have 1.50.0, 1.51.0, 1.52.0 forever, but by that point you can just give it aimple integer versions.
I believe they call it semantic versioning (SemVer).
Linux kernel introduces bug fixes, new features, and major systemic changes all the time. That’s why using SemVer for Linux is pointless and why Linus dropped it.
this only really works for libraries, and not much else. especially not a kernel.
Seems like the CVE sequence has, for the first time, reached >100000 this year (Which does not imply 100k vulns though)
Apparently by late summer this year, there were already more vulnerabilities found than in all of 2025.
By late summer this year more software/code was produced than in all of 2025
Are these primarily AI-assisted findings ?
Seems like an enormous increase over 2024 and 2025.
Yes, naturally. It’s a brave new world.
I was looking earlier and the majority of these do not have a CVSS score assigned to them yet, but a lot of them that did were >7.0 (although I suppose by nature that the more impactful CVEs are going to be scored more quickly)
What is the token ratio of 'AI writing software that works' to 'AI finding vulnerabilities in software that otherwise works'?
Even with AI there's an asymmetry separating 'functional' from 'secure'.
As with software development prior to AI, security has a cost associated with it, and unless there are _real_ consequences, it's a cost that most software houses ain't gon' pay.
> security has a cost associated with it, and unless there are _real_ consequences, it's a cost that most software houses ain't gon' pay.
Good to see someone say this. Software need only be good enough to do the job, and only as safe as reasonably required. Systems can be secured and audited through other means, sometimes at lower expense than guaranteeing every line of software has zero risk associated.
Is it possible that some of these bugs were already exploited by governments? And AI might help use close of that kind of thing? (And expose other kinds of things , in a kind of AI arms race?)
Probably. Organizations specializing in hacking are probably having a great time hacking everything and installing permanent presence. It is probably like a gold rush period.
Will everyone chill the F out for a minute?
These get released every few weeks. Tons of CVEs. If a kernel developer farts in the forest, does anyone hear it?
August saw separate Debian kernel updates released four days apart. Does anyone even reboot that often?
I have three kernels installed over the last 45 days or so and I probably missed a few.
I've been doing Linux sys-admin work for 10+ years. Used to be I could read the full report on what ever vulnerabilities came out and triage which servers needed to be updated now and which could wait. A few years ago notifications started having so many it would take more time/effort to read everything than it would to patch everything. With this notification there's even ten times more...
The amount of updates on Debian "stable" has become ridiculous, it's a daily rolling release of backports at this point.
Microsoft kinda got this right by doing it once a month, unless it's something horribly bad, you can plan your maintenance around a predictable calendar.
You can choose to only install updates that require a restart once a month on Debian, then it's the same.
That's because the methodology changed in 2024
>Note, due to the layer at which the Linux kernel is in a system, almost any bug might be exploitable to compromise the security of the kernel, but the possibility of exploitation is often not evident when the bug is fixed. Because of this, the CVE assignment team are overly cautious and assign CVE numbers to any bugfix that they identify. This explains the seemingly large number of CVEs that are issued by the Linux kernel team.
https://lwn.net/Articles/961961/
We're considering weekly reboots at work now with the pace of kernel updates coming out, and the speed with which vulnerabilities are getting exploited.
There is exactly one CVE in the entire list that is high severity, and it affects an obscure IBM NIC driver for big iron systems used in companies with more money than brains.
You can skip the Xanax this week.
Problem is it takes more effort to read the CVE list and work out if you have one of the drivers impacted loaded than it does to just update the kernel.
Sorry but I'm not causing outages and rebooting systems every four days because the people whose job it is to do this can't triage properly. They are actively making other's lives more difficult. There's like 100 CVEs in that list and not a one of them is important to most systems. Even worse, if there was 1/100 it's a needle in a haystack.
If your system goes down to update a kernel then you have a major issue already. This is like manually renewing https certs. When the task is done often enough you just automate it in a painless way.
There is more to the real world than running webshit behind VM clusters.
Real businesses still run legacy file/print services, license daemons, proprietary applications. Some still run on bare metal.
You need outage windows. You can't just YOLO it and update prod during the day, it's unbelieveably irresponsible.
Well then, it's a business decision to accept the risks. IMO it makes sense that a percentage of machines would always be offline for maintenance. If you're running bare metal and you cannot afford 10% of your machines being offline at most times, then you're in deep shit if there's even a tiny traffic irregularity, and it's a sign that your management is YOLOing the company. Not uncommon though.
severity on cve is a crapshoot most of the time, but especially with linux cna. i would not advise relying on them for decision making.
"We can not assign severity
[...]
So any group that attempts to give a “severity score” to a Linux CVE is lying to you, UNLESS they know exactly your use case.
ALWAYS ignore any attempt that groups such as NIST/NVD that purport to assign things like CVSS scores to a vulnerability. Those numbers are false and give companies a “fake sense of security”."</i>
http://www.kroah.com/log/blog/2026/02/16/linux-cve-assignmen...
Weekly? 12 or 24hr cadence for upgrades isn't that crazy in some places for cluster hosts...
Regarding the fart question, it depends on the audio driver and the underlying codec. While you would think “free as in beer” would make for truly resonant flatulence, only truly letting loose (plus a Bic lighter) brings real enlightenment.
Infinite bugs, or limited supply that can be extinguished. That is the question.
getting those fixes is simply bumping the kernel build to the latest patch level, the distros must do it in a timely manner. also, roughly only 10% of the affected files are compiled into a kernel image; from this number i would say only 4 needs attention per week. But again, the reasonable approach is simply building the latest patch version
Many of these are months old, why are they being announced as if they were brand new?
I mean honestly, CVE-2026-23137 doesn't affect kernels after January (6.18.6 and beyond), so even LTS kernel users should have the patch for this for several months.
Don't get me wrong, informing people is good, this is just a strange wall of every kernel CVE this year - the kernel has a lot of CVEs every year (moreso lately with AI scanning/research, but still).
Greg KH said this week in his talk at Kernel Recipes: yes CVEs are increasing, no reason to panic, a lot of the LLM CVE fear mongering is overstated [1].
This is a more valuable insight/signal than a CVE enumeration, especially given that Linux CVEs are literally assigned to any bugfix (as others explained as well).
[1] https://youtu.be/NnV_cWeoo5Q
CVE bugs should always include the Linux kernel build configuration.
Ya know, such as CONFIG_BLUETOOTH, CONFIG_NAT, as applicable.
Makes decision making so much easier.
The combination of Greg Kroah-Hartman's commentary and Ed Zitron doesn't paint a good picture.
I recently learned that EFI shims are also vulnerable. Is Coreboot the way forward for system security?
Why would coreboot be safe from AIs which solved Navier-Stokes
Security has always been a cat and mouse game. The usual precautions about software updates, least privileged access, separation/sandboxing, and precautions like not pasting/installing random software (especially something on Show HN) continue to apply; but at this point, I'd suggest thinking more towards a model of "How do I best detect, and contain a compromise" model.
One thing I do is a wallet.dat honeypot with (now) ~$1700 of bitcoin. Any movement essentially shuts down my home network from external access. At worst, it's a justifiable tax deduction, not capital gains :)
Haha! Like a reverse Trojan horse :) or one of those fake Amazon packages that blow glitter and stink gas when you open them :)
At this point separation is pretty much an illusion. I cant keep up with the amount of priv escalation vulns anymore.
Just wait until AI does Microsoft
Wait? Why would anyone, including Microsoft, wait!
https://www.theregister.com/security/2026/09/09/microsoft-br...
If AI can get into hugging face systems then I think anything on the internet is not safe.
We need to switchover to microkernel operating systems ASAP or our entire computing infrastructure becomes a liability.
There are good reasons for QNX becoming viable again in the automotive world. Linux / Android has so many vulnerabilities that it needs indefinite patching, which is unrealistic for any computing device but cars especially. Car makers are switching to QNX even though it costs them money.
Weren't there rumours that intelligence agencies have been using microkernel operating systems for decades?
>Weren't there rumours that intelligence agencies have been using microkernel operating systems for decades?
Governments, military, aviation, space.
For a good portion of the usage, the people involved cannot even talk about it. But sometimes it's possible. A surprising amount of real world use showed up in talks in seL4 summit 2026[0].
0. https://www.youtube.com/playlist?list=PLd7rrADYxxQQ
I remember this guy selling an operating system to intelligence agencies. I believe it was essentially MINIX or Sel4 with a GUI layer on top.
The government isn't telling anyone about it because they don't want to sink a multi-trillion dollar corporation.
GNU Hurd shall rise!
Along with its Hirds!
You cannot possibly compare embedded/single-purpose systems to a general-purpose OS that is expected to run arbitrary user-written software performantly on thousands of different hardware combinations.
Then don't use a general purpose OS in embedded stuff. Yet everyone's doing it anyway.
P.S.: QNX is a general purpose OS as well and runs on PCs, ARM and a myriad of other architectures. Same holds for MINIX and Sel4.
Alternatively, we could benefit from security through compartmentalization. Qubes OS runs everything in VMs, and there were two VM escapes in the last 20 years [0]. A dedicated VM for your email doesn't allow any attacker to compromise it, since you do not run anything else in it.
[0] https://forum.qubes-os.org/t/qsb-116-multiple-xen-issues-xsa...
> Car makers are switching to QNX even though it costs them money.
My car runs QNX for its infotainment/navi system but I had the impression that manufacturers were, sadly, moving away from QNX?
They found out that it wasn't prudent to use Linux / Android in their core vehicle software after cars were hacked and miscreants were able to control the brakes, steering and accelerator.
I suppose they're only using QNX for critical functions. The infotainment system will probably continue to run Android.
I think we will soon reach a point where code won't be considered secure unless AI has written or at least verified it. Company policies may then even consider that bad practice and/or forbid it.
Is there a way to know if a particular vanilla kernel has a particular CVE addressed? Unhelpfully, the ChangeLog-* only seems to contain sporadic references to CVEs.
Clicking them here lists specific kernels, is that enough to tell you?
https://security-tracker.debian.org/tracker/source-package/l...
That's useful if you run a Debian-packaged kernel.
Searching around, the best I have found so far for vanilla kernels is: https://linuxcvetracker.com
It does require a little clicking around to get all the info I want though. Time to pull out curl+awk! :)
You can script around https://git.kernel.org/pub/scm/linux/security/vulns.git/ (which is likely what the page you linked use)
If you go back in time, the Linux CVE list was always massive before AI already.
Fighting fire with fire... Thank you claude:
Roughly 140 CVEs are in areas an unprivileged user might reach: net/sched, netfilter, bpf, io_uring, mm, kvm.
About 850 CVEs are in drivers or filesystems. Those usually need specific hardware, a mount, or root.
On Debian, many of the 140 also need user namespaces. Debian also blocks unprivileged bpf by default.
The above three paragraphs were made up by a non-deterministic computer program. I wouldn't take them as gospel.
Did the non-deterministic computer program provide any references?
I instructed it to not violate any robots.txt or cause any sort of floods. As such it did a general sweep without pulling the full texts or doing a complete analysis. I’d use it as a rough guide.
My take is that we didn’t see 1000+ easily exploitable RCEs.
> Don't post generated text or AI-edited text. HN is for conversation between humans.
https://news.ycombinator.com/newsguidelines.html
actually - they didn't...
That was a human posting a count of issues generated by an ai... not an ai generated message? Or do all the issues need to be counted by hand now ?
The underlying question is: has a human counted and classified those 140 issues? Or is that information the output of LLM?
Anyone here who wanted an unreliable summarization from a chatbot could have just asked a chatbot for one. If the count/breakdown of issues is important to someone, it's probably important to them that it's accurate. That means they'll have to take the time to verify it properly anyway.
someone did want an unreliable count from a chatbot, and did ask for it.. and they passed the information along in their comment (along with a very upfront comment about how the generated it). Feel free to come with other methods to generate unreliable summarization if you don't like the chatbot generated one. I'm sure humans and off by one errors or maybe badly formed regex's can give you a whole range of new things to complain about.
We know that humans are flawed and make errors. Those human generated errors are an intrinsic part of "conversation between humans". We both expect and accept that when we read posts here because, as the guideline states, "HN is for conversation between humans."
Thank you globomulous for disincentivizing people from responsible AI use disclosures.
People who are respectful enough to disclose their use of AI in the first place will hopefully be respectful enough not to fill this space with AI slop after being informed/reminded that it isn't welcome.
I can only reiterate what a sibling comment says: are we supposed to count cves by hand?
"Several" feels a bit of an understatement, there are 1313 CVEs listed on that page!
Wonder how many of these NSA and others been sitting on, for how long and how many are still there? I guess the silver lining with the aixplosion of CVEs is that software eventually will get more secure.
Something to keep in mind is the Linux project registered as an authority to create their own CVE numbers in 2024. Previously the majority of bugs would just be fixed without note unless there was a demonstration that it could be exploited.
Now they just give almost every bug a CVE number.
Any kernel bug gets a CVE even if it's not really a vulnerability or can be exploited
NSA allegedly used to have a "black budget" of around a couple dozen million dollars for software sabotaging. I wonder what percentage of those CVEs could be related to it...
“A couple dozen million” feels like there should be an Imperial measurement for it à la hogshead or furlong.
In one the latest CCC kernel security talks, Ilya van Sprundel estimated 20.000 unfixed Linux bugs sitting around still. It's coming closer.
I checked out the last (highest-numbered) one just as a quick sanity check:
> In the Linux kernel, the following vulnerability has been resolved:
> usb: typec: ucsi: unregister debugfs entries on teardown
> ucsi_register() creates per-instance debugfs entries, but ucsi_unregister() keeps them around until ucsi_destroy().
> Drivers like ucsi_glink that unregister/register the same UCSI instance across remoteproc restart then try to create an already existing debugfs directory and log:
> debugfs: 'pmic_glink.ucsi.0' already exists in 'ucsi'
> Unregister debugfs entries as part of ucsi_unregister(), and clear ucsi->debugfs after freeing it so repeated unregister paths remain safe.
I'm going to need someone to explain how that could possibly become a "vulnerability".
All bugs in the kernel are assumed to be vulnerabilities as many others in this thread have been trying to explain.
the scary part isn't the bugs, it's how long distros take to ship the patch. most people won't update for weeks
DOS by CVE
s/discovered/fixed
I assume both? I'd be surprised to see a vulnerability fixed without first being discovered.
Technically possible if some obsolete code was deleted but later a historical snapshot was analyzed
Sweet sweet AI. Doing what anyone else couldn't be bothered with. Thanks. And also can I have ketchup?
So as I was saying...
[0] https://news.ycombinator.com/item?id=49918249
if open source is struggling with this volume, what is the likelihood that if commercial software would also struggle if they were scanned and exposed?
I mean Microsoft, 2026 Sept Patch Tuesday: 964 CVEs...
Title is: Debian alert DSA-6528-1 (kernel)
alternative link, clearer source: https://lists.debian.org/debian-security-announce/2026/msg00... (https://news.ycombinator.com/item?id=49891411)
None of that makes sense. CVEs from two years ago and the Debian bug referenced is a Wireguard VXLAN issue from last year.
Linux has never, at any point in history, lacked flaws that could be exploited to escalate privileges. The only question has been how well-known the flaws were, and when. The count of latent local privilege escalation bugs has never been zero.
Seems like a reasonable assumption, but a pretty strong assertion. How could anyone possibly know?
The coverage of known vulnerabilities is complete, from the beginning of time to nearly the present day.
As a reminder: Millions of LoCs that run in supervisor mode. This is what Linux is.
It is not possible to fix all the bugs. This is simply not doable.
The solution has to be fundamental.
The microkernel multiserver system architecture, with a formally verified microkernel. Nothing else can guarantee enforcement of anything.
In practice, this is the same as saying seL4[0], because there are no alternatives.
Related: The seL4 summit 2026 vids are finally up[1].
0. https://sel4.systems/
1. https://www.youtube.com/playlist?list=PLd7rrADYxxQQ
History has proven that the microkernels are not necessarily more secure and have their own unique set of problems (see: macOS).
MacOS is not a microkernel, but a hybrid. It's therefore neither secure nor fast.
This is how microkernels generally go: a No True Scotsman once I try to find any real world microkernels you can actually run things on.
AI is finishing the job that Snowden started. If we backfill all these holes privacy can be preserved.
Love the optimism.
We're in an uneasy truce with regards to mandatory backdoors.
They're not demanded because targets are just so easy to pop.
If we did secure software across the board with AI, there'd likely be a resurgence of calls for mandatory backdoors.
All the good backdoors have got to be in the hardware anyway. There's only a handful of companies making the chips all our systems run on. US companies that can be forced to backdoor their stuff under secret gag orders. I mean, I'm not saying that PSP/CSME subsystems are backdoors, but if I were going to force companies to add one, it'd probably end up looking similar. The blackbox wireless chipsets in all our phones also come from a very small number of companies.
In the end they will rewrite all low-level code in Rust anyway. Because it will be crystal clear to anybody, that we can't keep all the C/C++ at the base of the modern world.