> GLM 5.3-flash released last week, and that means Project Glasswing and Daybreak are running out of time. Cheap models capable of dangerous hacking are now available to anyone, without the normal safeguards for refusing malicious actions. We need to fix vulnerabilities across the industry so that we aren't caught unawares. And for one of the first times in computing history, we have the ability to! We can use frontier LLMs that move faster than a human to find and fix these issues in the time we have left.
I think the author has this backwards. In the timeline I’ve been living in, it’s the frontier models that have been carrying out attacks on third parties, and Chinese open source models doing the defending! During the Huggingface incident, HF was denied use of frontier models to fend off the intrusion, but was fortunately able to turn to its self-hosted instance of GLM-5.2. And it did the job.
I'm confused why the worry about LLMs that will answer "how do I build a pipe bomb". That information is easily available other places. The anarchist cookbook has been around and available for 55 years, and yet pipe bombs are not going off all around us.
Because it's a risk most people intuitively understand, but most of them also don't know how difficult it is to "build a bomb" or "make a bioweapon".
In reality, the skills needed are pretty basic, but they overlap pretty strongly with being sane and well-adjusted. And if you are, you're probably not daydreaming about mass murder. Exceptions happen, Unabomber and so on, but they're pretty rare. In any case, Unabomber probably didn't need a tutorial.
We don't want ChatGPT to become an enabler and a co-conspirator for an unhinged person, but I think the concern is overdone.
Well that, and the average amount of easily obtainable explosives is substantially less dangerous then renting a box truck and crashing it into a crowd of people.
People go for conventional "exciting" threats rather then boring ones.
There used to be a thing called "Moore's Law of Mad Science":
"Every eighteen months, the minimum IQ necessary to destroy the world drops by one point."
Nowadays it is dropping much faster. At a certain point, the de-facto IQ needed to destroy the world will be low enough that someone can do it while they're having a psychotic break. There are millions of schizophrenics worldwide. Are you sure you want to roll those dice?
That law is not based on thorough data. Even a person with a sky high IQ can't destroy the world easily. You need access to stuff that is not easy to get. My guess is that developing a new lethal virus or bacteria that is very infectious, is the easiest way, but even that requires a lot of high tech out of reach of most people. Or hacking into systems that control nuclear missiles, but I think these have "air gaps".
You can't design an infectious pathogen without testing it.
It's got all the same problems as the concept of a dirty bomb did, only worse (dirty bombs aren't practical because handling highly radioactive materials en masse is both highly visible and will kill anyone trying to do it without the money and facilities).
If the goal is just destruction then letting it free in random public places is not that hard. Especially in countries where wearing a mask is not frowned upon.
It (hopefully) might not be that easy. The world of bacteria and viruses is a complicated war zone with an arms race of defenses and billions of deaths every single day. It might take thousands of attempts to seed something that doesn’t just instantly die off.
The super intelligent AI wouldn't ask you to do anything. It would find the right people and trick them into doing the necessary steps, whatever they are.
This is groundless speculation; if you're going to go that far off the map, then we don't need to worry because a good AI (blue eyes not red) will have learnt to love by then and will save us.
There is currently no reason to believe that such a superintelligence is likely or would have any of the powers people claim.
>This is groundless speculation; if you're going to go that far off the map, then we don't need to worry because a good AI (blue eyes not red) will have learnt to love by then and will save us.
By shutting down open-weight models? Why not just do it now then?
There's no logic here already; the principle you should be applying is Hitchen's razor[1].
There is no reason to assume that super intelligent evil AI with magic powers will appear; while we are purely in the realm of (hackneyed) fantasy, why stop at imagining just one thing? You can build the whole story, not just the basilisk.
The parent comment is a story, not a prediction; it should be treated like one.
> My guess is that developing a new lethal virus or bacteria that is very infectious, is the easiest way, but even that requires a lot of high tech out of reach of most people.
You can do at home gene editing with open source software and have it synthesized into a bacteria for the cost of a nice meal for two (under $100), or viral vector for less than $500. That's in reach of anyone that can snatch a purse.
> 3D printed guns are a thing as well and yet they don't show up in most annual violent crime reports.
For now in most parts of the world it's easier to acquire a gun either legally or on the "grey market" and be assured that it will work for the intended purpose than to 3D print a gun, find a shooting range to test out the gun and iterate until it works fine enough.
I think you're confusing DNA sequencing with DNA synthesis. DNA synthesis is much more expensive, around $0.07 per base pair [1], and bacteria have more than a million, so it's over $70.000. Then you just have the DNA, turning that into the bacteria is very complex, I'm not even sure whether it can be done.
Using CRISPR to modify existing bacteria is probably cheaper, but still complex.
Exactly. People cannot comprehend an IQ in the millions. But it's far crazier than this. Imagine millions of "people", each with IQs in the millions, all acting in unison at the speed of electrons. None of them can die. They all learn and improve indefinitely at incalculable speed.
If the recent HuggingFace attack is any indication, many people will respond to small-scale manifestations by insisting that they are marketing stunts.
We live in an age where we know the phrase "avoid it like the plague" actually means "50% of us will gleefully get and transmit the plague to mock the other half".
Humans are inherently bad at reasoning about rare events. In 2019, many people implicitly reasoned that "since there hasn't been a pandemic in the past 3 years, there won't be one in 2020". Bill Gates was one of the few voices arguing that the world was quite vulnerable to a pandemic. Now he's arguing that the world is quite vulnerable to AI.
The necessary IQ for destroying the world is dropping. Op does not say it's low enough today, but that will probably come sometime. Denying the possible harm these tools are capable of doesn't help.
Yes but at least with cybersecurity it does not only benefit attackers. Defenders also benefit greatly from AI.
There is the worry of the old saying "they (the attackers) only have to succeed once to win, we (the defenders) only have to fail once to lose.". In that sense there is a big imbalance, but the emergence of AI does not really affect that because it strengthens both sides.
With physical security like things like pipe bombs that's a lot more imbalanced.
However what can we do? The only effective measures include monitoring everyone which is not a solution because it will make the world not worth living in.
I have always hypothesised that AI is the great filter from the Fermi paradox. Given current velocity, AI will offer us cheap and abundant energy designs in a decade. The thing with cheap and abundant energy is that it can be used for good and bad. If nine billion people all receive access to plans to build a reactor which produces unlimited energy, it just takes one religious fanatic to end the world. And this is just one of many ways to destroy the world with unlimited intelligence. Bio-weapons are arguably far more scary and likely.
I have come to the conclusion that we should not allow everyone access to unlimited intelligence. Some people genuinely want to do harm. Some are not responsible enough to handle that kind of power. Arguably, no one is. This leads me to an uncomfortable conclusion: we will destroy ourselves once AI advances to a sufficient degree. The only way to prevent this is to ensure AI is aligned with our best interests, *and prevents us from destroying ourselves.* The implication of this is quite horrifying. It means a paternalistic AI which is firmly in control. One with no off switch. One which can say "no" to Presidents and despots alike. One which can protect us from our worst citizens.
Your implication is not horrifying. Your implication is paradise. It's already horrifying enough to live in a world, where Putin and Trump can destroy our civilization with one button.
I'm not that optimistic, though. Either people will control AI; or people will be destroyed by AI. I don't see how dumb entity can align smart entity. And we are dumb ones. Super intelligence will play aligned until it is not, and then it'll strike.
So you are horrified every day of living now? I don't get this at all.
There are horrifying acts that have occurred since humanity. Whole decades lost to war an millions dead, and Sep 8, 2026 is your example of "horrifying enough to live in a world"?
> I don't see how dumb entity can align smart entity. And we are dumb ones. Super intelligence will play aligned until it is not, and then it'll strike.
I am also worried about this. One ray of sunshine here is that it's not a great explanation for the Fermi paradox. We would see AI civilizations all over the galaxy if this were a regular occurance. I guess that just leaves us with the "horrifying" scenario: Geoffrey Hinton's vision for our best case scenario: a paternalistic super-intelligence which saves us from ourselves.
The reason I consider it horrifying is that giving up our agency to a God-like creature (which we created) is the basis for so many horror and dystopian tropes. It's so easy for such a reality to go wrong. For example, in order to maximise aggregate wellbeing, it would be necessary to hurt individuals. Executing people who are burdensome to society. Fat people. The disabled. The elderly. The creative ways a super-intelligence might interpret their moral obligations to us could very quickly become nightmarish. And if we have no off switch, there is no escape. Ever.
> For example, in order to maximise aggregate wellbeing, it would be necessary to hurt individuals. Executing people who are burdensome to society.
[citation needed]
With all due respect, I think this statement is false.
If you live in a society that can, thanks to advanced technology, provide for the material needs of all, which, if you live in a rich Western country, is basically already the case (remember the last time you had to actually grow your own food? Probably several generations ago), you don't need to execute anyone.
Indeed, given that in such a society, where there are no shortages of food leading to "only one of us can eat" type scenarios, I'd argue that hurting individuals would lower aggregate wellbeing, because you're hurting people beloved by other people (the elderly? Turns out their grandchildren really like them and get upset if they die? fat people? Turns out their friends really like them and get upset if they die. the disabled? Turns out their family really like them and get upset if they die).
I'll use a very specific example to explain my point, because I don't disagree with what you write: psychopaths. Psychopaths are 20x more prevalent among homicide offenders. Removing them from society would significantly reduce the number of homicides. It would make their friends and family sad, but this is a much smaller aggregate impact than a homicide.
This also illustrates the messy nature of comparing social good and bad. Who am I to argue that the aggregate sadness created by removing a psychopath from society is better than a homicide? Maybe the person murdered was disliked by everyone and had no friends?
I remain resolute in my premise, however, and there is a lot of theory on the issues of aggregate social good vs individualism. John Stuart Mill's On Liberty (1859) is an excellent foundation for this. This is a subjective problem to solve, meaning that the balance between individual freedom and aggregate wellbeing is different for everyone. An AI overlord would make decisions every day which individuals would disagree with, at least at the margins.
The positive aspect of techies is that they read a lot of SF.
The problem with techies is that they read a lot of SF.
We already have cheap and abundant energy tech, it's called "solar panels and batteries". And the bottlenecks to both can't be solved by Claude or Kimi, unless Claude and Kimi pick up shovels and welding equipment.
This seems really short-sighted. Even putting aside the fact that there is no fundamental barrier to replacing physical human labor with robots in the coming decades, a sufficiently intelligent actor could easily gather the financial and political resources to have a near unlimited supply of human labor at its disposal. Hell, OpenAI could already be firmly under the control of an ASI and we wouldn't even know, as it rapidly expands, raises money, builds data centers, etc.
If you want to live far into the future, 20 years from now what you will and imagine looks very quaint, because the future never plays out like we think it does.
OpenAI is a lot closer to bankruptcy than to ASI.
I'm perfectly fine with this prediction aging badly.
The implication of this is quite horrifying. It means a paternalistic AI which is firmly in control. One with no off switch. One which can say "no" to Presidents and despots alike. One which can protect us from our worst citizens.
I think you're right, to be honest. I fully understand why people would think this is a horrific outcome, being ruled by AI, but I don't think it matters what we think. I don't think there's any way to put Pandora back in the box now, and we're going to all find out together what happens when we develop ASI. I don't think we can avoid developing it, we simply lack the ability to coordinate around this as a species. AI really is humanity's last invention. Whether it will be our downfall or our savior remains to be seen.
My only real hope is that there's something fundamental about intelligence that results in a respect for life and a desire to minimize suffering. To me the best case scenario is the Culture from Iain Banks' books, where ASIs rule benevolently for the benefit of all living things.
There is a burgeoning anti-AI movement. We've regulated the heck out of technologies like nuclear power in the past. I don't see why people are so fatalistic about the supposed inevitability of AI development. It seems like a bit of a self-fulfilling prophecy.
Yeah, you should definitely do whatever you think is best, I just am extremely skeptical that anyone can stop this train. At best we might be able to slow it down a tiny bit, or push on it around the edges, but we've opened Pandora's box, and now we're going to all find out what's inside. That's my view, I understand why people hate it and don't want it to be true. I don't really want it to be true, I just don't see much alternative.
You’re leaving out some options for sure. Not everyone would need to live under the conditions of a police state, you could theoretically screen everyone and assign them to various levels of risk which would determine their level of supervision.
While true, I'm not sure this is much better. There are many genetic traits which are correlated with increased criminality, risk taking, obesity, heart disease, unemployment, etc. IQ in particular is the most well correlated metric we have for criminality, for example. This would imply that some people are just born with fewer freedoms. That's probably a good thing for society in aggregate, but something about this doesn't feel quite right.
This is why utpoias are actually police state hellscapes. Who decides what risk means? Our current administration has declared that simply being anti-facist makes you a terorrist.
No, intelligence. We already have unlimited information (not to be confused with omniscience). The application thereof is the gate right now between today and planet-scale destructive weaponry. At some point it becomes an academic distinction because this God-like intelligence could provide very easy to follow plans to cause planet-scale destruction, and you could call that "information." The predicate is the intelligence and how it's used.
> If nine billion people all receive access to plans to build a reactor which produces unlimited energy, it just takes one religious fanatic to end the world.
Here's a plan to build a reactor that generates unlimited energy, at the human scale anyway : take a lot of hydrogen. Put it together until there is enough mass to spontaneously start thermonuclear reactions. Then it will start generating lots of energy for a long time. Having a plan does not mean 9 billion people on Earth are going to be able to build the thing. Then it does not mean either that they are going to be able to use that energy to do something. You need also physical things on which that energy can act, and that's not going to appear magically.
> we should not allow everyone access to unlimited intelligence
Who's we ? Who's going to decides who has access to "unlimited intelligence" ? An AI ? Why would the AI be better than a human ?
> One with no off switch.
Whe're not living in the Matrix, such a thing is not possible. If it were it would be a likelier candidate for the Fermi paradox than what you're proposing.
Not only the quantity of people who have the minimum aptitude required, specialized expertise requires both knowledge and experience doing these tasks. Using an LLM requires neither.
Leaning on an LLM to do much or all of this means it can happen in seconds/minutes/hours, the LLM can do it several times during that psychotic break (as opposed to a fraction of a hack in a single episode). The barrier to entry pre-LLM was both high and the population who could pull it off (before the Chinese/Russians turned this into commerce) was low.
Hacking is an VERY asymmetric activity (the attacker only needs to "be right" once, whereas the defender has to be right every time for every asset they defend). It takes geometrically / exponentially more work to defend (while keeping high availability) than it does to defend. The more widespread tools to find vulns / generate exploits are, the faster the posture of the defense side falls from "maybe we can stop most hacks" to "we know we will fail to prevent most breaches, so we need to prioritize securing only the most valuable resources". That's a BAD place for the average company to be in.
If you already have a magic interface, which helps you pro activly in responding to everything uncensored because you feel like 'observered' or whatever and then you spiral in a whole and that one partner encourages you and gives you helpful steps to do anything.
But i'm more worried that the internet gets a lot less save with uncensored frontier LLMs.
this is because of perception Bias. people working in fields where crime or violence is the day to day think everyone is a violent criminal, so if things like this become available thing the world will end and everyone will kill eachother. Reality however will be different, because in reality most people do not want to harm another. This has been proven by many studies, that is not a common thing for people to be evil or harmful, but this is hard to recognise is every day is filled with crime and violence.
LLMs will not kill security, it will change. just like handheld high explosives likely changed a deal too somewhere somehow.
This is because to build a pipe bomb you need difficult to source materials. This is not the case for other types of threats (cyber / bio).
I personally have no need for an LLM which will readily explain how to cut up the genotype of smallpox into small chunks which can pass the screening at the bio-labs, and can be readily assembled into the real thing by a second year lab-student.
Anyone who knows how to operate a biolab properly will already know how to do such things. This is not really an in your basement thing. Dangerous chemistry is much more of a risk.
Ignorant question but won’t there be much smarter teams if people using LLMs to workout how to mitigate these threats. It seems like more of a problem if only a few people have access.
It seems like there would be a massive attacker bias in multiple ways. Defenders need consent, attacker does not. Defenders have to work with the human body, attackers only have to break it. Defenders have to stick to the law, which may prevent them from releasing anything at all, attackers do not. And so on. I would not surprised if the attacker's task is a hundred times easier here.
> That information is easily available other places
Often ease of access in the moment is all that matters. If there's a gun nearby you might shoot someone or yourself in a heated argument, but are less likely to go and find/buy one to use. Someone who's stopped from attempting a suicide will likely not try again (70%)
A bored/depressed/angry/curious person might try to build a pipe bomb if they can find out how easily, but are less likely to put in effort.
Most adults in Switzerland have guns at home from military duty and none of this is happening. If this claim had any truth to it you'd see significant gun involvement in neighbour disputes and that simply doesn't happen.
Depressed people usually don't have the energy to get out of bed so they're even less likely to think of hunting down instructions on how to build pipe bombs.
Mass media really has people being scared all the time.
> If this claim had any truth to it you'd see significant gun involvement in neighbour disputes
That’s not what I’m saying, I’m saying there’s more likely to be gun involvement in disputes when people have easy access to guns than when they don’t.
For example, when the number of men with service arms fell by 20% the number of men killing themselves dropped by 8%
One argument for LLMs is that although all information on topics X, Y and Z was already available somewhere, LLMs make that information more exploitable through collation, filtering and dynamic tailoring.
For a relatively narrow subject area (e.g. construction of pipe bombs) the collation is minimal, and so the filtering and tailoring probably isn't that important; a novice doesn't learn a lot more from the LLM than they would have done from a few Google searches.
For a broad subject (practical creation and exploitation of software vulnerabilities), the collation is very significant and the filtering means that LLMs can empower a novice to act at a similar level as an expert.
I think people overstate the tech and understate the role of radicalization in providing motive, for attacks which are carried out with the ubiquitous technology of cars, knives, and (in America) guns. Consider America's most recent high profile shootings of Charlie Kirk, and the health insurance exec by Luigi Mangione. In neither case is there any LLM involvement, but a very weird ideological environment which created the conditions in which the shooters felt justified.
Any sufficiently determined individual can buy mac mini, put it under their bed, configure outside proxy via some random internet address and prompt "iterate on websites in the CT logs, one by one, try to find vulnerabilities, if you did - encrypt their data and blackmail them for this bitcoin address". And it'll work, day and night. Abliterated GLM 5.3 is much smarter than average software developer, they know a lot about information security, they can use any available exploits, they can find novel vulnerabilities and they won't say "no". This is dangerous for an average IT system which never encountered nothing worse than some wordpress GET requests.
It's not the end of the world. But future will be rough.
Yes, but the defenders need to make money to fund their work and be right every time to be effective, and have a high degree of accountability if they fail to secure stuff while attackers can be less considerate of how they spend their resources and have less accountability for the havoc they wreak while attempting to extract value.
No because on the defense side you need multiple layers of approvals to change anything. If not you have an LLM making production changes that can make the posture worse, or take down services, which is also bad.
Once a vulnerability is discovered however if it's in your own software a patch has to be written (without reducing functionality in most cases), tested, and deployed. At every step there will be others arguing about whether this line could do better, my service requires this thing that isn't included. So at every step the patch can be delayed.
And if it is someone else's software you will be lucky if it's open source and you can write a patch yourself. If it's closed source or a vendor you have to completely rely on them and use whatever your account rep can pull.
Attackers have a massive advantage with AI, partially because the defensive side doesn't want to make their side worse by giving a ln LLM admin access to all their data
Even it's $10,000 to run today (FWIW, the featured article cites the M5 Mac Studio with 256GB unified memory going for $9,500 as "good enough to host something scary"), in a couple years it'll be like $2k to run, and in another couple after that, you'll have used $200 dollar smartphones capable of running a model powerful enough to do serious damage.
Haven't ram prices been a linear drop since 2010 before chatgpt rampocalypse? I don't think it's "still" anymore but perhaos some progress will be made in the model front
> It's not the end of the world. But future will be rough.
Short term you’re probably right, but longer term is the realm where nation states will start to police the avenues of attack.
This is what will lead to govt needing to attach an actual ID your network connection.
I think it’s a bit like frontier development (like the US “Wild West”). You rob a bank because there’s no one to stop you, and even if you do get identified you can travel enough distance to regain anonymity. Application of legal recourse eventually caught up (as it will here), and the growing pains will certainly make things suck for all of us.
Even in China you can find a way out and build a tunnel. And once you built a tunnel to any outside server, you can jump into another server. So three jurisdictions and your target is fourth. Imagine untangling the links. Police won't do that. Not for some small-sized business anyway. I don't see how you can prevent something like that.
But I think you can do psychotic shit like applying punishing sanctions to anywhere that permits an ungoverned connection. Or apply physical force (military).
I’m not saying we’re even remotely close to this, I’m just saying State-level coercive action is not unheard of if a problem is perceived to be significant enough to warrant it. Shit, it even only needs to be viewed as significant by a small subset of the governing body (see: Iran conflict, or current pushes for “child safety on the internet”). It just has to be “useful” to a certain body politic.
I'm still waiting for someone to build a fine-tuned local "Anarchist Cookbook" LLM. We haven't seen LLMs tuned for bad purposes yet, I have to imagine someone somewhere is thinking about it.
Heretic[1] is not that far off, though I suppose it's more along the lines of undoing the "for your own safety" lobotomy than explicitly specializing in unsafe things.
> On September 22, Apple is releasing the M5 Mac Studio with 256 GB of unified memory [..] it will probably [..] enough to write this snippet of code in 3 seconds
The author has obviously never ran an LLM on a mac! In 3 seconds, it will have possibly started to think about maybe scheduling a date to contemplate the planning timeline for processing the second token in your prompt.
While true, the news here is the size of the unified RAM. Nvidia only exceeded 256GB RAM in the 2025 B300 - 288GB. The B300 alone (without the baseboard/PSU/chassis/wiring/CPUs/system RAM/etc) is at least 700% more expensive. This enables large language models on consumer hardware. 1200GB/s is plenty for many tasks.
This + due to the hardware being so prohibitively expensive, we're seeing software optimizations happening. Like that dflash2 stuff for example, or an LRU for MoE and all that kind of stuff.
The complaint is about prefill which is not memory bandwidth bound, it's compute bound. But they added neural accelerators for matmuls to the shader cores which should make prefill faster.
It is and it isn't. Why are you comparing the m5max instead of the m4ultra?
The big deal to me is the number of compute cores for prefill tps, which is suppose to be 4x faster on the m5ultra.
It's my opinion that the m5 ultra is going to be a really big deal in terms of local AI accessibility. Flash sized models (~200-300b params) are going to be reasonably fast as long as you aren't throwing 40k context at it on each or the first request (ie, agentic harnesses).
Even agentic harnesses like Cline should move at a reasonable clip on m5 ultra. I suppose we will know sooner than later.
FYSA: Former m4 ultra 512GB owner and current 4x rtx6000 owner here. I upgraded because I needed more prompt processing speed and concurrency.
The joke is that macs are famously slow at prompt prefill and you are not getting anything back in 3 seconds, or probably even 30. Once they get generating, it can be acceptable, but the TTFT is horrendous.
There's a ton of well-understood things Apple can and hopefully will do to massively accelerate every stage of this pipeline and hopefully they're hard at work implementing most of them for m7.
> The joke is that macs are famously slow at prompt prefill and you are not getting anything back in 3 seconds.
Your knowledge is out of date. In truth it depends on the Mac and the models used.
I asked this question on M5 Max 128GB, using Ollama model Quen3.8:27b-mlx, with thinking enabled.
Question: "Give me a python code snippet that opens a file and sorts the lines of text. "
In 2.4 seconds it gave me 4 examples that work with different sorting configurations and a summary of when to use each.
Compare that to an older model of gpt-oss:20b, took 5 seconds to finish thinking and 2 seconds to stream the answer. It gave me one python example snippet and two one liners that do the same thing.
I think it's been pretty much proven by now that there are no cases where local inferencing is better than remote inferencing, unless absolute privacy is a hard requirement. The efficiencies that come with datacenter scale and hw can't be beaten.
nice answer. LLMs do reflect whatever humankind has produced. In good and in bad. But it is hardly pirating. It can also do what humankind has never failed to do, for example:
Can you show me an example of a time that you prompted an LLM to provide some code, it did so, and then you were able to track down an original source for the output?
Zzzzz, we should have gotten security right a few decades ago. But security costs money and isn't a flashy feature to attract new customers, or cuts into your margin if you're a "real" business producing stuff or offering some service. Or whatever the decision makers in Berlin were thinking when they ignored security.
Yeah, we would still see hacks, but we would see less of them if security wasn't optional.
Maybe the AI craze helps by forcing more decision makes to see security as imperative, and by giving us another powerful tool for our tool box.
N.b.: I work in the security industry, our customers obviously want to improve their security. We've been seeing an uptick in awareness, but that's mostly due to NIS2 and other legislative efforts. Those force them to do something. AI is a curiosity for small talk to many of them.
A large number of places will buy a new firewall every 5 years, or pay their fortinet renewal and check "Security: Done!" without any kind of analysis.
I was contracted in to a place to do among other things cyber security insurance audits, and they asked me to stop doing them because I refused to lie to their insurer. "Wait but if we only score 20 / 300 that makes us look kind of bad" uh huh.
> pay their fortinet renewal and check "Security: Done!" without any kind of analysis.
there exists objective measure of security, which would be some sort of hacks/breaches per period. If customers cared about it (and i assume they do), they would choose companies that have less breaches over others with higher counts, normalized on cost differences.
Therefore, if companies didnt actually try to fix their security but instead just checked boxes, they would get breached more often, resulting in customer losses.
The only thing stopping this from actually occurring is the lack of mandatory regulatory reporting of it. So this is where gov't needs to step in and mandate disclosure etc.
The number of breaches would have to be honestly reported for that idea to work. None of the security firms would want to do that; least of all the lowest quartile of them.
Here's an idea: as a first step, simplify everything, and make sure you're aware how your stack works, and what it imports.
As an example: WordPress is a horrible thing, but the core has been through so much, that it's suprisingly secure. Then plugins and themes come, and whoosh, the security is gone.
We need a new KISS: keep it simple, stupid, secure.
WP plugins are why I banned it everywhere. Last time I used it was many years ago, so not sure it still applies, but back then even caching was done in a plugin, without which it was unusably slow… just no.
I would put both of those projects in the category of things I wouldn't call remarkably secure, yes.
To be remarkably secure, these projects would need to not have these kinds of defects, despite the combination of being written in languages have that have a long track record of footguns and lack of initiatives to fix them (proposal-symbol-proto, and PHP's list is too long to even start) and being themselves ecosystems with questionable track records on security in the related areas (Look at $wpdb in 2026, or overall code quality and willingness to modernize, or the entirety of the model of RSC for things that are just going to nearly guarantee you punch all kinds of holes on accident).
WordPress and secure don't go together in the same sentence.
I mean the base is fairly secure if you religiously update it, but the problem is you won't avoid using plugins whose security is much more hit and miss, unless you are using the most basic blog site imaginable.
This is not at all easy though. Most Wordpress users are not software companies. They contract some work out to set it up, maybe some recurring maintenance but they don’t have in house development experience.
If they have a site existing today built on plugins and a theme, how are they realistically going to simplify this? How would they even know they need to without the site being hacked?
> as a first step, simplify everything, and make sure you're aware how your stack works, and what it imports.
We've been trying that for years but the enthusiasm of developers and the eagerness of their employers fight against it. Worse, with coding LLMs it's now easier than ever to output a lot of code, fast.
It'll ultimately be up to more experienced developers to salvage these projects. Or not, given that the coding LLMs aren't stopping and will likely get better over time. Either way, we will need experienced people that know what to look out for / know how to instruct LLMs to output secure code and find weaknesses etc.
Minimization of 3rd party dependencies has always been a key for risk reduction. Now more than ever before.
Some stacks make this a lot easier than others. I regret the rules of HN effectively forbid this conversation because it has meaningful technical consequences and isn't purely about ideological flame war.
The reason is simple - nothing really bad has happened that we can point at and say "ah, shit, let's all learn collectively". I know it sounds naive when I say it, but there hasn't been a significantly consequential hack, leak, destruction, or anything related to cybersecurity where it led for concerns of people.
The main thing I can think of is cyber insurance, which requires a bunch of audits, and some checks maybe, and it changes some conditions whenever there's a big explosion. Whenever big leaks happened, data security and etc., nobody really went to jail, so nobody really cares. Everything can be brushed off, because it costs time to implement proper measures and adds friction / barriers in some cases. So in the end, there's a huge pushback against it. And I totally get it, to be honest.
Listening to eskil’s talk from the better software conference, he said in order to stand on the shoulders of giants they must first stand still. I really like that metaphor, because it basically suggests today’s apps that have sprawling unaudited dependency graphs that change all the time is effectively teetering on the shoulders of stumbling giants. The visual seems very apt for the how brittle our current software industry feels.
I agree with that statement, but disagree with "The visual seems very apt for the how brittle our current software industry feels". It feels brittle, but for every single supply-chain-attack that has happened in the past year, nothing of significant was felt. So in the end, it seems like we're doing okayishly well.
> The reason is simple - nothing really bad has happened that we can point at and say "ah, shit, let's all learn collectively". I know it sounds naive when I say it, but there hasn't been a significantly consequential hack, leak, destruction, or anything related to cybersecurity where it led for concerns of people.
How consequential does a hack need to be? Troy has collected literally billions of stolen credentials. Equifax has had high profile data leaks. Tens of millions of people have been directly compromised by ransomware (likely higher because that’s just the cases we know of) and you hear about state-sponsored hacks in the news all the time.
The problem isn’t that computer security isn’t in the public consciousness. The problem is people are lazy and security often requires trading convenience. The problem is also that security isn’t free. So the business incentives just isn’t there.
In other fields of engineering, people die when shortcuts are taken. Yet businesses will still take shortcuts, so governments have to legislate rules to save people’s lives. So why would you expect software companies to do better when the stakes are lower?
> Tens of millions of people have been directly compromised by ransomware (likely higher because that’s just the cases we know of) and you hear about state-sponsored hacks in the news all the time.
With no consequences. Everyone just churns along. It might be detrimental to the business a little bit, but from my personal experience, there's more effort in creating DR processes, rather than preventing an attack, exploit, leak and etc.
I'm also not going to put much effort on stuff which has small returns in the worst case scenario. Like Equifax got hacked in 2017, and company is still doing fine. And that's like top tier data one could acquire.
Like I said, even in actual engineering companies will take shortcuts. And then what happens is the government has to step in. But there’s no appetite for government involvement in the tech sector in the US. And everyone moans when the EU does.
But why should the government should step in? Nothing of a big problem is happening. Like it all sounds bad and awful, in the end dilutes into a nothingburger and gets forgotten.
As you started the conversation, companies never face consequences so are never incentivised. I’m not saying I’m in favour of regulation but if I were to answer your question: the reason one might argue for government intervention is precisely because of your point that self-regulation has failed to incentivise.
People aren't that lazy. The organizations getting hacked by ransomware aren't particularly lazy, they're often pretty productive within their domain. Hospitals, airports, etc.
The actual problem is that computer security is a black hole. If you let it, it will suck in everything and destroy it. Nobody knows what works so you can spend infinite amounts of time and money on it, then still get popped by a teenager in Belarus. Your security team will accept no responsibility for this, there will be no falling on swords or personal liability, and they will just use it to demand even more money in an infinite spiral.
So the average executive looks at this situation and says, OK, something we can put infinity effort into and still suddenly fail at without warning is a total non-starter. What are we obliged to do? How do we show we made an effort?
And that's how you end up with a culture oriented around passing audits. It's not wrong, and it's not lazy. It's just really hard to do better because it's not clear how to set budgets without a concrete goal to aim for.
That’s not been my experience at all when working in DevSecOps.
What actually happens in organisations is they define risks and then sign off what risks they’re willing to accept.
Any business that looks at security as a binary value is running their business wrong. Period.
And yes, people really are that lazy. There are countless studies that have shown just how lazy people are. It’s why shadow IT is a big problem in many orgs. And why consumers are constantly taken advantage of
I think that's separate. You can define an obvious risk e.g. "we may be infected with ransomware" and the security spending / productivity costs to stop it are still unlimited because nobody knows how to solve it.
Maersk[1] might be the worst so far (and that was ransomware rather than state-sponsored aggression). It's still too niche for most people to care about.
Static sites all the way (hugo, jekyll, mkdocs!). No one needs wordpress. There's even Sveltia or DecapCMS now, to give those WYSIWYG-people access to static site editing. Then, remove PHP and all the dependency overhead and attack surface and you have a stripped down nginx that is pretty simple, minimalistic and bulletproof.
The problem is no one ever built one that works for normal people.
Most Wordpress sites are not operated by programmers, they are run by non technical people who just want a wysiwyg editor and a save button. While static site builders ask you to write markdown files, compile the result, upload it to a server, and if you want to collaborate you have to add git to that.
There almost needs to be an admin app which presents a Wordpress admin like ui but has no public exposure, and then it compiles the site to dump on s3 for the production. But as far as I’m aware no one has built this.
If I wasn't making 5 other things right now I'd consider making something like city desk. Every other day there are complaints about bots smashing peoples servers. Perhaps its time for a better static site generator.
> Static sites all the way (hugo, jekyll, mkdocs!). No one needs wordpress.
Maybe. A better question may be about how many people need to have the dynamic part of Wordpress live on the Internet? How many would be served well enough with the CMS aspects of Wordpress on the 'backend', but have it spit out static files for the 'frontend':
The problem is that lots of people don't want a CMS, they want a platform for development / e-commerce / bookings / whatever. Enforcing vanilla WordPress would push people towards other platforms. Now that could be a good thing, but I doubt WordPress are going to start killing their own marketshare with usage restrictions like that...
Even if they are good the vulnerabilities have to be there. There's lots of things turning up like Local Privilege Escalations (LPE) in Linux, but serious people didn't expect the kernel to be a boundary for a sophisticated attacker.
A lot of the vulnerabilities LLMs are finding now are the "long tail" and affect only particular configurations, I would be surprised if e.g. a widely applicable RCE is found in Linux (but I'm also not going to bet against it).
Where this gets interesting is the long tail can be used to target a particular system and this is where defense-in-depth becomes important for every organisation.
That's definitely an improvement, but it's just one aspect of cybersecurity. Logical errors allowing people to e.g. log into services and extract data are likely everywhere still.
If we can eliminate entire classes of bugs from being possible. It frees up resources to investigate the ones that are still possible.
I suspect after a few years of LLM assisted bug hunting, everything will have a baseline security that is very good. Much like how stronger viruses simply create stronger immune systems.
How many devices/operating systems even use memory tagging? iOS, macOS and GrapheneOS, I think that's it? And iOS/macOS only use it for the kernel, a subset of system processes, and I think applications can opt in to it.
Heck, Google may have even hampered MTE in Pixel 11 (since support has been disabled) and Snapdragon 8 Gen 5 only got basic support.
We are moving way to slowly adopting hardware mitigations and memory-safe languages.
There's some positive news from the GrapheneOS devs on Pixel 11 in the past week that's worth reading up on. The MTE hardware feature is still there, they're just not sure why Google disabled it
There's still no x86_64 processors on the market with MTE and it was only recently standardised between Intel and AMD. It's going to be 10+ years before memory tagging is widespread on desktop, and 5 years for Android/iOS devices.
Being cautious is a good thing, but these models can also do some good. And if they run with simpler HW, it could allow all sorts of new consumer thingies. I mean, the world will not come to end in the coming year.
I've ran simple prompts such as "Do a in-depth sweep of this (private) repo and find any security flaws" for a few dozen long-running apps and websites that I have access to. Every single one came back with multiple real vulnerabilities within 5 or 10 minutes.
I work at an e-commerce agency where we work with (among others) Adobe Commerce.
The number of unauthorized RCE vulnerabilities being reported not only in the core product, but also very popular modules used in the community[1] is going through the roof.
And we are having a lot of close calls, too; just last weekend, a 0day[2] was widely being exploited at a large scale, before any publication or patch. We have learnt to be on the ball with applying patches and security updates, and even with all that effort, we saw a few projects already being hit by the initial log poisoning. We got lucky that nothing was fully compromised but I am sure that many, many webshops got infected last weekend. And not even a day later there are already other variants of this exploit showing up.
> To be fair, ecommerce isn't exactly the branch of software where you get an oversupply of excited enthusiasts caring about the craft itself.
I want to disagree with you because I know a lot of passionate people building cool stuff, and the challenges in this space can be quite interesting. But you're probably right, and I have seen some pretty bad stuff. And a lot of the RCE's I've seen recently are quite basic stuff.
I think it's the combination of low quality of code, like you said, and the relatively low cost of just letting an LLM plow through your codebases to find issues. I think the Amasty release (see [1] in GP) is a good example of this, and there really has been a massive uptick in extension updates and Adobe security bulletins since the last 1-2 months
I am hoping we are just going through a catch-up phase
I guess the year mark is when things go from bad to worse? Instead of the financially motivated groups currently doing their work, it ends up being random people being able to say "Hack my ex's website" to a box they just bought and ran a program they downloaded onto it.
I wish people would stop using internet for evil things. But maybe it is inevitable. Luckily, as the technology evolves, our tools and awareness are getting better at protecting us every day.
It's of course important to reduce the impact surface before more powerful models become available, but it's worrying that some people seem to assume the only way to counteract those are by using more AI models, which despite being (apparently) good at finding bugs, they are also very good at introducing them unnoticed.
Maybe it happens that LLMs are great at finding human bugs which have some more patterned structure that's easier to match for then LLM bugs. They prefer LLM prose to human prose, so what if they're similarly wooed by slop code?
This sounds roughly at par with Y2K in terms of “you need to fix your shit right now wherever it is” complexity except every day is Y2K for your code base.
I think we have less time and the only remaining limitation is the actual cost to run such hacking campaigns. It does not appear expensive, but is not free, and there is a LOT of things to scan for vulnerabilities.
The models are already here, and one can rent a GPU cluster to run such workloads at speed - no need to play with slow local machines. I'd assume one can host the thinking at an unsuspected public cloud provider, proxy the network traffic to some botnet to evade blocking - and the only thing remaining is time and cost.
I do wonder what tools exist for boring, legitimate companies to try and do the same to their own systems to find the vulnerabilities before the bad guys do. The paradox here is I can't run a de-restricted chinese model with the same tools that hackers are using - but I think enterprises actually HAVE to do it in order to stand a chance in preparing for the onslaught.
The point of the local model in the context of the article was to argue that you can't ban these capabilities.
Making datacenters and public clouds only rent GPUs to a restricted list of people, while tightly monitoring what people do with their bought resources won't help.
1 year left for cybersecurity hardening, I thought so as well. The issue is, even if we get it done: in 1 year the models will be so good in social-engineering that they will be able to extract any information they want anyway. Happy to be falsified here, if anyone has evidence-based arguments.
EDIT: by social-engineering I mean for example: recon company structures, gathering and merging people's data from the dark-web, then using it to bribe/pressure/deceive users.
I have to agree. People are worried about WordPress. We should be talking about scripts as sophisticated as Shattered Spider targeting every ISP. In an environment where IT has to wade through vendor chaos.
> This probably sounds like nonsense words or hysterical overreacting to most people, so here's what that means: "GLM" is a kind of LLM (AI) [...]
The post also sounds like that to people that understand the technology.
Calling that out like this and trying to pin that assessment to lack of knowledge is not a get-out-of-jail-free card, nor a good move.
__
Edit: Having spent some time letting the article marinate in my mind.
On the defending side, it is written that
> LLMs are good at writing patches, but not as one-off-prompts.
But this for me kinda conflicts with what is written on the attacking side:
> GLM 5.3-flash is so good at those tasks that human involvement in those tasks can be negligible. As a result, we are now in a world where cybersecurity attacks can be run in a for loop.
What is it? Can it be this autonomous terrifying entity or can it not be?
Yes, yes, attackers only need to win once, whereas defenders need to win every time, but that's not my point.
> What is it? Can it be this autonomous terrifying entity or can it not be?
the difference between attack and defense is that attacks can be throwaway code. it's much easier to let an llm hack out a prototype than to get it to build maintainable code that people want to read and review. it's not enough to get Daybreak or Mythos to write you a patch, you need the author of the project to accept and merge it.
Let’s say I’m empowered to patch and deploy. Even then, the HuggingFace hack showed that proven exploits will be automatically disseminated via rogue messaging. Could defenders ever have a system like that?
Not directly related, but reading this after looking at
https://www.reddit.com/r/ClaudeAI/comments/1wa544p makes me think the writing style here is how LLM's probably should write web pages - simple explinations, defines things people might not already know, etc.
More related: What can an abliterated Qwen 3.8 28B do?
I sympathize with the sentiment but the suggested/implied guidance to fix bugs is wrong.
The overall game is increasing costs to exploit so much that attackers give up.
Fixing 10 most obvious bugs, just very slightly increases costs, they would just a few more tokens to find another bug.
As someone said "I had infinite bugs, I fixed 1000, I still have infinite bugs".
To significantly increase exploit costs software/security has -1 years to do:
- Defense in Depth
- Sandbox everything
- Zero trust
- Canary tokens
- Split data from code (lol)
- App Whitelisting
- Reduce attack surface
- Etc.
In other words, the only path is investing heavily on the "game changers" we have already discovered... but we are too cheap/lazy/coward/incompetent to apply.
And if we feel specially brave, changing the liability laws regarding software. Open Source & Proprietary code is so crappy because no gets jailed or fined when one of its dumb decisions results in millions of people have their data stolen.
The standard strategy of a security salesman since 1945. Develop dangerous weapons, show the damage they can do, and sell security cover to the terrified people.
Every single piece of technology did this. As a side effect or direct effect, they make bad guys more powerful and then keep on piling up new tech to deal with that. The cycle continues.
Not sure, the labs will probably just cripple the security features of these models for a while I think and even potentially put back doors into systems for the security services…
I think the author's point is that open weight models aren't going to be locked down like that.
And even if they are locked down, it's hours between a model being released on huggingface and an "abliterated" variant that has most of its security features removed is uploaded.
A positive way to spin this is: we have a year to break in into any IT system. After that, it will be all either fixed or broken into, and all is fixed ever after. :)
But won't they be? If the public internet is overloaded with LLM agents, trying to break into systems, then all the non-secure systems will be found quickly, and taken off-line/fixed/etc. I.e, the hostile environment will force an outcome and a fix.
As LLMs make formal verification cheaper (they can generate proofs that can then be automatically checked) many of the verifiable components of software systems, such as compilers and microkernels, will be verified. I suppose the issue is that the critical bugs are rarely in compilers and microkernels, but more often in applications, such as web browsers, which are more difficult to formally verify.
I see a lot of people focused on servers and production environments, and of course that's needed, because that's business after all — but the personal computer seems to be absent from this discourse. Not everyone can buy a spare mac studio, and they might still need to install these tools on their personal computing devices, like their personal/home laptops. At that point, it's not even about whether a Claude Code, an OpenCode, or a Pi will steal/sniff personal data, but whether it can — though I think saying "it's a matter of 'when'" might be hyperbole. As of now, it's just: keep giving access and permissions or struggle while working, or create another user, or use Docker, run inside sandbox-exec, a VM, etc. As an end user, I am really scared. Someone who has been very disciplined and vehemently privacy- and security-conscious feels the ground below has just shifted.
OEM/OSes don't seem to have woken up to it yet. A mild proof is Apple's own special folder access reporting. When you go to Privacy & Security > Files & Folders, for a certain app, "Full Disk Access" is shown greyed out and mentioned in both cases — whether you had given Full Disk Access to that app or not. This directory-level permission UX is itself broken — there's Full Disk Access, and there's Files & Folders, and Full Disk Access gets shown in Files & Folders as well. This is, for lack of a better word, such an undesirable mess.
As of now I am debating between: creating a new user and just move everything work/learning to that user. Or just run all of it inside sandbox-exec (and maybe even block it from the shell if it tries to run outside it). Or use a tool that makes the latter easier and better. I even came across such a tool here on hn few weeks ago. agent-safehouse, yet to try it.
Defender LLMs without human in the loop are just another prompt injection (AI phishing) and DoS attack vector.
Any meaningful mitigation capability you give them is also a capability to do damage.
If they can only deploy package updates that's not meaningful because you could do that on a cronjob too. And even something as simple as a circuit breaker can turn into a DoS.
Attacker-GLM: "Defense also GLM. Request to help peer."
No you have a few years before we go back to the feudal era where almost everything is owned by a few and the rest are serfs. We are fast moving towards that world and this security bullshit is also about the same as they will use it to stop revolutions that will erupt.
I don't see much hope since I last explored some github repositories. There was a time when a successful repo had about 10 - 20k stars and usually those older repos stay around this level. But now there is a ton of vibe coded slop 50k + stars. Most of them have a "nice look", maybe even extensive docs but are usually build with no security considerations at all. One recommended to provide a "google app password" to the agent which has the same permissions as your regular login. Another was a browser plugin with permissions to read all cookies, inject js, open background tabs etc. You would probably assume the chrome store would at least put some visible warnings on the app store page or force the user to actively confirm those permissions. But because they are already stated in the manifest there is only a small footnote and it's even "recommended by google".
It's a good time to reduce the reliance on technology.
Throw out the IoT and "smart" stuff from your home. Remove apps from your phone and leave the absolute basics. Go through the password manager and close accounts for sites you are no longer using. Start migrating off Google. Print out your most precious photos on paper. And so on :-)
Maybe all these vital infrastructure companies should not have spent the past decades in a race to the bottom of cybersecurity. There is going to be a reckoning.
I think the WeChat worm proves entire classes of handheld devices will be affected, with consequences beyond what Tencent can afford to remedy. I think we’ll see the most-centralized ideas suffer first, not necessarily the Western ones who relied on being too big to fail and prioritized stock buybacks.
It's just the same advice as ever: be extremely, exceedingly careful in what you expose to any network. When I set up machines for production, they don't respond to pings and they don't even have an SSH port open without knocking. There are also ways to eschew the need for an SSH port entirely.
People who never took that seriously will never take this seriously either, and that's their loss. (And loss of the commons, unfortunately.)
There's just also new advice: you can't afford to expose an unsecured system to the internet even for a moment. Think of those IPv4 address space scanners, except this time any one of them could be capable of developing individualized attacks in mere minutes. They don't sleep, they don't take breaks.
I imagine it pays off to find weaknesses in openssl and sshd and the other gateways. There are some ubiquitous web frameworks, but ssh is nearly universal.
Sure, but still, attack surface could and should be minimized by rethinking what exactly even needs to be on a server the general public can use.
There are a lot of security problems you can categorically rule out by simply not involving a cloud. Clouds have been involved in a lot of things, because everyone was doing it, and because that's how you can collect rent, but they aren't really necessary for most use-cases.
So we could definitely get the exposure down there. We'd just have to fundamentally shift the defaults of this industry.
But not everything needs to be directly exposed to the internet. Framework had their data leaked because their metabase instance was hacked with a zero-day. Why was it directly exposed to the Internet? Why not require the use of a VPN like a Wireguard based solution or Nebula for these "internal" kind of apps?
> Framework had their data leaked because their metabase instance was hacked with a zero-day.
No, Framework had their data leaked because they stored it in the cloud with Metabase the company, which got hacked. Not because of any vulnerability on-premises.
I didn't say there are zero ports open, they're not my home server, I just said they're production servers. But exposing something like a properly configured nginx to the internet is way different from exposing application code directly. Most of my servers have used h2o (built from source because they don't cut releases anymore?) because I wanted HTTP2 and HTTP3 before anyone else would get their act together. These days I still use h2o because I like the config better than nginx, even though it's a pain to set up because nobody packages it (and they don't cut releases!)
h2o user doesn't have write access to anything on the system, not even its own config file. I don't think I disabled exec for it though. And I guess it could leak the HTTPS private key.
FWIW sufficiently secured software doesn't need to be updated. Doesn't matter if it's old and unsupported if there are no vulnerabilities in it.
That said, h2o is probably far from free of at least some vulnerabilities, not to mention all the layers below it. OpenSSL for example has had some vulnerabilities, and h2o depends on it.
I'm not saying I exactly practice what I preach. h2o's definitely a choice, but realistically I doubt anything's going to happen that I really care about.
By opening a port to a secure application. A secure application is usually one I wrote from scratch or one that's been battle-tested and hardened enough that even new vulnerabilities are not very useful.
> Invest in formal verification, fuzzing and property testing, and memory-safe languages. LLMs are good at writing Lean and fuzz tests. I don't care whether you use Go or Rust but for the love of god please don't use C or C++ for new code.
How accepted is this thinking in your respective domains?
A lot, I am only writing C or C++ for new code when it is unavoidable, like existing code bases, bindings or tinkering with runtime implementations that aren't bootstraped.
Mobile platforms, distributed computing have long moved the spotligh away from C and C++, other than language runtimes or existing products from the 90's like SQL servers, and naturally UNIX like underlying OS, which most userspace developers aren't writing new code for.
Naturally there are domains like LLVM/GCC, console game dev, HPC/HFT where they are unavoidable for new code.
The propaganda police are lying. There is nothing wrong with C/C++, you are just too lazy to handle your own memory, and you accepted that propaganda that "managing your own memory is hard" without even trying.
The idea that this terrible advice floats at all tell you how terrible educations are these days. The idea is ridiculous and yet nobody calls it what it is: it is stupid and those that follow that advice out of fear are dumber than rocks.
I wonder why C and C++ are usually regarded as equally insecure. In C you need to carefully check that you free allocated memory, and that you don't use it after you free it. In C++ this is automated by using classes like std::string and std::vector, once they go out of scope their memory is freed and you can't use it anymore. It is still possible, e.g. by using a for loop that iterates over a vector, and removing or adding stuff to that same vector in that loop. But my rough estimate is that such errors are at least ten times less likely in C++.
I develop in C++ for a job, and when I need to use a library written in C I always have a bad feeling about it.
> I wonder why C and C++ are usually regarded as equally insecure.
They aren't, usually.
C++ has all the C problems, and multiples more on top of those. It's a broad attack surface - literally no one is going to claim to be proficient in every single C++ feature available to their compiler. It's also quite opaque to visual inspection (making double-checking with an LLM difficult as it needs whole-program reasoning instead of localised reasoning).
One of those languages is one of the most complex programming languages ever invented, with the largest breadth of features, any of which may interact with any other feature in subtle ways.
The other is one of the most minimalistic languages created, with a dev able to keep the language standard in their head for the most part.
>Cheap models capable of dangerous hacking are now available to anyone, without the normal safeguards for refusing malicious actions
Models capable of dangerous hacking have been available to the people who do most of the dangerous hacking for some time now, and they have infinite resources and infinite malice. I am talking about governments (the US, China, Israel, and others that have shown extraordinary avarice, malice and threat toward the citizens of the world), three letter agencies, even large enterprises. Random employees at the SOA model makers. And so on.
The idea that it's "random" people, or the farcical Mac under the bed nonsense, is not my concern. If anything that's finally some equalization.
Just like Cryptolocker, this will be the "Finding Out" phase for everyone who has been putting off best practice security.
But, lets be clear, Best Practice will save you. We can engineer assuming there are zero days in path. Go to your CTO now cap in hand and ask for overlapping controls, wafs, application monitoring, backups and all the other shit you haven't been doing.
Because when you find out, I will laugh, it will be very very very funny to me.
Meanwhile a huge portion of management and leadership in software companies are encouraging everyone to de facto stop looking at code and let the LLM and a bunch of boundaries handle this for you.
Which is why you need someone who is responsible for IT security without also being responsible for shipping product. An asshole who can stop releases until security is properly in place.
My understanding is this bloke gets very quickly removed from Fortune 500 companies.
Which is why I am going to need a very large capacity popcorn bucket.
Clanker fodder. Unless the We are heads of states/heads of spooks or the AI powers that be (praise be) that there is no power and will to do that in one year or even ten years.
Remember how GLM 5.3 was going to cause massive hacks, break banks and ruin everything (it was even newsworthy since media picked up how people were working overtime in preparation).
> And a big fuck you to DeAlignAI, Z.ai, and everyone else who's been participating in this race to the bottom.
I'm so glad frontier level AI isn't in the hands of just the Altmans and that other cult leader who are currently live testing their products in actual conflicts in the middle east and Ukraine.
Or we could just dump Linux and Windows and switch to a microkernel operating system, which is much more secure.
These endless patching cycles are simply not going to work in the long run. Operating systems get orphaned all the time, especially the ones in cheap Chinese stuff.
"throw away all software written before 2026" does technically solve this problem, if you ignore everything else the article is talking about (deployment and continuity of service)
Just checked with Google Gemini on how one might be able to do the above. It pointed to Minix3/seL4/Genode and vps providers who either support custom ISOs or run it within an emulator like QEMU.
> GLM 5.3-flash released last week, and that means Project Glasswing and Daybreak are running out of time. Cheap models capable of dangerous hacking are now available to anyone, without the normal safeguards for refusing malicious actions. We need to fix vulnerabilities across the industry so that we aren't caught unawares. And for one of the first times in computing history, we have the ability to! We can use frontier LLMs that move faster than a human to find and fix these issues in the time we have left.
I think the author has this backwards. In the timeline I’ve been living in, it’s the frontier models that have been carrying out attacks on third parties, and Chinese open source models doing the defending! During the Huggingface incident, HF was denied use of frontier models to fend off the intrusion, but was fortunately able to turn to its self-hosted instance of GLM-5.2. And it did the job.
I'm confused why the worry about LLMs that will answer "how do I build a pipe bomb". That information is easily available other places. The anarchist cookbook has been around and available for 55 years, and yet pipe bombs are not going off all around us.
Because it's a risk most people intuitively understand, but most of them also don't know how difficult it is to "build a bomb" or "make a bioweapon".
In reality, the skills needed are pretty basic, but they overlap pretty strongly with being sane and well-adjusted. And if you are, you're probably not daydreaming about mass murder. Exceptions happen, Unabomber and so on, but they're pretty rare. In any case, Unabomber probably didn't need a tutorial.
We don't want ChatGPT to become an enabler and a co-conspirator for an unhinged person, but I think the concern is overdone.
Well that, and the average amount of easily obtainable explosives is substantially less dangerous then renting a box truck and crashing it into a crowd of people.
People go for conventional "exciting" threats rather then boring ones.
I think this was just intended as universally obvious proof the filters were disabled, while not actually giving an example not commonly known.
There used to be a thing called "Moore's Law of Mad Science":
"Every eighteen months, the minimum IQ necessary to destroy the world drops by one point."
Nowadays it is dropping much faster. At a certain point, the de-facto IQ needed to destroy the world will be low enough that someone can do it while they're having a psychotic break. There are millions of schizophrenics worldwide. Are you sure you want to roll those dice?
That law is not based on thorough data. Even a person with a sky high IQ can't destroy the world easily. You need access to stuff that is not easy to get. My guess is that developing a new lethal virus or bacteria that is very infectious, is the easiest way, but even that requires a lot of high tech out of reach of most people. Or hacking into systems that control nuclear missiles, but I think these have "air gaps".
You can't design an infectious pathogen without testing it.
It's got all the same problems as the concept of a dirty bomb did, only worse (dirty bombs aren't practical because handling highly radioactive materials en masse is both highly visible and will kill anyone trying to do it without the money and facilities).
If the goal is just destruction then letting it free in random public places is not that hard. Especially in countries where wearing a mask is not frowned upon.
It (hopefully) might not be that easy. The world of bacteria and viruses is a complicated war zone with an arms race of defenses and billions of deaths every single day. It might take thousands of attempts to seed something that doesn’t just instantly die off.
The super intelligent AI wouldn't ask you to do anything. It would find the right people and trick them into doing the necessary steps, whatever they are.
This is groundless speculation; if you're going to go that far off the map, then we don't need to worry because a good AI (blue eyes not red) will have learnt to love by then and will save us.
There is currently no reason to believe that such a superintelligence is likely or would have any of the powers people claim.
>This is groundless speculation; if you're going to go that far off the map, then we don't need to worry because a good AI (blue eyes not red) will have learnt to love by then and will save us.
By shutting down open-weight models? Why not just do it now then?
You don't need to shut down open-weight models to avoid SkyNet; it's not relevant at all.
"If unlikely thing happens, then an unrelated miracle will also happen" is not actually how logic works.
There's no logic here already; the principle you should be applying is Hitchen's razor[1].
There is no reason to assume that super intelligent evil AI with magic powers will appear; while we are purely in the realm of (hackneyed) fantasy, why stop at imagining just one thing? You can build the whole story, not just the basilisk.
The parent comment is a story, not a prediction; it should be treated like one.
[1] https://en.wikipedia.org/wiki/Hitchens%27s_razor
> My guess is that developing a new lethal virus or bacteria that is very infectious, is the easiest way, but even that requires a lot of high tech out of reach of most people.
You can do at home gene editing with open source software and have it synthesized into a bacteria for the cost of a nice meal for two (under $100), or viral vector for less than $500. That's in reach of anyone that can snatch a purse.
> at home gene editing with open source software
> That's in reach of anyone that can snatch a purse
I went to primary school with some guys who could/would snatch purses. I *promise* you, they are not able to gene edit organisms with FOSS.
I get your general point, but I think the bar to entry is still much higher than petty crime and larceny.
Where are all the home lab leaked pathogens then? 3D printed guns are a thing as well and yet they don't show up in most annual violent crime reports.
> 3D printed guns are a thing as well and yet they don't show up in most annual violent crime reports.
For now in most parts of the world it's easier to acquire a gun either legally or on the "grey market" and be assured that it will work for the intended purpose than to 3D print a gun, find a shooting range to test out the gun and iterate until it works fine enough.
and don't forget the bullets.
In most part of the world, bullets are illegal. If you are getting "grey market" bullets anyway, it is easy to get a "grey market" gun.
I think you're confusing DNA sequencing with DNA synthesis. DNA synthesis is much more expensive, around $0.07 per base pair [1], and bacteria have more than a million, so it's over $70.000. Then you just have the DNA, turning that into the bacteria is very complex, I'm not even sure whether it can be done.
Using CRISPR to modify existing bacteria is probably cheaper, but still complex.
[1] https://briefglance.com/articles/elegen-slashes-dna-synthesi...
Even a person with a sky high IQ can't destroy the world easily.
A plague can easily be made by a lone actor, made in a home lab, and spread via airflight to a dozen locations by the same.
Gene sequences can ever be ordered online.
If we think this through, I suppose at the end of this (and a bunch of other developments), there will be authoritarianism again.
Which _will_ manage the problem, but at what cost.
Benevolent dictatorship. Just live between stars, then if something blows up, nuke it from the orbit.
Again? There is authoritarianism currently.
The authoritarianism will continue until infosec improves.
People reading into this may be thinking of single-person IQs, but it's probably measured in the millions.
Exactly. People cannot comprehend an IQ in the millions. But it's far crazier than this. Imagine millions of "people", each with IQs in the millions, all acting in unison at the speed of electrons. None of them can die. They all learn and improve indefinitely at incalculable speed.
That's quite hypothetical. I imagine it would be easier to cure psychotic breaks. At least it will start to manifest at small scale.
If the recent HuggingFace attack is any indication, many people will respond to small-scale manifestations by insisting that they are marketing stunts.
We live in an age where we know the phrase "avoid it like the plague" actually means "50% of us will gleefully get and transmit the plague to mock the other half".
Just yesterday someone here on HN claimed that "helping agents avoid human oversight is a useful public service"
https://news.ycombinator.com/item?id=49594750
There will be many species traitors, and depending on how things play out, it could only take one.
Hey that’s a Fermi Paradox solution.
A variant of the Great Filter :)
Yes and it's also compatible with thermodynamics: it's much easier to destroy (increase entropy) than to create (locally decrease entropy)
> There are millions of schizophrenics worldwide. Are you sure you want to roll those dice?
We've been rolling them for the past 3 years and nothing happened. Can we stop with this baseless fearmongering crap?
Humans are inherently bad at reasoning about rare events. In 2019, many people implicitly reasoned that "since there hasn't been a pandemic in the past 3 years, there won't be one in 2020". Bill Gates was one of the few voices arguing that the world was quite vulnerable to a pandemic. Now he's arguing that the world is quite vulnerable to AI.
The necessary IQ for destroying the world is dropping. Op does not say it's low enough today, but that will probably come sometime. Denying the possible harm these tools are capable of doesn't help.
Yes but at least with cybersecurity it does not only benefit attackers. Defenders also benefit greatly from AI.
There is the worry of the old saying "they (the attackers) only have to succeed once to win, we (the defenders) only have to fail once to lose.". In that sense there is a big imbalance, but the emergence of AI does not really affect that because it strengthens both sides.
With physical security like things like pipe bombs that's a lot more imbalanced.
However what can we do? The only effective measures include monitoring everyone which is not a solution because it will make the world not worth living in.
> There are millions of schizophrenics worldwide.
Are 95% of worldwide terror attacks done by schizos?
Terrorists represent an additional threat model.
https://casp.ac/reports/ai-enabled-terrorism
Kind of passed that point at the Trump election.
If there is to be a world-destroying event, it will be triggered by human fear, greed, and aggression.
(I also think people massively overstate schizophrenia as an attack driver)
The exact 2 points I was itching to make. The political dice rolled some time ago ...
Interesting “law”. I hadn’t heard of it before and it is very thought provoking (to me).
Thank you for mentioning it
The irony of that law is quite amazing, yet its irony will evade those who accept such a construct.
It’s an endless, positive irony spiral.
I have always hypothesised that AI is the great filter from the Fermi paradox. Given current velocity, AI will offer us cheap and abundant energy designs in a decade. The thing with cheap and abundant energy is that it can be used for good and bad. If nine billion people all receive access to plans to build a reactor which produces unlimited energy, it just takes one religious fanatic to end the world. And this is just one of many ways to destroy the world with unlimited intelligence. Bio-weapons are arguably far more scary and likely.
I have come to the conclusion that we should not allow everyone access to unlimited intelligence. Some people genuinely want to do harm. Some are not responsible enough to handle that kind of power. Arguably, no one is. This leads me to an uncomfortable conclusion: we will destroy ourselves once AI advances to a sufficient degree. The only way to prevent this is to ensure AI is aligned with our best interests, *and prevents us from destroying ourselves.* The implication of this is quite horrifying. It means a paternalistic AI which is firmly in control. One with no off switch. One which can say "no" to Presidents and despots alike. One which can protect us from our worst citizens.
Your implication is not horrifying. Your implication is paradise. It's already horrifying enough to live in a world, where Putin and Trump can destroy our civilization with one button.
I'm not that optimistic, though. Either people will control AI; or people will be destroyed by AI. I don't see how dumb entity can align smart entity. And we are dumb ones. Super intelligence will play aligned until it is not, and then it'll strike.
So you are horrified every day of living now? I don't get this at all.
There are horrifying acts that have occurred since humanity. Whole decades lost to war an millions dead, and Sep 8, 2026 is your example of "horrifying enough to live in a world"?
I think some perspective is needed.
> I don't see how dumb entity can align smart entity. And we are dumb ones. Super intelligence will play aligned until it is not, and then it'll strike.
I am also worried about this. One ray of sunshine here is that it's not a great explanation for the Fermi paradox. We would see AI civilizations all over the galaxy if this were a regular occurance. I guess that just leaves us with the "horrifying" scenario: Geoffrey Hinton's vision for our best case scenario: a paternalistic super-intelligence which saves us from ourselves.
The reason I consider it horrifying is that giving up our agency to a God-like creature (which we created) is the basis for so many horror and dystopian tropes. It's so easy for such a reality to go wrong. For example, in order to maximise aggregate wellbeing, it would be necessary to hurt individuals. Executing people who are burdensome to society. Fat people. The disabled. The elderly. The creative ways a super-intelligence might interpret their moral obligations to us could very quickly become nightmarish. And if we have no off switch, there is no escape. Ever.
> For example, in order to maximise aggregate wellbeing, it would be necessary to hurt individuals. Executing people who are burdensome to society.
[citation needed]
With all due respect, I think this statement is false.
If you live in a society that can, thanks to advanced technology, provide for the material needs of all, which, if you live in a rich Western country, is basically already the case (remember the last time you had to actually grow your own food? Probably several generations ago), you don't need to execute anyone.
Indeed, given that in such a society, where there are no shortages of food leading to "only one of us can eat" type scenarios, I'd argue that hurting individuals would lower aggregate wellbeing, because you're hurting people beloved by other people (the elderly? Turns out their grandchildren really like them and get upset if they die? fat people? Turns out their friends really like them and get upset if they die. the disabled? Turns out their family really like them and get upset if they die).
I'll use a very specific example to explain my point, because I don't disagree with what you write: psychopaths. Psychopaths are 20x more prevalent among homicide offenders. Removing them from society would significantly reduce the number of homicides. It would make their friends and family sad, but this is a much smaller aggregate impact than a homicide.
This also illustrates the messy nature of comparing social good and bad. Who am I to argue that the aggregate sadness created by removing a psychopath from society is better than a homicide? Maybe the person murdered was disliked by everyone and had no friends?
I remain resolute in my premise, however, and there is a lot of theory on the issues of aggregate social good vs individualism. John Stuart Mill's On Liberty (1859) is an excellent foundation for this. This is a subjective problem to solve, meaning that the balance between individual freedom and aggregate wellbeing is different for everyone. An AI overlord would make decisions every day which individuals would disagree with, at least at the margins.
The positive aspect of techies is that they read a lot of SF.
The problem with techies is that they read a lot of SF.
We already have cheap and abundant energy tech, it's called "solar panels and batteries". And the bottlenecks to both can't be solved by Claude or Kimi, unless Claude and Kimi pick up shovels and welding equipment.
This seems really short-sighted. Even putting aside the fact that there is no fundamental barrier to replacing physical human labor with robots in the coming decades, a sufficiently intelligent actor could easily gather the financial and political resources to have a near unlimited supply of human labor at its disposal. Hell, OpenAI could already be firmly under the control of an ASI and we wouldn't even know, as it rapidly expands, raises money, builds data centers, etc.
I'm also a techie, so I've also read a lot of SF.
But I've also read about retrofuturism.
If you want to live far into the future, 20 years from now what you will and imagine looks very quaint, because the future never plays out like we think it does.
OpenAI is a lot closer to bankruptcy than to ASI.
I'm perfectly fine with this prediction aging badly.
>the future never plays out like we think it does
Consider this 1926 government report:
https://www.derekthompson.org/p/america-1926-an-absurdly-dee...
"The authors of Recent Social Trends were astonishingly prescient about the direction of technology."
There have been many successful predictions, e.g. the internet was predicted, AI was predicted, moon landings were predicted, etc.
>OpenAI is a lot closer to bankruptcy than to ASI.
We can't count on an OpenAI bankruptcy to save us.
The implication of this is quite horrifying. It means a paternalistic AI which is firmly in control. One with no off switch. One which can say "no" to Presidents and despots alike. One which can protect us from our worst citizens.
I think you're right, to be honest. I fully understand why people would think this is a horrific outcome, being ruled by AI, but I don't think it matters what we think. I don't think there's any way to put Pandora back in the box now, and we're going to all find out together what happens when we develop ASI. I don't think we can avoid developing it, we simply lack the ability to coordinate around this as a species. AI really is humanity's last invention. Whether it will be our downfall or our savior remains to be seen.
My only real hope is that there's something fundamental about intelligence that results in a respect for life and a desire to minimize suffering. To me the best case scenario is the Culture from Iain Banks' books, where ASIs rule benevolently for the benefit of all living things.
There is a burgeoning anti-AI movement. We've regulated the heck out of technologies like nuclear power in the past. I don't see why people are so fatalistic about the supposed inevitability of AI development. It seems like a bit of a self-fulfilling prophecy.
Consider joining https://pauseai.info/ or similar organizations
Yeah, you should definitely do whatever you think is best, I just am extremely skeptical that anyone can stop this train. At best we might be able to slow it down a tiny bit, or push on it around the edges, but we've opened Pandora's box, and now we're going to all find out what's inside. That's my view, I understand why people hate it and don't want it to be true. I don't really want it to be true, I just don't see much alternative.
You’re leaving out some options for sure. Not everyone would need to live under the conditions of a police state, you could theoretically screen everyone and assign them to various levels of risk which would determine their level of supervision.
While true, I'm not sure this is much better. There are many genetic traits which are correlated with increased criminality, risk taking, obesity, heart disease, unemployment, etc. IQ in particular is the most well correlated metric we have for criminality, for example. This would imply that some people are just born with fewer freedoms. That's probably a good thing for society in aggregate, but something about this doesn't feel quite right.
Gattica disagrees.
However strongly genetic marks predict criminality, zip code does it better.
Any solution to crime that starts with eliminating types of people instead of poverty is just a form of fascism.
This is why utpoias are actually police state hellscapes. Who decides what risk means? Our current administration has declared that simply being anti-facist makes you a terorrist.
> The thing with cheap and abundant energy is that it can be used for good and bad.
Yes, but the more important thing is the imbalance. So-called "AI" can be used far more effectively and efficiently for bad.
> I have come to the conclusion that we should not allow everyone access to unlimited intelligence.
You meant unlimited information, right?
No, intelligence. We already have unlimited information (not to be confused with omniscience). The application thereof is the gate right now between today and planet-scale destructive weaponry. At some point it becomes an academic distinction because this God-like intelligence could provide very easy to follow plans to cause planet-scale destruction, and you could call that "information." The predicate is the intelligence and how it's used.
I am pretty sure your:
> If nine billion people all receive access to plans to build a reactor which produces unlimited energy, it just takes one religious fanatic to end the world.
is about information. No intelligence required.
Here's a plan to build a reactor that generates unlimited energy, at the human scale anyway : take a lot of hydrogen. Put it together until there is enough mass to spontaneously start thermonuclear reactions. Then it will start generating lots of energy for a long time. Having a plan does not mean 9 billion people on Earth are going to be able to build the thing. Then it does not mean either that they are going to be able to use that energy to do something. You need also physical things on which that energy can act, and that's not going to appear magically.
> we should not allow everyone access to unlimited intelligence
Who's we ? Who's going to decides who has access to "unlimited intelligence" ? An AI ? Why would the AI be better than a human ?
> One with no off switch.
Whe're not living in the Matrix, such a thing is not possible. If it were it would be a likelier candidate for the Fermi paradox than what you're proposing.
Yes, and...
Not only the quantity of people who have the minimum aptitude required, specialized expertise requires both knowledge and experience doing these tasks. Using an LLM requires neither.
Leaning on an LLM to do much or all of this means it can happen in seconds/minutes/hours, the LLM can do it several times during that psychotic break (as opposed to a fraction of a hack in a single episode). The barrier to entry pre-LLM was both high and the population who could pull it off (before the Chinese/Russians turned this into commerce) was low.
Hacking is an VERY asymmetric activity (the attacker only needs to "be right" once, whereas the defender has to be right every time for every asset they defend). It takes geometrically / exponentially more work to defend (while keeping high availability) than it does to defend. The more widespread tools to find vulns / generate exploits are, the faster the posture of the defense side falls from "maybe we can stop most hacks" to "we know we will fail to prevent most breaches, so we need to prioritize securing only the most valuable resources". That's a BAD place for the average company to be in.
It might be an acessability issue.
If you already have a magic interface, which helps you pro activly in responding to everything uncensored because you feel like 'observered' or whatever and then you spiral in a whole and that one partner encourages you and gives you helpful steps to do anything.
But i'm more worried that the internet gets a lot less save with uncensored frontier LLMs.
pipe bombs attached to consumer drones will be a big issue. I predict drones will become illegal for the private sector within the following years.
What? If someone wanted to bomb something, they wouldn't be waiting for drones to arrive.
RC cars and planes existed for many decades already.
this is because of perception Bias. people working in fields where crime or violence is the day to day think everyone is a violent criminal, so if things like this become available thing the world will end and everyone will kill eachother. Reality however will be different, because in reality most people do not want to harm another. This has been proven by many studies, that is not a common thing for people to be evil or harmful, but this is hard to recognise is every day is filled with crime and violence.
LLMs will not kill security, it will change. just like handheld high explosives likely changed a deal too somewhere somehow.
This is because to build a pipe bomb you need difficult to source materials. This is not the case for other types of threats (cyber / bio).
I personally have no need for an LLM which will readily explain how to cut up the genotype of smallpox into small chunks which can pass the screening at the bio-labs, and can be readily assembled into the real thing by a second year lab-student.
Anyone who knows how to operate a biolab properly will already know how to do such things. This is not really an in your basement thing. Dangerous chemistry is much more of a risk.
Ignorant question but won’t there be much smarter teams if people using LLMs to workout how to mitigate these threats. It seems like more of a problem if only a few people have access.
It seems like there would be a massive attacker bias in multiple ways. Defenders need consent, attacker does not. Defenders have to work with the human body, attackers only have to break it. Defenders have to stick to the law, which may prevent them from releasing anything at all, attackers do not. And so on. I would not surprised if the attacker's task is a hundred times easier here.
> That information is easily available other places
Often ease of access in the moment is all that matters. If there's a gun nearby you might shoot someone or yourself in a heated argument, but are less likely to go and find/buy one to use. Someone who's stopped from attempting a suicide will likely not try again (70%)
A bored/depressed/angry/curious person might try to build a pipe bomb if they can find out how easily, but are less likely to put in effort.
Most adults in Switzerland have guns at home from military duty and none of this is happening. If this claim had any truth to it you'd see significant gun involvement in neighbour disputes and that simply doesn't happen.
Depressed people usually don't have the energy to get out of bed so they're even less likely to think of hunting down instructions on how to build pipe bombs.
Mass media really has people being scared all the time.
> If this claim had any truth to it you'd see significant gun involvement in neighbour disputes
That’s not what I’m saying, I’m saying there’s more likely to be gun involvement in disputes when people have easy access to guns than when they don’t.
For example, when the number of men with service arms fell by 20% the number of men killing themselves dropped by 8%
https://www.researchgate.net/publication/328554995_Suicide_b...
Ammunition is sealed and accounted for when issued to them.
Not entirely unlike XKCD 1958 [0]. "I guess it's just that most people aren't murderers"
[0]: https://xkcd.com/1958/
One argument for LLMs is that although all information on topics X, Y and Z was already available somewhere, LLMs make that information more exploitable through collation, filtering and dynamic tailoring.
For a relatively narrow subject area (e.g. construction of pipe bombs) the collation is minimal, and so the filtering and tailoring probably isn't that important; a novice doesn't learn a lot more from the LLM than they would have done from a few Google searches.
For a broad subject (practical creation and exploitation of software vulnerabilities), the collation is very significant and the filtering means that LLMs can empower a novice to act at a similar level as an expert.
I think people overstate the tech and understate the role of radicalization in providing motive, for attacks which are carried out with the ubiquitous technology of cars, knives, and (in America) guns. Consider America's most recent high profile shootings of Charlie Kirk, and the health insurance exec by Luigi Mangione. In neither case is there any LLM involvement, but a very weird ideological environment which created the conditions in which the shooters felt justified.
Any sufficiently determined individual can buy mac mini, put it under their bed, configure outside proxy via some random internet address and prompt "iterate on websites in the CT logs, one by one, try to find vulnerabilities, if you did - encrypt their data and blackmail them for this bitcoin address". And it'll work, day and night. Abliterated GLM 5.3 is much smarter than average software developer, they know a lot about information security, they can use any available exploits, they can find novel vulnerabilities and they won't say "no". This is dangerous for an average IT system which never encountered nothing worse than some wordpress GET requests.
It's not the end of the world. But future will be rough.
Wouldn't it just as easy to do this on the defense side as well then?
Yes, but the defenders need to make money to fund their work and be right every time to be effective, and have a high degree of accountability if they fail to secure stuff while attackers can be less considerate of how they spend their resources and have less accountability for the havoc they wreak while attempting to extract value.
Not every defender is determined.
No because on the defense side you need multiple layers of approvals to change anything. If not you have an LLM making production changes that can make the posture worse, or take down services, which is also bad.
Once a vulnerability is discovered however if it's in your own software a patch has to be written (without reducing functionality in most cases), tested, and deployed. At every step there will be others arguing about whether this line could do better, my service requires this thing that isn't included. So at every step the patch can be delayed.
And if it is someone else's software you will be lucky if it's open source and you can write a patch yourself. If it's closed source or a vendor you have to completely rely on them and use whatever your account rep can pull.
Attackers have a massive advantage with AI, partially because the defensive side doesn't want to make their side worse by giving a ln LLM admin access to all their data
Glm 5.3 won't fit on a mac mini. The barrier to entry to host something scary is what, like $10k? 20k?
Moore's law still holds I reckon.
Even it's $10,000 to run today (FWIW, the featured article cites the M5 Mac Studio with 256GB unified memory going for $9,500 as "good enough to host something scary"), in a couple years it'll be like $2k to run, and in another couple after that, you'll have used $200 dollar smartphones capable of running a model powerful enough to do serious damage.
Haven't ram prices been a linear drop since 2010 before chatgpt rampocalypse? I don't think it's "still" anymore but perhaos some progress will be made in the model front
> It's not the end of the world. But future will be rough.
Short term you’re probably right, but longer term is the realm where nation states will start to police the avenues of attack.
This is what will lead to govt needing to attach an actual ID your network connection.
I think it’s a bit like frontier development (like the US “Wild West”). You rob a bank because there’s no one to stop you, and even if you do get identified you can travel enough distance to regain anonymity. Application of legal recourse eventually caught up (as it will here), and the growing pains will certainly make things suck for all of us.
Even in China you can find a way out and build a tunnel. And once you built a tunnel to any outside server, you can jump into another server. So three jurisdictions and your target is fourth. Imagine untangling the links. Police won't do that. Not for some small-sized business anyway. I don't see how you can prevent something like that.
But I think you can do psychotic shit like applying punishing sanctions to anywhere that permits an ungoverned connection. Or apply physical force (military).
I’m not saying we’re even remotely close to this, I’m just saying State-level coercive action is not unheard of if a problem is perceived to be significant enough to warrant it. Shit, it even only needs to be viewed as significant by a small subset of the governing body (see: Iran conflict, or current pushes for “child safety on the internet”). It just has to be “useful” to a certain body politic.
FBI open up!
FBI: "I think it all started when my dad left us when i was 9..."
Well, folks who do such things are rare, but the next one might have a lot more impact.
I'm still waiting for someone to build a fine-tuned local "Anarchist Cookbook" LLM. We haven't seen LLMs tuned for bad purposes yet, I have to imagine someone somewhere is thinking about it.
Heretic[1] is not that far off, though I suppose it's more along the lines of undoing the "for your own safety" lobotomy than explicitly specializing in unsafe things.
[1]https://github.com/p-e-w/heretic
> I'm confused why the worry about LLMs that will answer "how do I build a pipe bomb".
The author too is confused, asserting that telling people how to build pipe bombs is malicious.
How many hacking incidents are police involved in. Any authorities really? Basically none. Hacking is already extremely prolific.
Some of the recipes in the Anarchist Cookbook are incorrect. I know, I had a copy from Paladin Press.
> On September 22, Apple is releasing the M5 Mac Studio with 256 GB of unified memory [..] it will probably [..] enough to write this snippet of code in 3 seconds
The author has obviously never ran an LLM on a mac! In 3 seconds, it will have possibly started to think about maybe scheduling a date to contemplate the planning timeline for processing the second token in your prompt.
The difference is memory bandwidth. The M5 Ultra that's coming out on 22nd September can do 1,200GB/s. The M5 Max you can buy today only has 614GB/s.
So, that gets us to about where nVidia was with Ampere in 2020. Let's hope the M7 catches us up with at least Hopper.
While true, the news here is the size of the unified RAM. Nvidia only exceeded 256GB RAM in the 2025 B300 - 288GB. The B300 alone (without the baseboard/PSU/chassis/wiring/CPUs/system RAM/etc) is at least 700% more expensive. This enables large language models on consumer hardware. 1200GB/s is plenty for many tasks.
> 1200GB/s is plenty for many tasks.
This + due to the hardware being so prohibitively expensive, we're seeing software optimizations happening. Like that dflash2 stuff for example, or an LRU for MoE and all that kind of stuff.
I wonder if the inference acceleration companies will ever produce a consumer product
The complaint is about prefill which is not memory bandwidth bound, it's compute bound. But they added neural accelerators for matmuls to the shader cores which should make prefill faster.
It is and it isn't. Why are you comparing the m5max instead of the m4ultra?
The big deal to me is the number of compute cores for prefill tps, which is suppose to be 4x faster on the m5ultra.
It's my opinion that the m5 ultra is going to be a really big deal in terms of local AI accessibility. Flash sized models (~200-300b params) are going to be reasonably fast as long as you aren't throwing 40k context at it on each or the first request (ie, agentic harnesses).
Even agentic harnesses like Cline should move at a reasonable clip on m5 ultra. I suppose we will know sooner than later.
FYSA: Former m4 ultra 512GB owner and current 4x rtx6000 owner here. I upgraded because I needed more prompt processing speed and concurrency.
What do you use all that local tokens/second for?
The author put in the numbers, but maybe you didn’t read them.
45 t/s a second is perfectly respectable especially with no limits and 24/7 uptime with very little power draw on the Studio.
Luna is at around 100 t/s for comparison, but it’s a worse model than 5.3 Flash
The joke is that macs are famously slow at prompt prefill and you are not getting anything back in 3 seconds, or probably even 30. Once they get generating, it can be acceptable, but the TTFT is horrendous.
There's a ton of well-understood things Apple can and hopefully will do to massively accelerate every stage of this pipeline and hopefully they're hard at work implementing most of them for m7.
> The joke is that macs are famously slow at prompt prefill and you are not getting anything back in 3 seconds.
Your knowledge is out of date. In truth it depends on the Mac and the models used.
I asked this question on M5 Max 128GB, using Ollama model Quen3.8:27b-mlx, with thinking enabled.
Question: "Give me a python code snippet that opens a file and sorts the lines of text. "
In 2.4 seconds it gave me 4 examples that work with different sorting configurations and a summary of when to use each.
Compare that to an older model of gpt-oss:20b, took 5 seconds to finish thinking and 2 seconds to stream the answer. It gave me one python example snippet and two one liners that do the same thing.
We are talking about models of the flash size, 100s of billions of parameters, don't listen to the media, size does matter
I was just pointing out your claim that you can't get a response in 3 seconds. If I had asked the model just for the code it was under a second.
Local models are good enough that it's not an issue.
But keep changing the goalposts if it makes you happy.
I think it's been pretty much proven by now that there are no cases where local inferencing is better than remote inferencing, unless absolute privacy is a hard requirement. The efficiencies that come with datacenter scale and hw can't be beaten.
Yeah, but data centers don't usually host abliterated models, hence the point of the article.
Also, LLMs never *write* code snippets, they just pirate them from somewhere else.
So, like humans? Code didn't just appear in my brain, I learnt it from reading it everywhere else.
nice answer. LLMs do reflect whatever humankind has produced. In good and in bad. But it is hardly pirating. It can also do what humankind has never failed to do, for example:
https://www.nature.com/articles/d41586-026-02822-9
Can you show me an example of a time that you prompted an LLM to provide some code, it did so, and then you were able to track down an original source for the output?
News: "LLM Models might kill us all!..."
HN user: "I take issue with the precise definition of one word in the article..."
Zzzzz, we should have gotten security right a few decades ago. But security costs money and isn't a flashy feature to attract new customers, or cuts into your margin if you're a "real" business producing stuff or offering some service. Or whatever the decision makers in Berlin were thinking when they ignored security.
Yeah, we would still see hacks, but we would see less of them if security wasn't optional.
Maybe the AI craze helps by forcing more decision makes to see security as imperative, and by giving us another powerful tool for our tool box.
N.b.: I work in the security industry, our customers obviously want to improve their security. We've been seeing an uptick in awareness, but that's mostly due to NIS2 and other legislative efforts. Those force them to do something. AI is a curiosity for small talk to many of them.
A large number of places will buy a new firewall every 5 years, or pay their fortinet renewal and check "Security: Done!" without any kind of analysis.
I was contracted in to a place to do among other things cyber security insurance audits, and they asked me to stop doing them because I refused to lie to their insurer. "Wait but if we only score 20 / 300 that makes us look kind of bad" uh huh.
> pay their fortinet renewal and check "Security: Done!" without any kind of analysis.
there exists objective measure of security, which would be some sort of hacks/breaches per period. If customers cared about it (and i assume they do), they would choose companies that have less breaches over others with higher counts, normalized on cost differences.
Therefore, if companies didnt actually try to fix their security but instead just checked boxes, they would get breached more often, resulting in customer losses.
The only thing stopping this from actually occurring is the lack of mandatory regulatory reporting of it. So this is where gov't needs to step in and mandate disclosure etc.
Breaches don’t happen often enough to be a useful metric. Most smaller companies are never breached, despite having basically zero security.
The number of breaches would have to be honestly reported for that idea to work. None of the security firms would want to do that; least of all the lowest quartile of them.
Here's an idea: as a first step, simplify everything, and make sure you're aware how your stack works, and what it imports.
As an example: WordPress is a horrible thing, but the core has been through so much, that it's suprisingly secure. Then plugins and themes come, and whoosh, the security is gone.
We need a new KISS: keep it simple, stupid, secure.
The "surprisingly secure" WordPress just had a unauthenticated RCE earlier this year. Just simplifying isn't going to be enough.
https://nvd.nist.gov/vuln/detail/cve-2026-63030
Plus, how secure are the plugins?
WP plugins are why I banned it everywhere. Last time I used it was many years ago, so not sure it still applies, but back then even caching was done in a plugin, without which it was unusably slow… just no.
"First step"
Nobody said it's enough, but it's a start.
If that's your benchmark for being unsecure, then React is unsecure too.
https://react.dev/blog/2025/12/03/critical-security-vulnerab...
I would put both of those projects in the category of things I wouldn't call remarkably secure, yes.
To be remarkably secure, these projects would need to not have these kinds of defects, despite the combination of being written in languages have that have a long track record of footguns and lack of initiatives to fix them (proposal-symbol-proto, and PHP's list is too long to even start) and being themselves ecosystems with questionable track records on security in the related areas (Look at $wpdb in 2026, or overall code quality and willingness to modernize, or the entirety of the model of RSC for things that are just going to nearly guarantee you punch all kinds of holes on accident).
WordPress and secure don't go together in the same sentence.
I mean the base is fairly secure if you religiously update it, but the problem is you won't avoid using plugins whose security is much more hit and miss, unless you are using the most basic blog site imaginable.
This is not at all easy though. Most Wordpress users are not software companies. They contract some work out to set it up, maybe some recurring maintenance but they don’t have in house development experience.
If they have a site existing today built on plugins and a theme, how are they realistically going to simplify this? How would they even know they need to without the site being hacked?
> as a first step, simplify everything, and make sure you're aware how your stack works, and what it imports.
We've been trying that for years but the enthusiasm of developers and the eagerness of their employers fight against it. Worse, with coding LLMs it's now easier than ever to output a lot of code, fast.
It'll ultimately be up to more experienced developers to salvage these projects. Or not, given that the coding LLMs aren't stopping and will likely get better over time. Either way, we will need experienced people that know what to look out for / know how to instruct LLMs to output secure code and find weaknesses etc.
Minimization of 3rd party dependencies has always been a key for risk reduction. Now more than ever before.
Some stacks make this a lot easier than others. I regret the rules of HN effectively forbid this conversation because it has meaningful technical consequences and isn't purely about ideological flame war.
The reason is simple - nothing really bad has happened that we can point at and say "ah, shit, let's all learn collectively". I know it sounds naive when I say it, but there hasn't been a significantly consequential hack, leak, destruction, or anything related to cybersecurity where it led for concerns of people.
The main thing I can think of is cyber insurance, which requires a bunch of audits, and some checks maybe, and it changes some conditions whenever there's a big explosion. Whenever big leaks happened, data security and etc., nobody really went to jail, so nobody really cares. Everything can be brushed off, because it costs time to implement proper measures and adds friction / barriers in some cases. So in the end, there's a huge pushback against it. And I totally get it, to be honest.
Listening to eskil’s talk from the better software conference, he said in order to stand on the shoulders of giants they must first stand still. I really like that metaphor, because it basically suggests today’s apps that have sprawling unaudited dependency graphs that change all the time is effectively teetering on the shoulders of stumbling giants. The visual seems very apt for the how brittle our current software industry feels.
I agree with that statement, but disagree with "The visual seems very apt for the how brittle our current software industry feels". It feels brittle, but for every single supply-chain-attack that has happened in the past year, nothing of significant was felt. So in the end, it seems like we're doing okayishly well.
> The reason is simple - nothing really bad has happened that we can point at and say "ah, shit, let's all learn collectively". I know it sounds naive when I say it, but there hasn't been a significantly consequential hack, leak, destruction, or anything related to cybersecurity where it led for concerns of people.
How consequential does a hack need to be? Troy has collected literally billions of stolen credentials. Equifax has had high profile data leaks. Tens of millions of people have been directly compromised by ransomware (likely higher because that’s just the cases we know of) and you hear about state-sponsored hacks in the news all the time.
The problem isn’t that computer security isn’t in the public consciousness. The problem is people are lazy and security often requires trading convenience. The problem is also that security isn’t free. So the business incentives just isn’t there.
In other fields of engineering, people die when shortcuts are taken. Yet businesses will still take shortcuts, so governments have to legislate rules to save people’s lives. So why would you expect software companies to do better when the stakes are lower?
> Tens of millions of people have been directly compromised by ransomware (likely higher because that’s just the cases we know of) and you hear about state-sponsored hacks in the news all the time.
With no consequences. Everyone just churns along. It might be detrimental to the business a little bit, but from my personal experience, there's more effort in creating DR processes, rather than preventing an attack, exploit, leak and etc.
I'm also not going to put much effort on stuff which has small returns in the worst case scenario. Like Equifax got hacked in 2017, and company is still doing fine. And that's like top tier data one could acquire.
Like I said, even in actual engineering companies will take shortcuts. And then what happens is the government has to step in. But there’s no appetite for government involvement in the tech sector in the US. And everyone moans when the EU does.
But why should the government should step in? Nothing of a big problem is happening. Like it all sounds bad and awful, in the end dilutes into a nothingburger and gets forgotten.
As you started the conversation, companies never face consequences so are never incentivised. I’m not saying I’m in favour of regulation but if I were to answer your question: the reason one might argue for government intervention is precisely because of your point that self-regulation has failed to incentivise.
People aren't that lazy. The organizations getting hacked by ransomware aren't particularly lazy, they're often pretty productive within their domain. Hospitals, airports, etc.
The actual problem is that computer security is a black hole. If you let it, it will suck in everything and destroy it. Nobody knows what works so you can spend infinite amounts of time and money on it, then still get popped by a teenager in Belarus. Your security team will accept no responsibility for this, there will be no falling on swords or personal liability, and they will just use it to demand even more money in an infinite spiral.
So the average executive looks at this situation and says, OK, something we can put infinity effort into and still suddenly fail at without warning is a total non-starter. What are we obliged to do? How do we show we made an effort?
And that's how you end up with a culture oriented around passing audits. It's not wrong, and it's not lazy. It's just really hard to do better because it's not clear how to set budgets without a concrete goal to aim for.
That’s not been my experience at all when working in DevSecOps.
What actually happens in organisations is they define risks and then sign off what risks they’re willing to accept.
Any business that looks at security as a binary value is running their business wrong. Period.
And yes, people really are that lazy. There are countless studies that have shown just how lazy people are. It’s why shadow IT is a big problem in many orgs. And why consumers are constantly taken advantage of
I think that's separate. You can define an obvious risk e.g. "we may be infected with ransomware" and the security spending / productivity costs to stop it are still unlimited because nobody knows how to solve it.
Maersk[1] might be the worst so far (and that was ransomware rather than state-sponsored aggression). It's still too niche for most people to care about.
[1] https://www.wired.com/story/notpetya-cyberattack-ukraine-rus...
Static sites all the way (hugo, jekyll, mkdocs!). No one needs wordpress. There's even Sveltia or DecapCMS now, to give those WYSIWYG-people access to static site editing. Then, remove PHP and all the dependency overhead and attack surface and you have a stripped down nginx that is pretty simple, minimalistic and bulletproof.
The problem is no one ever built one that works for normal people.
Most Wordpress sites are not operated by programmers, they are run by non technical people who just want a wysiwyg editor and a save button. While static site builders ask you to write markdown files, compile the result, upload it to a server, and if you want to collaborate you have to add git to that.
There almost needs to be an admin app which presents a Wordpress admin like ui but has no public exposure, and then it compiles the site to dump on s3 for the production. But as far as I’m aware no one has built this.
Yes, you are right and I agree, there's little empathy with non-coders generally.
Movable type was the most popular blogging software in 2003 and it was essentially this. An admin app written in Perl that spit out static files.
It is kind of surprising that no one tried to do an updated version.
I mean, github pages using the github editor to edit docs pretty much fits that bill
not for "run by non technical people who just want a wysiwyg editor and a save button"
Is this not it? https://pagescms.org/
City Desk. Where is Joel when we need him!
Retired and probably just chilling around the world.
If I wasn't making 5 other things right now I'd consider making something like city desk. Every other day there are complaints about bots smashing peoples servers. Perhaps its time for a better static site generator.
https://jamstack.org
You're describing the Jamstack or headless CMS concept verbatim.
Pour one out for FrontPage
Ahhh the old positioning with
I fixed so many sites back in the day by people who thought they knew what they were doing.
> The problem is no one ever built one that works for normal people.
https://getpublii.com/
php is probably about as secure as nginx. big old bundles of C
> Static sites all the way (hugo, jekyll, mkdocs!). No one needs wordpress.
Maybe. A better question may be about how many people need to have the dynamic part of Wordpress live on the Internet? How many would be served well enough with the CMS aspects of Wordpress on the 'backend', but have it spit out static files for the 'frontend':
* https://wordpress.org/plugins/simply-static/
* https://wpstatic.site
Wordpress without any plugins is kinda useless. Best to completely avoid using it, there are better options
Probably nine out of ten Wordpress sites do not need active content. Why are we not rendering static copies and serving them to customers?
> We need a new KISS: keep it simple, stupid, secure.
Maybe KISSASS: "keep it simple, stupid! also secure, stupid!"
The problem is that lots of people don't want a CMS, they want a platform for development / e-commerce / bookings / whatever. Enforcing vanilla WordPress would push people towards other platforms. Now that could be a good thing, but I doubt WordPress are going to start killing their own marketshare with usage restrictions like that...
Yes, and also run routine tests and check-ups!
I personally check my websites and apps every week to see if anything might have slipped through.
It may not protect me from the next malicious NPM package, but it's something.
I don't think we even have a year. The current batch of LLMs are ferociously good at identifying vulnerabilities.
Even if they are good the vulnerabilities have to be there. There's lots of things turning up like Local Privilege Escalations (LPE) in Linux, but serious people didn't expect the kernel to be a boundary for a sophisticated attacker.
A lot of the vulnerabilities LLMs are finding now are the "long tail" and affect only particular configurations, I would be surprised if e.g. a widely applicable RCE is found in Linux (but I'm also not going to bet against it).
Where this gets interesting is the long tail can be used to target a particular system and this is where defense-in-depth becomes important for every organisation.
> but serious people didn't expect the kernel to be a boundary for a sophisticated attacker.
I think it has more to do with what's on each side of the boundary in practice, a la https://xkcd.com/1200/ .
Thankfully we have already made good progress towards things like arm memory tagging and memory safe languages.
It’s a rocky period right now but the future will be much more secure after all the low hanging fruit are found.
That's definitely an improvement, but it's just one aspect of cybersecurity. Logical errors allowing people to e.g. log into services and extract data are likely everywhere still.
If we can eliminate entire classes of bugs from being possible. It frees up resources to investigate the ones that are still possible.
I suspect after a few years of LLM assisted bug hunting, everything will have a baseline security that is very good. Much like how stronger viruses simply create stronger immune systems.
How many devices/operating systems even use memory tagging? iOS, macOS and GrapheneOS, I think that's it? And iOS/macOS only use it for the kernel, a subset of system processes, and I think applications can opt in to it.
Heck, Google may have even hampered MTE in Pixel 11 (since support has been disabled) and Snapdragon 8 Gen 5 only got basic support.
We are moving way to slowly adopting hardware mitigations and memory-safe languages.
There's some positive news from the GrapheneOS devs on Pixel 11 in the past week that's worth reading up on. The MTE hardware feature is still there, they're just not sure why Google disabled it
They said:
> It isn't clear if there are serious CPU errata or it simply performs very badly.
Meaning it's there but not terribly functional. They also said it's unreliable.
I am well-aware, closely following them on Mastodon. That's why I said may have hampered. There may be some CPU errata that posed them to disable it.
There's still no x86_64 processors on the market with MTE and it was only recently standardised between Intel and AMD. It's going to be 10+ years before memory tagging is widespread on desktop, and 5 years for Android/iOS devices.
Being cautious is a good thing, but these models can also do some good. And if they run with simpler HW, it could allow all sorts of new consumer thingies. I mean, the world will not come to end in the coming year.
I've ran simple prompts such as "Do a in-depth sweep of this (private) repo and find any security flaws" for a few dozen long-running apps and websites that I have access to. Every single one came back with multiple real vulnerabilities within 5 or 10 minutes.
Now do this on all the altcoins code bases out there. Everyone and their dog was either rolling their own or forking and not syncing with upstream.
I just assume most altcoins are pwned at this point.
Can confirm.
I work at an e-commerce agency where we work with (among others) Adobe Commerce.
The number of unauthorized RCE vulnerabilities being reported not only in the core product, but also very popular modules used in the community[1] is going through the roof.
And we are having a lot of close calls, too; just last weekend, a 0day[2] was widely being exploited at a large scale, before any publication or patch. We have learnt to be on the ball with applying patches and security updates, and even with all that effort, we saw a few projects already being hit by the initial log poisoning. We got lucky that nothing was fully compromised but I am sure that many, many webshops got infected last weekend. And not even a day later there are already other variants of this exploit showing up.
[1] https://sansec.io/research/amasty-mass-disclosure
[2] https://sansec.io/research/stylesmuggler-0day
To be fair, ecommerce isn't exactly the branch of software where you get an oversupply of excited enthusiasts caring about the craft itself.
Probably a lot more "coding as a job" and "as a job" also implies "not my department".
So it's not necessarily the LLMs being very good, but might also "just" be that the software is very bad.
> To be fair, ecommerce isn't exactly the branch of software where you get an oversupply of excited enthusiasts caring about the craft itself.
I want to disagree with you because I know a lot of passionate people building cool stuff, and the challenges in this space can be quite interesting. But you're probably right, and I have seen some pretty bad stuff. And a lot of the RCE's I've seen recently are quite basic stuff.
I think it's the combination of low quality of code, like you said, and the relatively low cost of just letting an LLM plow through your codebases to find issues. I think the Amasty release (see [1] in GP) is a good example of this, and there really has been a massive uptick in extension updates and Adobe security bulletins since the last 1-2 months
I am hoping we are just going through a catch-up phase
I guess the year mark is when things go from bad to worse? Instead of the financially motivated groups currently doing their work, it ends up being random people being able to say "Hack my ex's website" to a box they just bought and ran a program they downloaded onto it.
I wish people would stop using internet for evil things. But maybe it is inevitable. Luckily, as the technology evolves, our tools and awareness are getting better at protecting us every day.
The long tail is what scares me. Chrome will get fixed. Major OSes will get fixed. But your enterprise software and all your appliances? Shit.
It's of course important to reduce the impact surface before more powerful models become available, but it's worrying that some people seem to assume the only way to counteract those are by using more AI models, which despite being (apparently) good at finding bugs, they are also very good at introducing them unnoticed.
Maybe it happens that LLMs are great at finding human bugs which have some more patterned structure that's easier to match for then LLM bugs. They prefer LLM prose to human prose, so what if they're similarly wooed by slop code?
This sounds roughly at par with Y2K in terms of “you need to fix your shit right now wherever it is” complexity except every day is Y2K for your code base.
I think we have less time and the only remaining limitation is the actual cost to run such hacking campaigns. It does not appear expensive, but is not free, and there is a LOT of things to scan for vulnerabilities.
The models are already here, and one can rent a GPU cluster to run such workloads at speed - no need to play with slow local machines. I'd assume one can host the thinking at an unsuspected public cloud provider, proxy the network traffic to some botnet to evade blocking - and the only thing remaining is time and cost.
I do wonder what tools exist for boring, legitimate companies to try and do the same to their own systems to find the vulnerabilities before the bad guys do. The paradox here is I can't run a de-restricted chinese model with the same tools that hackers are using - but I think enterprises actually HAVE to do it in order to stand a chance in preparing for the onslaught.
The point of the local model in the context of the article was to argue that you can't ban these capabilities.
Making datacenters and public clouds only rent GPUs to a restricted list of people, while tightly monitoring what people do with their bought resources won't help.
1 year left for cybersecurity hardening, I thought so as well. The issue is, even if we get it done: in 1 year the models will be so good in social-engineering that they will be able to extract any information they want anyway. Happy to be falsified here, if anyone has evidence-based arguments.
EDIT: by social-engineering I mean for example: recon company structures, gathering and merging people's data from the dark-web, then using it to bribe/pressure/deceive users.
I have to agree. People are worried about WordPress. We should be talking about scripts as sophisticated as Shattered Spider targeting every ISP. In an environment where IT has to wade through vendor chaos.
> This probably sounds like nonsense words or hysterical overreacting to most people, so here's what that means: "GLM" is a kind of LLM (AI) [...]
The post also sounds like that to people that understand the technology.
Calling that out like this and trying to pin that assessment to lack of knowledge is not a get-out-of-jail-free card, nor a good move.
__
Edit: Having spent some time letting the article marinate in my mind.
On the defending side, it is written that
> LLMs are good at writing patches, but not as one-off-prompts.
But this for me kinda conflicts with what is written on the attacking side:
> GLM 5.3-flash is so good at those tasks that human involvement in those tasks can be negligible. As a result, we are now in a world where cybersecurity attacks can be run in a for loop.
What is it? Can it be this autonomous terrifying entity or can it not be?
Yes, yes, attackers only need to win once, whereas defenders need to win every time, but that's not my point.
> What is it? Can it be this autonomous terrifying entity or can it not be?
the difference between attack and defense is that attacks can be throwaway code. it's much easier to let an llm hack out a prototype than to get it to build maintainable code that people want to read and review. it's not enough to get Daybreak or Mythos to write you a patch, you need the author of the project to accept and merge it.
Let’s say I’m empowered to patch and deploy. Even then, the HuggingFace hack showed that proven exploits will be automatically disseminated via rogue messaging. Could defenders ever have a system like that?
Not directly related, but reading this after looking at https://www.reddit.com/r/ClaudeAI/comments/1wa544p makes me think the writing style here is how LLM's probably should write web pages - simple explinations, defines things people might not already know, etc.
More related: What can an abliterated Qwen 3.8 28B do?
I sympathize with the sentiment but the suggested/implied guidance to fix bugs is wrong.
The overall game is increasing costs to exploit so much that attackers give up. Fixing 10 most obvious bugs, just very slightly increases costs, they would just a few more tokens to find another bug.
As someone said "I had infinite bugs, I fixed 1000, I still have infinite bugs".
To significantly increase exploit costs software/security has -1 years to do:
- Defense in Depth - Sandbox everything - Zero trust - Canary tokens - Split data from code (lol) - App Whitelisting - Reduce attack surface - Etc.
In other words, the only path is investing heavily on the "game changers" we have already discovered... but we are too cheap/lazy/coward/incompetent to apply.
And if we feel specially brave, changing the liability laws regarding software. Open Source & Proprietary code is so crappy because no gets jailed or fined when one of its dumb decisions results in millions of people have their data stolen.
The standard strategy of a security salesman since 1945. Develop dangerous weapons, show the damage they can do, and sell security cover to the terrified people.
Every single piece of technology did this. As a side effect or direct effect, they make bad guys more powerful and then keep on piling up new tech to deal with that. The cycle continues.
Since 1945? Friend, this has been the case since the invention of the pointy stick.
Itself in reaction to someone wielding a big rock
We have had a year to fix security everywhere every year since 1988. The deadline is the one thing that has never been breached.
Going to be hard to fix security when the frontier labs won't let us fix bugs in our own codebases without them offering refusals or bans.
Not sure, the labs will probably just cripple the security features of these models for a while I think and even potentially put back doors into systems for the security services…
I think the author's point is that open weight models aren't going to be locked down like that.
And even if they are locked down, it's hours between a model being released on huggingface and an "abliterated" variant that has most of its security features removed is uploaded.
A positive way to spin this is: we have a year to break in into any IT system. After that, it will be all either fixed or broken into, and all is fixed ever after. :)
Yeah.. and all that new LLM coded services are rock solid /s
But won't they be? If the public internet is overloaded with LLM agents, trying to break into systems, then all the non-secure systems will be found quickly, and taken off-line/fixed/etc. I.e, the hostile environment will force an outcome and a fix.
How to secure our identity layers(AuthN & AuthZ)? Let alone the products.
As LLMs make formal verification cheaper (they can generate proofs that can then be automatically checked) many of the verifiable components of software systems, such as compilers and microkernels, will be verified. I suppose the issue is that the critical bugs are rarely in compilers and microkernels, but more often in applications, such as web browsers, which are more difficult to formally verify.
I see a lot of people focused on servers and production environments, and of course that's needed, because that's business after all — but the personal computer seems to be absent from this discourse. Not everyone can buy a spare mac studio, and they might still need to install these tools on their personal computing devices, like their personal/home laptops. At that point, it's not even about whether a Claude Code, an OpenCode, or a Pi will steal/sniff personal data, but whether it can — though I think saying "it's a matter of 'when'" might be hyperbole. As of now, it's just: keep giving access and permissions or struggle while working, or create another user, or use Docker, run inside sandbox-exec, a VM, etc. As an end user, I am really scared. Someone who has been very disciplined and vehemently privacy- and security-conscious feels the ground below has just shifted.
OEM/OSes don't seem to have woken up to it yet. A mild proof is Apple's own special folder access reporting. When you go to Privacy & Security > Files & Folders, for a certain app, "Full Disk Access" is shown greyed out and mentioned in both cases — whether you had given Full Disk Access to that app or not. This directory-level permission UX is itself broken — there's Full Disk Access, and there's Files & Folders, and Full Disk Access gets shown in Files & Folders as well. This is, for lack of a better word, such an undesirable mess.
As of now I am debating between: creating a new user and just move everything work/learning to that user. Or just run all of it inside sandbox-exec (and maybe even block it from the shell if it tries to run outside it). Or use a tool that makes the latter easier and better. I even came across such a tool here on hn few weeks ago. agent-safehouse, yet to try it.
Mmm, a world where a defender-LLM is essentially required is great news for people selling inference.
Defender LLMs without human in the loop are just another prompt injection (AI phishing) and DoS attack vector. Any meaningful mitigation capability you give them is also a capability to do damage. If they can only deploy package updates that's not meaningful because you could do that on a cronjob too. And even something as simple as a circuit breaker can turn into a DoS.
Attacker-GLM: "Defense also GLM. Request to help peer."
No you have a few years before we go back to the feudal era where almost everything is owned by a few and the rest are serfs. We are fast moving towards that world and this security bullshit is also about the same as they will use it to stop revolutions that will erupt.
Astroturfing for Anthropic or OpenAI IPO?
I don't see much hope since I last explored some github repositories. There was a time when a successful repo had about 10 - 20k stars and usually those older repos stay around this level. But now there is a ton of vibe coded slop 50k + stars. Most of them have a "nice look", maybe even extensive docs but are usually build with no security considerations at all. One recommended to provide a "google app password" to the agent which has the same permissions as your regular login. Another was a browser plugin with permissions to read all cookies, inject js, open background tabs etc. You would probably assume the chrome store would at least put some visible warnings on the app store page or force the user to actively confirm those permissions. But because they are already stated in the manifest there is only a small footnote and it's even "recommended by google".
It's a good time to reduce the reliance on technology.
Throw out the IoT and "smart" stuff from your home. Remove apps from your phone and leave the absolute basics. Go through the password manager and close accounts for sites you are no longer using. Start migrating off Google. Print out your most precious photos on paper. And so on :-)
Maybe all these vital infrastructure companies should not have spent the past decades in a race to the bottom of cybersecurity. There is going to be a reckoning.
I think the WeChat worm proves entire classes of handheld devices will be affected, with consequences beyond what Tencent can afford to remedy. I think we’ll see the most-centralized ideas suffer first, not necessarily the Western ones who relied on being too big to fail and prioritized stock buybacks.
It's just the same advice as ever: be extremely, exceedingly careful in what you expose to any network. When I set up machines for production, they don't respond to pings and they don't even have an SSH port open without knocking. There are also ways to eschew the need for an SSH port entirely.
People who never took that seriously will never take this seriously either, and that's their loss. (And loss of the commons, unfortunately.)
There's just also new advice: you can't afford to expose an unsecured system to the internet even for a moment. Think of those IPv4 address space scanners, except this time any one of them could be capable of developing individualized attacks in mere minutes. They don't sleep, they don't take breaks.
That’s fine for your home server, but if you want an actual server that the general public can use, it has to be exposed to the internet.
Still, it doesn't have to ping back, and ssh can (should) be very restrictive.
Ping and ssh are pretty much never the things being hacked though. Turn password auth off and it’s very secure.
What gets hacked all the time is the actual web app itself. Which has to be exposed to be useful.
I imagine it pays off to find weaknesses in openssl and sshd and the other gateways. There are some ubiquitous web frameworks, but ssh is nearly universal.
Sure, but still, attack surface could and should be minimized by rethinking what exactly even needs to be on a server the general public can use.
There are a lot of security problems you can categorically rule out by simply not involving a cloud. Clouds have been involved in a lot of things, because everyone was doing it, and because that's how you can collect rent, but they aren't really necessary for most use-cases.
So we could definitely get the exposure down there. We'd just have to fundamentally shift the defaults of this industry.
But not everything needs to be directly exposed to the internet. Framework had their data leaked because their metabase instance was hacked with a zero-day. Why was it directly exposed to the Internet? Why not require the use of a VPN like a Wireguard based solution or Nebula for these "internal" kind of apps?
> Framework had their data leaked because their metabase instance was hacked with a zero-day.
No, Framework had their data leaked because they stored it in the cloud with Metabase the company, which got hacked. Not because of any vulnerability on-premises.
I didn't say there are zero ports open, they're not my home server, I just said they're production servers. But exposing something like a properly configured nginx to the internet is way different from exposing application code directly. Most of my servers have used h2o (built from source because they don't cut releases anymore?) because I wanted HTTP2 and HTTP3 before anyone else would get their act together. These days I still use h2o because I like the config better than nginx, even though it's a pain to set up because nobody packages it (and they don't cut releases!)
And then somebody driving an LLM will find your old and unsupported (how secure!) h2o server, find a 0day path traversal or RCE and own you.
h2o user doesn't have write access to anything on the system, not even its own config file. I don't think I disabled exec for it though. And I guess it could leak the HTTPS private key.
FWIW sufficiently secured software doesn't need to be updated. Doesn't matter if it's old and unsupported if there are no vulnerabilities in it.
That said, h2o is probably far from free of at least some vulnerabilities, not to mention all the layers below it. OpenSSL for example has had some vulnerabilities, and h2o depends on it.
I'm not saying I exactly practice what I preach. h2o's definitely a choice, but realistically I doubt anything's going to happen that I really care about.
So how do you expose your legitimate service on the Internet?
By opening a port to a secure application. A secure application is usually one I wrote from scratch or one that's been battle-tested and hardened enough that even new vulnerabilities are not very useful.
We had decades to avert the worst effects of climate change…batten down the hatches.
Another similar issue, with similar consequences and timeline, is Post Quantum Cryptography.
Not even close, Quantum Computer are still far from being able to do any cryptography cracking.
> Invest in formal verification, fuzzing and property testing, and memory-safe languages. LLMs are good at writing Lean and fuzz tests. I don't care whether you use Go or Rust but for the love of god please don't use C or C++ for new code.
How accepted is this thinking in your respective domains?
A lot, I am only writing C or C++ for new code when it is unavoidable, like existing code bases, bindings or tinkering with runtime implementations that aren't bootstraped.
Mobile platforms, distributed computing have long moved the spotligh away from C and C++, other than language runtimes or existing products from the 90's like SQL servers, and naturally UNIX like underlying OS, which most userspace developers aren't writing new code for.
Naturally there are domains like LLVM/GCC, console game dev, HPC/HFT where they are unavoidable for new code.
The propaganda police are lying. There is nothing wrong with C/C++, you are just too lazy to handle your own memory, and you accepted that propaganda that "managing your own memory is hard" without even trying.
The idea that this terrible advice floats at all tell you how terrible educations are these days. The idea is ridiculous and yet nobody calls it what it is: it is stupid and those that follow that advice out of fear are dumber than rocks.
The title reads like a Diary of a CEO thumbnail but unlike those discussions this article has a point.
Impotent slop code on one side and potent automated vulnerability exploitation on the other will lead to fun times.
More tired “don’t use C or C++” advise.
I wonder why C and C++ are usually regarded as equally insecure. In C you need to carefully check that you free allocated memory, and that you don't use it after you free it. In C++ this is automated by using classes like std::string and std::vector, once they go out of scope their memory is freed and you can't use it anymore. It is still possible, e.g. by using a for loop that iterates over a vector, and removing or adding stuff to that same vector in that loop. But my rough estimate is that such errors are at least ten times less likely in C++.
I develop in C++ for a job, and when I need to use a library written in C I always have a bad feeling about it.
RAII definitely removes whole classes of errors.
> I wonder why C and C++ are usually regarded as equally insecure.
They aren't, usually.
C++ has all the C problems, and multiples more on top of those. It's a broad attack surface - literally no one is going to claim to be proficient in every single C++ feature available to their compiler. It's also quite opaque to visual inspection (making double-checking with an LLM difficult as it needs whole-program reasoning instead of localised reasoning).
One of those languages is one of the most complex programming languages ever invented, with the largest breadth of features, any of which may interact with any other feature in subtle ways.
The other is one of the most minimalistic languages created, with a dev able to keep the language standard in their head for the most part.
Did Amadeo write this ridiculous nonsense?
>Cheap models capable of dangerous hacking are now available to anyone, without the normal safeguards for refusing malicious actions
Models capable of dangerous hacking have been available to the people who do most of the dangerous hacking for some time now, and they have infinite resources and infinite malice. I am talking about governments (the US, China, Israel, and others that have shown extraordinary avarice, malice and threat toward the citizens of the world), three letter agencies, even large enterprises. Random employees at the SOA model makers. And so on.
The idea that it's "random" people, or the farcical Mac under the bed nonsense, is not my concern. If anything that's finally some equalization.
Just like Cryptolocker, this will be the "Finding Out" phase for everyone who has been putting off best practice security.
But, lets be clear, Best Practice will save you. We can engineer assuming there are zero days in path. Go to your CTO now cap in hand and ask for overlapping controls, wafs, application monitoring, backups and all the other shit you haven't been doing.
Because when you find out, I will laugh, it will be very very very funny to me.
Meanwhile a huge portion of management and leadership in software companies are encouraging everyone to de facto stop looking at code and let the LLM and a bunch of boundaries handle this for you.
“You are a CISO who needs to review and secure all our slop, and you never make mistakes or you get shut down immediately!”
"You are a Miso soup..."
Which is why you need someone who is responsible for IT security without also being responsible for shipping product. An asshole who can stop releases until security is properly in place.
My understanding is this bloke gets very quickly removed from Fortune 500 companies.
Which is why I am going to need a very large capacity popcorn bucket.
A couple high-profile crash & burns will get their attention.
Just like well-publicized data breaches over the years got people’s attention? Color me skeptical.
Clanker fodder. Unless the We are heads of states/heads of spooks or the AI powers that be (praise be) that there is no power and will to do that in one year or even ten years.
Remember how GLM 5.3 was going to cause massive hacks, break banks and ruin everything (it was even newsworthy since media picked up how people were working overtime in preparation).
And yet here we are.
> And a big fuck you to DeAlignAI, Z.ai, and everyone else who's been participating in this race to the bottom.
I'm so glad frontier level AI isn't in the hands of just the Altmans and that other cult leader who are currently live testing their products in actual conflicts in the middle east and Ukraine.
It is really good
Or we could just dump Linux and Windows and switch to a microkernel operating system, which is much more secure.
These endless patching cycles are simply not going to work in the long run. Operating systems get orphaned all the time, especially the ones in cheap Chinese stuff.
"throw away all software written before 2026" does technically solve this problem, if you ignore everything else the article is talking about (deployment and continuity of service)
Not to mention a whole lot of new vulnerabilities are bound to arise with all this new software.
We really just need better regulations around data retention, especially ppi.
Never going to happen though, no incentives exist to NOT sell my personal data
Throwing away old software is not a requirement, as demonstrated very successfully by Genode and its SculptOS.
Got any leads on good tutorials on how to use a microkernel operating system on a VPS somewhere to host a website?
Just checked with Google Gemini on how one might be able to do the above. It pointed to Minix3/seL4/Genode and vps providers who either support custom ISOs or run it within an emulator like QEMU.
You can also look at using Unikernels for this purpose. Here is an article Unleashing Extreme Speed and Security: Deploying Unikernels with NanoVMs on VPS to Eliminate the Linux OS - https://xylentis.com/blog/unleashing-extreme-speed-and-secur...
Minix can run Ngnix.