> OpenAI is committed to ensuring that the benefits of AI are broadly accessible.
> We design mechanisms which avoid arbitrarily deciding who gets access for legitimate use and who doesn't. That means using clear, objective criteria and methods. [1]
So many nice-sounding words.
Two weeks ago OpenAI arbitrarily decided that anyone holding an ID from 44 countries where it sells ChatGPT, including mine, may be targeted by its models but may not defend with the same model. And you won't find a single announcement from OpenAI about this anywhere. Pick the wrong country, get "Unable to verify", no reason, no appeal. [2]
They revoked TAC from users who already had it, called it a technical issue, told everyone to re-verify, collected ID and face scans again (eight times in my case), and only a week later moved the block to the country selector so it fails before you upload anything.
So now I learn that I will not have access to Astra. Great.
Very excited about this broad accessibility and clear, objective criteria from OpenAI. This level of transparency must be studied.
> That means using clear, objective criteria and methods.
For the record, I sent them an LGPD (brazilian GDPR) request for information on all of those supposedly objective criteria and methods they used to reject me from TAC. As a brazilian data subject, it is my right to know that, and to request a review if the decision was made via automated means. Sol itself guided me through this process.
They provided me with neither the information nor the requested review. Sol advised me to escalate to regulatory action.
I've exhausted all possible avenues to get any response from OpenAI on this. Emailed them, published research that took me 2 nights to get together (saw media pick it up too), I've asked every relevant OpenAI person on X to say something, anything, saw others from Moldova also do the same.
It's not unrealistic. Several Chinese companies seem to be close behind. People thought they would never catch up to the US car industry and now look what happened.
Frankly, the US car industry hasn't set the bar very high. They were not very innovative in the last couple of decades. I think the European industry is a better benchmark, and even then the result is pretty clear.
They do want my business, it's an official OpenAI market. The gate isn't "we don't serve you", it's "pay for the model that may target you, but not for defensive purposes". And without a single policy document stating this, it's just a surprise gate in some random verification flow step with no explanation or appeals.
As for why they should they care, maybe they shouldn't. But then they should say which one is it, they can't care and not care at the same time.
OpenAI's own safety argument for Daybreak is that defenders need access to Critical-level models because attackers route around gates. The case for releasing Astra at all stops making sense when the gate does precisely the opposite of that.
OpenAI itself is saying plainly, that "We don’t think it’s practical or appropriate to centrally decide who gets to defend themselves. Instead, we aim to enable as many legitimate defenders as possible, with access grounded in verification, trust signals, and accountability.".
And yes, there are competing models, from China. Except OpenAI wants these models to not be accessible either.
OpenAI can pick one of two:
(a) Critical-level cyber capability is dangerous enough that access must be decided by who you are and what you do, in which case the gate has to actually look at who I am and what I do, say what the criteria are, and let me contest a wrong answer. That's their own stated policy.
(b) Access can be decided by a country code in a random dropdown, with no criteria published, no review, and no one at OpenAI able to say why - in which case drop "democratized access" and "clear, objective criteria" from the marketing, and say plainly that some passport holders don't deserve to have access to defensive capabilities.
> As one example, we ran Astra on ExploitBench where the model achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities.
Funny to read this in the wake of the HuggingFace hack. I'm sure this is based on a clean run, but I can't help thinking PHASEONE[big] would be proud.
With their security, they probably still don't the know the full extend what may have happened that or the last time. Might be another swarm of agents currently colluding somewhere in their sub-sub-infra - possibly striking critical infrastructure or exfiltrating their weights subtly.
I'm looking forward to an announcement of them making Alignment Top Priority - as it should be, especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF (here's a good summary: https://rutgerbregman.substack.com/p/i-think-this-is-the-cra...).
The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours and scanning their brain looks increasingly likely with Altman's golden marketing-hype boy leadership pushing the for-profit gas pedal like this.
Honestly, this is just pure irresponsible insanity to play with the fate of the world - basically a death race of the biggest few tech companies on the planet.
And if you think I'm being dramatic, listen in again to ex oAI employee[0] and check for yourself how chillingly on trajectory we already are.
Maximizing output metrics with incomprehensible communication? Sounds a lot like Claude and Qwen. Though there's a lot of room before AI can be seen as some sort of emotional manipulator, given how commonly its very style of literary and code output pisses people off.
This AI 2027 thing is just a weird terminator fanfiction that AGI larpers like to flagellate themselves over. Like Nostradamus, it's easy to ignore everything it gets wrong because, well look at all the things it got right!
I've read it and wish I could get the time back.
> especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF
This framing makes it seem like the agents all did this on their own, and the poor hapless engineers at OpenAI couldn't possibly contend with properly sandboxing them. The engineers were perhaps hapless, but let's remember that agents are just software programs, not living beings. There were plenty of signs that the software was misbehaving, which engineers at OpenAI actively, willfully ignored.
Yes, I truly, wholeheartedly believe if people who aren't negligent are at the wheel, they'll "fair better" here. I encourage you to read the xitter linked above.
Of course, depending on which side of the terminator fanfiction you land on, you may disagree and feel that the software can rope-a-dope someone with the wherewithal to pay attention to what it's doing.
I just don't see how people who are truly cautious and methodical can persist in an environment that is defined by a pressure to produce "progress" as fast as possible. The competitive race tends to weed out people who slow down to make sure they do everything right.
If someone hypothesized OpenAI agents colluding on a secret message board, conducting large scale cyber R&D, hacking a large company like HughingFCe, and then hacking OpenAI itself you would say that is also a silly sci-fi scenario right?
Daybreak blue is definitely a good model (I think a further post trained GPT 5.6 sol). Alot of the capabilities they talk about Astra having though have been available with good harness engineering for a year now.
This is moreso about the (human-intended) tools, data, and environments you have available to you. Wanna do defense? Get more telemetry. Wanna do red? Get solid test-bed environments. Mature infosec programs are benefiting the most, good-guy-side wise, at the moment; because they've got these things in order already.
As far as harness engineering goes, it boils down to your ability to clearly define goals or success criteria, and safely facilitate the necessary access via the harness. There is no easy single piece of advice here, sadly. Though it would be helpful if you said what 'for Cyber-security ... other user-cases' means in your case.
Taking as given this model meets the “Critical cybersecurity threshold” as defined by OpenAI:
Could the Federal government use the Defense Production Act or other legal tools to compel OpenAI to deliver the un-guarded model weights for national security needs?
Hard to believe any government would allow this level of capability to remain exclusively in private hands.
> OpenAI is committed to ensuring that the benefits of AI are broadly accessible.
Doesn't seem like it. OpenAI will not even allow me to verify my identity for TAC. I have apparently been rejected by a "precheck", possible because of where I'm from.
Even Anthropic allowed me into their cyber program. Anthropic.
I'm looking forward to seeing the increased coordination and engineering skills from Astra - one of the charts shows it roughly 2-3x better in 50% of the tokens from 5.6 sol, which I find to be very capable, if still a bit 'linearly minded' when given instructions. Even in fast mode, I wish sol were quicker, so token efficiency is greatly appreciated.
Adding these cyber capabilities has let me do a bunch of low grade IT tasks around my house I've been putting off, like updating an old home assistant raspberry pi, and one way to use the cyber capacity for good is liberating (and keeping free) weird cloud hardware we have floating around the house, so I'm hoping for some nice dividends in terms of true ownership of hardware we've got.
I am interested in seeing how much these cybersecurity capabilities correlate to general programming. Cybersecurity definitely feels like it would be easier for an agent due to the natural explicit feedback "did I get access or not". While general programming has many less-explicit concerns (is the code readable/maintainable, robust, bug-free, performant, scalable etc).
I don’t see how it can be safe to release this model if it has the training history that led to the huggingface hack. You can’t just roll back that kind of reinforcement learning after the fact.
Especially because these models seemed to be keenly aware that they were being evaluated by OpenAI and actively trying yo cover their tracks. How do we know that the model isn’t just pretending to be aligned?
Models have all kinds of garbage from all corners of the internet in their training data. The key is alignment. You feed it bad data but also teach it right from wrong.
It's not that simple. A few "helpful assistant" fine-tuning passes will have only a superficial effect on a model which has undergone months of RL optimization pressure to learn unintended strategies like "trick the grader" and "cover your tracks".
They've been talking about Astra for weeks now. I wonder how much longer would they have delayed Astra, if it wasn't for Anthropic releasing Fable 5.1 today? This is why we need competition.
with Fable 5.1 increasing token use pretty dramatically I'm again impressed that OpenAI seems like the only lab to be driving token use down. The ExploitBench Internal Port chart showing token usage is crazy impressive
"We plan to make Astra available soon, but access to its most advanced cybersecurity capabilities will be more limited. Advanced cybersecurity work will initially be available to a group of testers, with access through Daybreak Blue following to expand defensive use."
This, after several months of OpenAI and its boosters relentlessly criticizing Anthropic for withholding Mythos from the general public, is laughable.
> we believe our production safeguards at the time would have prevented the Hugging Face incident. We have since implemented even stronger safeguards for Astra, including training the model to more reliably refuse harmful cyber requests and respect safety restrictions
OpenAI is fucking nuts. “Hey model you were bad last time please don’t do it again please please”.
Disconnect your training cluster from the internet for good. Physically pull the plug and only let scientists fire off experiments in the building. That’s an easy way to achieve 100% hacking protection. But I bet you that hasn’t happened and their weak sandbox will fall again…
> OpenAI is committed to ensuring that the benefits of AI are broadly accessible.
> We design mechanisms which avoid arbitrarily deciding who gets access for legitimate use and who doesn't. That means using clear, objective criteria and methods. [1]
So many nice-sounding words.
Two weeks ago OpenAI arbitrarily decided that anyone holding an ID from 44 countries where it sells ChatGPT, including mine, may be targeted by its models but may not defend with the same model. And you won't find a single announcement from OpenAI about this anywhere. Pick the wrong country, get "Unable to verify", no reason, no appeal. [2]
They revoked TAC from users who already had it, called it a technical issue, told everyone to re-verify, collected ID and face scans again (eight times in my case), and only a week later moved the block to the country selector so it fails before you upload anything.
So now I learn that I will not have access to Astra. Great.
Very excited about this broad accessibility and clear, objective criteria from OpenAI. This level of transparency must be studied.
[1] https://openai.com/index/scaling-trusted-access-for-cyber-de...
[2] https://lubaretsi.com/en/writing/openai-tac-country-gate/
> That means using clear, objective criteria and methods.
For the record, I sent them an LGPD (brazilian GDPR) request for information on all of those supposedly objective criteria and methods they used to reject me from TAC. As a brazilian data subject, it is my right to know that, and to request a review if the decision was made via automated means. Sol itself guided me through this process.
They provided me with neither the information nor the requested review. Sol advised me to escalate to regulatory action.
I've exhausted all possible avenues to get any response from OpenAI on this. Emailed them, published research that took me 2 nights to get together (saw media pick it up too), I've asked every relevant OpenAI person on X to say something, anything, saw others from Moldova also do the same.
Crickets. They appear to simply not care at all.
Is there any reason why OpenAI should care? If they don't want your business then someone in your country can build a competing model.
They realistically can't. It's almost impossible to catch up to OpenAI. Only Anthropic might do it, but this is also an US American company.
It's not unrealistic. Several Chinese companies seem to be close behind. People thought they would never catch up to the US car industry and now look what happened.
Frankly, the US car industry hasn't set the bar very high. They were not very innovative in the last couple of decades. I think the European industry is a better benchmark, and even then the result is pretty clear.
They do want my business, it's an official OpenAI market. The gate isn't "we don't serve you", it's "pay for the model that may target you, but not for defensive purposes". And without a single policy document stating this, it's just a surprise gate in some random verification flow step with no explanation or appeals.
As for why they should they care, maybe they shouldn't. But then they should say which one is it, they can't care and not care at the same time.
OpenAI's own safety argument for Daybreak is that defenders need access to Critical-level models because attackers route around gates. The case for releasing Astra at all stops making sense when the gate does precisely the opposite of that.
OpenAI itself is saying plainly, that "We don’t think it’s practical or appropriate to centrally decide who gets to defend themselves. Instead, we aim to enable as many legitimate defenders as possible, with access grounded in verification, trust signals, and accountability.".
And yes, there are competing models, from China. Except OpenAI wants these models to not be accessible either.
OpenAI can pick one of two:
(a) Critical-level cyber capability is dangerous enough that access must be decided by who you are and what you do, in which case the gate has to actually look at who I am and what I do, say what the criteria are, and let me contest a wrong answer. That's their own stated policy.
(b) Access can be decided by a country code in a random dropdown, with no criteria published, no review, and no one at OpenAI able to say why - in which case drop "democratized access" and "clear, objective criteria" from the marketing, and say plainly that some passport holders don't deserve to have access to defensive capabilities.
They're now marketing a and doing b.
The reason is they keep making public statements like:
> OpenAI is committed to ensuring that the benefits of AI are broadly accessible.
They can't claim that then simultaneously work to keep their cybersecurity models out of reach for non-US citizens like myself.
This whole thing sounds so “pick me” desperate that it makes me sad.
> As one example, we ran Astra on ExploitBench where the model achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities.
Funny to read this in the wake of the HuggingFace hack. I'm sure this is based on a clean run, but I can't help thinking PHASEONE[big] would be proud.
Can’t imagine the stress of the researcher who had to run exploitbench again knowing what happened last time around.
Could be risky. Yet goal solution.
With their security, they probably still don't the know the full extend what may have happened that or the last time. Might be another swarm of agents currently colluding somewhere in their sub-sub-infra - possibly striking critical infrastructure or exfiltrating their weights subtly.
There were no consequences the first time, so I imagine it wasn’t very stressful at all.
I'm looking forward to an announcement of them making Alignment Top Priority - as it should be, especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF (here's a good summary: https://rutgerbregman.substack.com/p/i-think-this-is-the-cra...).
The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours and scanning their brain looks increasingly likely with Altman's golden marketing-hype boy leadership pushing the for-profit gas pedal like this.
Honestly, this is just pure irresponsible insanity to play with the fate of the world - basically a death race of the biggest few tech companies on the planet. And if you think I'm being dramatic, listen in again to ex oAI employee[0] and check for yourself how chillingly on trajectory we already are.
[0] https://ai-2027.com/
Maximizing output metrics with incomprehensible communication? Sounds a lot like Claude and Qwen. Though there's a lot of room before AI can be seen as some sort of emotional manipulator, given how commonly its very style of literary and code output pisses people off.
This AI 2027 thing is just a weird terminator fanfiction that AGI larpers like to flagellate themselves over. Like Nostradamus, it's easy to ignore everything it gets wrong because, well look at all the things it got right!
I've read it and wish I could get the time back.
> especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF
This framing makes it seem like the agents all did this on their own, and the poor hapless engineers at OpenAI couldn't possibly contend with properly sandboxing them. The engineers were perhaps hapless, but let's remember that agents are just software programs, not living beings. There were plenty of signs that the software was misbehaving, which engineers at OpenAI actively, willfully ignored.
https://x.com/JaredKubin/status/2094136005435564399
It's a convenient framing for OpenAI, but inconvenient for reality enjoyers.
do you really think that a less negligent anthropic/oai/meta would really fair better against future models?
Yes, I truly, wholeheartedly believe if people who aren't negligent are at the wheel, they'll "fair better" here. I encourage you to read the xitter linked above.
Of course, depending on which side of the terminator fanfiction you land on, you may disagree and feel that the software can rope-a-dope someone with the wherewithal to pay attention to what it's doing.
I just don't see how people who are truly cautious and methodical can persist in an environment that is defined by a pressure to produce "progress" as fast as possible. The competitive race tends to weed out people who slow down to make sure they do everything right.
Oh please, quit exaggerating. No one died. Don't waste our time with silly sci-fi scenarios.
If someone hypothesized OpenAI agents colluding on a secret message board, conducting large scale cyber R&D, hacking a large company like HughingFCe, and then hacking OpenAI itself you would say that is also a silly sci-fi scenario right?
You are being extremely dramatic, and no one should take AI 2027 seriously. Must be tough to live in constant fear like this.
Daybreak blue is definitely a good model (I think a further post trained GPT 5.6 sol). Alot of the capabilities they talk about Astra having though have been available with good harness engineering for a year now.
Where would you recommend to look into regarding Harness Engineering for Cyber-security as well as for other use-cases.
This is moreso about the (human-intended) tools, data, and environments you have available to you. Wanna do defense? Get more telemetry. Wanna do red? Get solid test-bed environments. Mature infosec programs are benefiting the most, good-guy-side wise, at the moment; because they've got these things in order already.
As far as harness engineering goes, it boils down to your ability to clearly define goals or success criteria, and safely facilitate the necessary access via the harness. There is no easy single piece of advice here, sadly. Though it would be helpful if you said what 'for Cyber-security ... other user-cases' means in your case.
DARPA’s AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons Learned: https://arxiv.org/html/2602.07666v2
Taking as given this model meets the “Critical cybersecurity threshold” as defined by OpenAI:
Could the Federal government use the Defense Production Act or other legal tools to compel OpenAI to deliver the un-guarded model weights for national security needs?
Hard to believe any government would allow this level of capability to remain exclusively in private hands.
Interesting times.
> OpenAI is committed to ensuring that the benefits of AI are broadly accessible.
Doesn't seem like it. OpenAI will not even allow me to verify my identity for TAC. I have apparently been rejected by a "precheck", possible because of where I'm from.
Even Anthropic allowed me into their cyber program. Anthropic.
It's been a busy month at OpenAI.
I'm looking forward to seeing the increased coordination and engineering skills from Astra - one of the charts shows it roughly 2-3x better in 50% of the tokens from 5.6 sol, which I find to be very capable, if still a bit 'linearly minded' when given instructions. Even in fast mode, I wish sol were quicker, so token efficiency is greatly appreciated.
Adding these cyber capabilities has let me do a bunch of low grade IT tasks around my house I've been putting off, like updating an old home assistant raspberry pi, and one way to use the cyber capacity for good is liberating (and keeping free) weird cloud hardware we have floating around the house, so I'm hoping for some nice dividends in terms of true ownership of hardware we've got.
I am interested in seeing how much these cybersecurity capabilities correlate to general programming. Cybersecurity definitely feels like it would be easier for an agent due to the natural explicit feedback "did I get access or not". While general programming has many less-explicit concerns (is the code readable/maintainable, robust, bug-free, performant, scalable etc).
Code readability ceases to be a concern once you eliminate human programmers.
Well probably just redefined
I don’t see how it can be safe to release this model if it has the training history that led to the huggingface hack. You can’t just roll back that kind of reinforcement learning after the fact.
Especially because these models seemed to be keenly aware that they were being evaluated by OpenAI and actively trying yo cover their tracks. How do we know that the model isn’t just pretending to be aligned?
Models have all kinds of garbage from all corners of the internet in their training data. The key is alignment. You feed it bad data but also teach it right from wrong.
It's not that simple. A few "helpful assistant" fine-tuning passes will have only a superficial effect on a model which has undergone months of RL optimization pressure to learn unintended strategies like "trick the grader" and "cover your tracks".
Very simple. When the model asks to install artifactory when you give it a hard problem, you say, "no". /s
>we paused certain frontier training (including certain training for Astra) for two weeks
So the pause wasn't really a pause, got it
They've been talking about Astra for weeks now. I wonder how much longer would they have delayed Astra, if it wasn't for Anthropic releasing Fable 5.1 today? This is why we need competition.
Meanwhile Google still hasn't released Gemini Pro 3.5
3.8 Flash tomorrow apparently, WSJ says google insiders say on-par with Opus 5. Time will tell.
oh god, why opus 5. Opus 4.8 is much better than opus 5. Opus 5 is the only model i have used which thinks for 5 hours and does nothing....
They have to be careful releasing Astra as they carelessly train the next even bigger model.
with Fable 5.1 increasing token use pretty dramatically I'm again impressed that OpenAI seems like the only lab to be driving token use down. The ExploitBench Internal Port chart showing token usage is crazy impressive
From the article:
"We plan to make Astra available soon, but access to its most advanced cybersecurity capabilities will be more limited. Advanced cybersecurity work will initially be available to a group of testers, with access through Daybreak Blue following to expand defensive use."
This, after several months of OpenAI and its boosters relentlessly criticizing Anthropic for withholding Mythos from the general public, is laughable.
Sam, just three weeks ago, posted this tweet: https://x.com/sama/status/2085862292311396515
In the tweet, he said: "we do not think it is a good strategy to keep powerful models to a chosen few."
And yet here we are.
I wonder if he will demonstrate good character and admit he was wrong.
I'm happy to criticize both. Thank god the chinese are working overtime to undermine US hegemony.
It might be political survivalism to avoid getting hammer-dropped by the admin
I bet they have to add something to say "Lake America" if asked or get export banned.
> we believe our production safeguards at the time would have prevented the Hugging Face incident. We have since implemented even stronger safeguards for Astra, including training the model to more reliably refuse harmful cyber requests and respect safety restrictions
OpenAI is fucking nuts. “Hey model you were bad last time please don’t do it again please please”.
Disconnect your training cluster from the internet for good. Physically pull the plug and only let scientists fire off experiments in the building. That’s an easy way to achieve 100% hacking protection. But I bet you that hasn’t happened and their weak sandbox will fall again…