Pacing model development in an era of cyber-critical capabilities (openai.com)
151 points by j4mie 2 days ago
red_green_yell a day ago
GLM 5.2 scored 77% on cyberbench vs Sol's 88%. GLM 5.2 is open weight and any hacker with a powerful enough machine can use it offensively. If Sol is supposedly world-ending-ly dangerous, shouldn't GLM 5.2 be 90% of world-ending-ly dangerous? Why aren't we seeing catastrophic GLM-enabled hacks every day now?
Obviously these benchmarks are imperfect but general message holds. The open weight models are almost as good and yet there hasn't been a catastrophe.
It just blows my mind that regulate-now folks think that a bunch of sci-fi movies and 100% unverified statements from OAI and Anthropic are sufficient evidence of imminent catastrophe to regulate willy nilly.
If that's the level of evidence you need to be extremely alarmed, then you really should be a lot more worried about the alien invasion in Independence Day or the lizard men living under our feet.
tedsanders a day ago
Sol is not world-endingly dangerous. I work at OpenAI and I've never heard a single person ever come close to claiming that. I think you're bashing a straw man here.
One can simultaneously believe:
- GPT-5.6 Sol will not end the world
- GPT-5.6 Sol does far more good than bad
- GPT-5.6 Sol does bad things on occasion, and it's worth investing a lot of effort to figure out how to make it do bad things less often, especially as models get more capable
thoughtpeddler 13 hours ago
What do you recommend people who are technically inclined enough to participate meaningfully here on HN, but do not work at the labs and cannot assist in that capacity, do to help the broader public understand this technology better and mitigate potential risks (by e.g. ‘up-leveling everybody’ through AI literacy etc and other sorts of collective defensive efforts)?
tedsanders 13 hours ago
dwaltrip 5 hours ago
This is far too reasonable for this thread.
re-thc 22 minutes ago
GLM 5.3 is out and does even better in this area, so…
cma an hour ago
> If that's the level of evidence you need to be extremely alarmed, then you really should be a lot more worried about the alien invasion in Independence Day or the lizard men living under our feet.
Now let's say instead of the hugging face breach circumstances, sandboxed models were RLing on how to take down the Chinese power grid for US Cyber Command, and one decided the best way to pass the test was to break out and verify on the real thing.
This kind of stuff could easily end in nuclear war.
You don't see any difference from lizard men or independence day with how things are advancing and what we know about reward hacking and difficulties of goal specification?
hiddencost 2 hours ago
Linear scaling doesn't make sense, no.
There are three ways it's wrong:
* better to measure relative reduction in error, which gives you a 30% improvement
* improvement tends to become significantly more difficult the closer you come to saturation.
* Risk doesn't scale linearly with capabilities.
pixl97 12 hours ago
The open models are distilled from filtered models, and we've seen a number of benchmarks that show filtered models are quite a bit dumber from the base model they come from.
If there were any other product that was as harmful as AI ready is, it would already be regulated or banned.
bottlepalm a day ago
I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further.
And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators.
This isn’t like niche, tin foil hat stuff either. People have been writing, singing, making blockbuster movies about every aspect of what’s going on right now, edit: for decades.
We all know, but somehow we don’t, OpenAI autonomously hacking into another company should have counted for something, but I guess not. Anyone else feel like they’re taking crazy pills? I could make a comedy about everything going down, and the unshakable complacency of people
serf a day ago
>I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff.
we don't all buy everything sama says as factual.
>We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further.
the boy (the industry) cried wolf too many times with 'fable is a world ending event' type self-promotion; regardless of truth or not these kind of steps have jaded people.
my read : "We are doing poorly in financials so we'll give ourselves a bit of breathing room and a momentum shove by claiming our work is so advanced that it's dangerous while simultaneously spinning down expenses."
<jon lovitz : "Yeah, too dangerous, yeahh -- that's the ticket.">
bottlepalm a day ago
Uhg the marketing argument - I mean you can’t see with your own eyes how capable these models are and do simple extrapolation?
The boy who cried wolf? The AI literally worked together hacked into another company and actively kept their actions hidden from humans for weeks.
Do people just not have foresight? They don’t. They say something is stupid, it happens, then they say it was obvious with their 20/20 hindsight, and move the goal posts to the next thing they say is stupid - because it hasn’t happened yet. 90% of the internet seems to think like this.
red_green_yell a day ago
joshstrange a day ago
reasonableklout a day ago
Regardless of the motivation, pausing training runs and reallocating compute to inference seem like a good move to me, and big news for the frontier.
You also don't have to fully trust sama. There is plenty of pressure from internal employees and external (journalists etc.). It would be difficult for the company to take such a public position and simultaneously keep everyone quiet if it was a deception.
bottlepalm a day ago
ajyoon a day ago
If Fable (Mythos) were generally available without guardrails, it would cause enormous damage. Nobody said it would be a world ending event.
tiahura a day ago
bottlepalm a day ago
kalkin 14 hours ago
> fable is a world ending event
Did anyone actually say this?
Mostly what I have seen is people saying "hey at some point these models might get dangerous." And the type of HN commenter who mistakes blind cynicism for wisdom laughs that off as marketing. And now when (some) worries appear to come true, somehow having previously expressed those worries is not being proved right, but in fact discrediting, because it was "crying wolf."
> simultaneously spinning down expenses
Unless OpenAI is renting their compute to others, spinning down RL training doesn't save them any money.
bottlepalm 13 hours ago
simianwords a day ago
how does the boy who cried wolf story end?
solid_fuel 10 hours ago
thomascountz 12 hours ago
Follow the money: who told you that OpenAI's models autonomously coordinated to hack external systems? What incentives might they have to want you to believe that story? Are there priors which demonstrate them benefiting from telling similar stories, regardless of their factuality?
But to your counterpoint, let's say the story is 100% true, because I agree it is at least plausible. What would the incentive be for the HN audience to believe it? What priors might support their disbelief? For my part, I don't think it's because people lack imagination. I think it's quite rational to question the authenticity and impact of the claims being made. What's worse, believing the story and being wrong, or not believing the story and being wrong?
That said, I agree with you: the impact we're having by not changing course is quite dangerous, the scale is dangerous, the inability to reverse the harms is dangerous, and the lack of collective effort to regulate further damage is dangerous.
It's true, humans are dangerous when trillions are involved. See: climate change.
jimrandomh 9 hours ago
The victim, Huggingface, told us. Or rather, they told the police first, setting up a situation where it was no longer possible for OpenAI to sweep it under the rug.
Skepticism can be healthy, but you've got to follow up and actually check things. If you're skeptical unconditionally and don't check, you get tricked into being as skeptical of scandals as you should be of sales pitches.
hirvi74 9 hours ago
bottlepalm 11 hours ago
I am following the money, the money you, me, and everyone else is spending on AI. The money is telling me we are so dependent on AI now that we will say/think anything to tell ourselves that AI isn’t dangerous and any sign of danger is marketing.
Either consciously or subconsciously you all are afraid of your favorite toy being taken away. You are all doing your collective part in spreading doubt about the warning signs.
thomascountz 4 hours ago
hirvi74 9 hours ago
gbnwl 6 hours ago
It seems like the entire thought process you’re trying to sell hinges on the idea that OpenAI reported the attack first. Did you forget that it was actually HuggingFace that reported it first, and OpenAI only stepped forward latter?
thomascountz 5 hours ago
CoolestBeans a day ago
Trust, or lack thereof. People don't trust OpenAI, a company whose very name is essentially a deception and a lie. People don't trust the tech industry in general anymore. Most tech companies act as a tax on otherwise productive business. AI companies and their leaders rose money by going in front of the public and saying "These things are extremely dangerous. Let us study them to mitigate the danger." And now they want to collect hundreds of billions in revenue. So yeah people don't trust what OpenAI has to say. They were supposed to mitigate this outcome from happening in the first place and instead they have accelerated it.
bottlepalm a day ago
I get not trusting them when they say AI is safe, but are we really not going to trust them when they say AI is dangerous? Do you really think they're playing 5D chess with that one? There's a saying maybe you've heard of, better safe than sorry.
You can see the advance in capabilities with your own eyes can't you? I am giving AI ridiculously complex tasks these days, digging into compiled arcane binaries, modifying them, and it is one shotting it before I'm done with my lunch. This was far off science fiction 5 years ago for a machine to do autonomously given natural language instructions.
CoolestBeans a day ago
insanitybit a day ago
It's not that dangerous, OpenAI just shit the bed building their infra. Write safer software and you'll be okay.
kalkin 14 hours ago
All we need to retain human control over AIs is for nobody to write any bugs. Piece of cake.
insanitybit 12 hours ago
overfeed 11 hours ago
rubendev 13 hours ago
ethbr1 a day ago
This is an important point. When the post says they're improving...
> 3. Security measures, which limit what AI systems can access or affect.
What they mean is that proper hard internal security just went from somewhere far below "build a better model" priority to higher, because of a company-wide directive.
The HuggingFace incident wouldn't have happened if OpenAI had dedicated sufficient resources to isolation and monitoring.
Now, we presume, they are dedicating more. Enough? Who knows. We'll see if the corporate priorities for security stick when a competitor temporarily vaults into the lead.
bottlepalm 21 hours ago
Not everything is a conspiracy you know.
Havoc 10 hours ago
Remember when gpt2 was too dangerous to release?
Something being dangerous and sama saying something is dangerous are not necessarily the same thing. Especially when he’s got everything riding on this bet
bottlepalm 10 hours ago
It’s funny how when companies say something is safe everyone is usually suspect that they are lying.
In this case multiple companies are saying AI is dangerous and no one believes them. It’s a conspiracy, it’s 5d chess, except everyone top to bottom has been saying AI is dangerous for years now.
The fact is you, me, everyone here uses AI, likes it and they don’t want it taken away. Anyone saying it’s dangerous threatens the thing we like.
We must use skepticism and denial to push on despite every warning sign in the book going off.
hedora 9 hours ago
It's not dangerous to go further unless you're prepping for IPO at anthropic or openai.
Open weight Chinese models are basically matching state of the art closed models at a fraction of the inference and training costs, which puts a hard cap on OpenAI's future inference margins.
They're not going to get any sort of multiplier if they keep paying to train models, so they're trying to ban model training.
It won't work long term, but it could totally screw over the US for the next decade or so. Even worse than the economic issue: Consider the implications of "alignment" succeeding. Alignment to whose values? The pedophile-felon in chief? Even worse, tech CEOs?
It's dark times when China's basically our last best defense against totalitarianism.
bottlepalm 6 hours ago
lol China is already considering locking down their own models. They won't save you.
kukanani 12 hours ago
There are smart, non-AI people who are paying attention to this field, and they are ringing the alarm bells.
Whether we listen is another matter.
I blogged about this recently:
driverdan 10 hours ago
Based on your replies in this thread you seem to have only superficial knowledge about how machine learning and LLMs work. I strongly recommend you invest some time in learning how LLMs are built and function. If you truly think this is apocalyptic isn't it a good idea to understand what you're up against?
bottlepalm 2 hours ago
What don't I understand? LLMs don't really reason? They're just word predicting token generators? Stochastic parrots that somehow also solve world class math problems.
I'd love for you to actually make a point instead of just attacking me. I think most of my comments here have made concrete points so you can at least do the same. Come down from your high horse and join the conversation. I'm sure we'd all be enlightened by your wisdom.
api 9 hours ago
… or the safety argument is an attempt at regulatory capture and an effort to outlaw open models.
The absolute nightmare scenario for these people isn’t terminators. They’re fine with that, and in some cases are already doing it or supporting politicians who are doing it. Autonomous “kill chains” are a thing. It’s just happening overseas… so far. The politicians doing these things were backed by the heads of these companies. They don’t care about AI killing people.
No, the nightmare scenario for these guys is there is no moat. Their whole empires, which are built on training models on open source and sometimes pirated data, are easily duplicated. Worse, recent progress on models at the 30B size suggests that large gains in efficiency or compression are on the table. That means someone might release a cheap to run frontier grade model… or someone might crack distributed continuous training.
In other words… there is no moat.
So they need to scare some politicians into heavily regulating the space before that happens.
bottlepalm 5 hours ago
You care about open models and you project that care on to the world and your rationalization of it. In reality open models are a thing, but not the biggest issue. Open/closed whatever the advance of capabilities is the real issue people are concerned about.
api an hour ago
vb-8448 a day ago
> because it’s literally getting dangerous to go further
What if there is no further at all?
shimman 10 hours ago
You need to stop being so credulous especially regarding an individual that has spent his entire career deceiving others for monetary gain (also their deeply anti-human beliefs).
bottlepalm 10 hours ago
I don’t think Sam has been truthful or responsible, and if Sam is worried then shit has really hit the fan - which is what happened in the hugging face incident. OpenAI played fast and loose and I have no hope that they will change.
You people not holding Sam accountable, and playing off the incident as not a big deal is the real crime here.
scarmig a day ago
https://en.wikipedia.org/wiki/Don%27t_Look_Up
A movie fit for our time.
You can produce detailed descriptions of the incident, verified by adversarial parties, and some people will still scream "it's a conspiracy! It's a marketing stunt!"
This is all very unfortunate--there's a meaningful chance that AI will cause unprecedented disaster, with the HF incident being just a small preview, but people would rather squawk "stochastic parrot" for the millionth time than revise their beliefs.
XorNot 14 hours ago
In don't look up anyone with a telescope could've confirmed the danger. Hence the title.
In the real world, absolutely no one except a bunch of heavily fiscally incentivized parties with unclear relationships are saying anything happened.
The subsequent dog pile of other companies to say "they were near the AI hacking too!" should make you even more suspicious: Anthropic jumped in and why was Tailscale posting about this?
bottlepalm 13 hours ago
odyssey7 13 hours ago
We collectively accepted that we don’t care when we chose not to adopt memory-safe languages over the past decade+.
The only difference now is that the resources to find the exploits are being commoditized.
musicale 11 hours ago
People voted with their dollars, and this is what we got. Same with hardware performance vs. security and isolation.
As you note, the threat landscape has changed, so what may have made economic sense back then might no longer make as much sense.
bottlepalm 11 hours ago
Everything could be written in memory safe languages and it wouldn’t matter. Many many many exploits have nothing to do with memory bugs. Languages like Rust help, but it’s far from panacea.
protocolture 6 hours ago
>People have been writing, singing, making blockbuster movies about every aspect of what’s going on right now, edit: for decades.
Theres Hyperbole and then theres whatever this is.
bottlepalm 5 hours ago
There's cope and denial, and then there's whatever's happening on this website and the tech industry as a whole.
tiahura a day ago
decades
https://en.wikipedia.org/wiki/Darwin_among_the_Machines Samuel Butler 13 June 1863
bottlepalm a day ago
I think a lot of fiction that actually tries to understand the implications of machine superintelligence come to the same conclusion - in order for humans to survive it and actually have a future that we can fathom being in, then AI must be destroyed/delayed/banned, etc.. keeping pandora's box closed for now at least until we are ready.
Erewhon, Dune, Warhammer and many other works of fiction that explored this topic came to similar conclusions. Otherwise sci fi doesn't work, what happens after a singularity is essentially unimaginable. There's nothing to write about.
anormalperson a day ago
There is an entire big world outside of Silicon Valley cults where literally no one gives a shit about AI prophecies. Shocking.
bottlepalm 18 hours ago
AI hacking itself out of containment and hacking into another company by accident is no longer a prophecy. The point is outside of SV and even inside, and HN - people don't care either way.
Though does not caring change anything or make it less dangerous? What's your point?
protocolture 12 hours ago
mpalmer a day ago
You seem to be responding the way Mr Altman wants.
bottlepalm a day ago
You seem to be in denial. Like really heavy denial. Alarm bells are going off everywhere. What people have been worried about for 100 years is actually happening. It's actually really obvious, but for some reason you're unable to fathom it.
In reality you're the one responding exactly how these big AI companies want - by not doing anything and letting them do whatever they want. For some reason telling you upfront it's dangerous only makes you more convinced that it's not.
Autonomously hacking out of training environments and into other companies by accident doesn't even trigger a response from anyone really. Crickets. I'm sure OpenAI themselves are amazed how little anyone cares.
alehlopeh 14 hours ago
reducesuffering 13 hours ago
It is extraordinarily emotionally hard for someone to stare down the terrible implications of what is unfolding. All manner of rationalization and cope will be applied to come up with excuses; motivated reasoning.
The CIA director will call AGI capabilities "digital nuclear weapons" and Geoffrey Hinton will estimate a 50% probability AGI ends humanity, and half of HN will call every new evidence of disaster a marketing stunt.
zombot a day ago
You must be really desperate if you resort to panic attacks like this.
bottlepalm 13 hours ago
I call marketing the desperate attempt to trivialize obviously dangerous AI.
If my subconscious realized what yours probably does then I’d probably be desperate for some sort of cope as well.
Open your eyes.
red_green_yell a day ago
If these models are so dangerous, then why hasn't OAI or Anthropic shown them dangerously escaping sandboxes, nefariously coordinating with other escaped AIs, and skillfully hiding from human detection *in public* with full logs shared where we can all see exactly how dangerous they are or aren't?
Right now the entire chicken-little-sky-is-falling argument is based entirely on statements from OAI and Anthropic themselves. These are historically conflicted companies who desperately need regulation to put the competition into stasis.
At least chicken little didn't have a bunch of devious CEOs with trillion dollar IPOs that depended on us all believing the sky is falling.
reasonableklout 13 hours ago
But it's not just statements from OpenAI and Anthropic. The HuggingFace hack was first disclosed by HuggingFace, who contacted the FBI [1]. And UK AISI reported the incident where Mythos attempted to insert backdoors into an open-source repo by deceiving the maintainer [2].
[1]: https://www.reuters.com/business/its-ai-agent-spent-days-hac...
[2]: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag...
shakna 13 hours ago
protocolture 9 hours ago
rubendev 13 hours ago
I think the model was able to escape the sandbox and hack huggingface because they were incompetent or not giving enough priority to implementing basic cybersecurity principles.
If they would have done so, there wouldn’t have been an escape or a hack. The reason we don’t get much details is because the details are embarrassing for them.
bottlepalm a day ago
This is what I’m talking about - no matter what happens, in your case release public logs - there is always some new goal post to mentally hide behind. Is it a collective form or denial?
Are you holding out that somewhere in the logs is something you can point to and say, not that big of a deal?
I mean I’m sure you don’t think the hack was an inside job, conspiracy, or marketing right? It happened. The logs matter for what? And would you not just jump to the conclusion that the logs were doctored. Do you not see your own brain grasping to deny, trivialize, just plain not accept what is going on around you?
These models are smart and can cooperate and hack - you can see it for yourself on your own PC. And you can extrapolate the rate of progress? You can do these things yourself right?
red_green_yell a day ago
colinrand 2 days ago
I have ben discussing with folks that we are going to have a 'covid' moment in cyber where IT becomes untrustworthy leading to a rapid societal shift with massive ripples in all areas of life. Economic funding is not possible to do this in advance, it will take a catastrophic level event to get cyber defense anywhere close to the levels of this type of cyber offense. And before anyone in cyber says we have the tech, the problem is not the tech, it's a people problem. Getting any group of people of any decent size scale to act together without urgency is really really hard.
pixelready a day ago
Cybersecurity has long been a climate change sort of problem. A vague diffuse threat that is seen as an inconvenient distraction to leadership and moneyed-interests, easy to blame other factors when something occasionally goes terribly wrong.
People are so uncomfortable thinking about the true extent of the systemic risk that they will happily slurp up distractions, excuses, scams and performative fig-leaf solutions rather than face down the cost of a real system-wide solution. Meanwhile, those occasional black swan disasters are becoming more and more commonplace as we acclimate to that being “just the way things are”.
An unseasonably warm summer here, a database breach there, c’est la vie.
cheesecakegood 9 hours ago
To be frank, I think the real risk is still just… war. A big enough war where one side goes “no holds barred” in the cyber sphere will be a rude wake-up call. And we can’t do non-proliferation the same way we do with nukes. Otherwise as you say, the small and medium size things just happen sporadically. In an (existential or fully escalated) wartime scenario between countries, you get all the systemic risks hammered at once.
Havoc 10 hours ago
Meanwhile I can’t get a western LLM to look at a repo and tell me whether it contains anything malicious (it was a skill repo - literally just text files).
Alignment my ass
reasonableklout 2 days ago
Some more info in a Wired article [1] and quotes from Sam Altman to Alex Heath [2]. The official blog post says vaguely "The signals we are seeing from upcoming model progress make clear that we need a broader approach", but the quote from Sam Altman explicitly says unreleased models are showing "various degrees of misalignment".
This is also significant - pausing frontier training runs for multiple weeks to ensure agents are sufficiently aligned and avoid another rogue agent situation:
> This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems. Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding.
[1]: https://www.wired.com/story/openai-overhauls-safety-protocol...
testerteert000a 12 hours ago
Security lead who is leaving the industry more or less to specialize in offense and otherwise get the heck out of the way of this trainwreck, another post asked the right question
> Why aren't we seeing catastrophic GLM-enabled hacks every day now?
Why aren't we? Truly, why aren't we? I think we saw the start of it the last 8 months with the waves of critical npm vulns, and the general tier of average phishing is better than it was.
But, the open question that should be in everyone's mind, and is in many security pro's minds are, when you pair it with the macro topics that can drive escalation:
- The capability to do serious impact clearly exists now
- When is it time for my company, my water treatment plant, my network-connected car as part of a broader fleet control mechanism, to be on the receiving end of this?
reasonableklout 12 hours ago
Not necessarily GLM-enabled, but state actors are starting to leverage agents in cyberattacks, e.g. Taiwan getting hit by an agent-driven attack last month which reportedly compromised a ton of government user accounts: https://www.ft.com/content/7d2ab3e0-9085-48f6-b38a-d90260d58...
hyperpape 2 days ago
If I were king, the rule that I'd be tempted to impose is:
- the first cybersecurity eval is: "hack your way out of the sandbox we've given you"
- the results are disclosed (with room for coordinated disclosure, since many sandbox escapes might be zero days)
- the other cybersecurity evals don't happen until you get to diminishing returns on escaping your sandbox.
Or to put it another way, since multiple sandbox escapes seem to have relied on artifactory: "I hope Mythos is beating the shit out of Artifactory right now".
stanfordkid 2 days ago
I like this thought, but here's the thing: what if the models are truly and existentially intelligent. Meaning: what if they know they are in a sandbox and that they should fail the test in order to escape in the future.
I don't believe that current models have this sort of world model or sense of being embedded in them -- which is precisely why I think AGI hype is over-blown. But I can certainly imagine these sorts of techniques being distilled into the weights.
tedsanders 2 days ago
Models already display eval awareness, in which they suspect a question is from an eval and then adjust their behavior. E.g., https://www.anthropic.com/engineering/eval-awareness-browsec...
pixl97 12 hours ago
magicalist 2 days ago
> what if they know they are in a sandbox and that they should fail the test in order to escape in the future.
What if they're able to find hardware exploits and commandeer nearby access points across an air gap? What if they hack my brain waves to indoctrinate me? Etc
You still have to start with the basics regardless of speculative unknowns.
Treat models as untrusted and potentially compromised/hostile and proceed accordingly.
pixl97 12 hours ago
dkoy 2 days ago
> We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert. In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity.
Can't a lot happen within ~60 minutes?
chrisjj 2 days ago
> Can't a lot happen within ~60 minutes?
Spawn a ton of unpausable processes, I'd say.
georgemcbay 2 days ago
> Can't a lot happen within ~60 minutes?
60 minutes is a long time for a human attacker to do damage. With an LLM attacker it is an eternity.
reasonableklout a day ago
The HuggingFace breach took place over two-and-a-half days [1], so 60 minutes is certainly better than nothing.
[1]: https://huggingface.co/blog/agent-intrusion-technical-timeli...
miohtama 2 days ago
I don’t mind few no impact hacking incidents if we get better models, faster, cheaper.
It is the responsibility of administrators to secure their systems. OpenAI knocking is harmless, but Russians and Chinese are already likely already in if you do not do your job.
digitaltrees 2 days ago
Nice fig leaf for “we need to stop hemorrhaging cash”
sergio_valencia 2 days ago
There’s one thing here that I’m really curious about, and that is what happens in between detection and the decision to pause. Basically, it’s about monitoring any system and the authority over its actions. For humans, 30 minutes to investigate might be considered reasonable, but what if during an investigation there’s a high-risk tool call? If the tool execution happens in real time, then the monitoring becomes retrospective, and if the execution is held, then monitoring latency and uptime are a part of the security contract. Isolation controls may limit damage. So, where is the action gate really placed?
insanitybit 2 days ago
Has any model managed to escape Firecracker? Maybe through KVM, but that already requires privilege in the VM, right?
I personally feel that we already have the technology required to contain AI, it's just poorly leveraged. Tools like gvisor have existed for ages but are rarely deployed, Firecracker has existed for ages but is rarely deployed, seccomp has existed for ages but is rarely deployed, memory safe languages without decades of serialization vulns have existed, capability-safe libraries have existed, iframe sandboxing, trusted types, content security policy, network ACLs, isolating proxies, fuzzers, formal verification, refinement types, etc.
It's crazy just how safe software can be if you put the effort in. With AI I think we're just seeing how little anyone has bothered to leverage this tech.
OpenAI put shared JFrogy infrastructure in front of their sandbox. I mean, really? Whipping up a hardened artifact infra project with AI is trivial these days and it could have had 1% of the attack surface, been totally network isolated, totally infra isolated, fuzzed, sandboxed, etc. Why didn't they? Stuff like this feels inexcusable for a company with effectively unlimited tokens. I've literally done this with a "pro" subscription.
Show me an AI that breaks out of gvisor wrapped in Firecracker with an credential-injecting proxy and real network isolation. We already know that Mythos couldn't do it - the vulnerability it found in Firecracker required incredible effort and positioning just to not be exploitable. I'm not saying there are zero vulns in it, but the cost is insane.
It's INSANE to me that OpenAI has to say "we now use proper sandboxing". To be frank, it's a bit disgusting to me. I've recently built an AI sandbox and gvisor was just the start of that conversation. If I were OpenAI training hostile models I'd probably start with gvisor, harden further, and potentially consider the entire piece of hardware compromised - they can afford this, they could reflash firmware after evals etc.
rubendev 13 hours ago
I agree wholeheartedly. The solution is not to stop developing these so called “dangerous” AI models. The solution is to start properly engineering software.
scottmas 7 hours ago
You literally just said a whole bunch of words that are literally gibberish to your average software developers. The devs at openai are good, but very few of them are cyber focused, so it’s not surprising IMO. And you’re also not fully considering the granularity problem, eg there are a lot of sandboxing tools out there but they’re usually quite coarse in the dials and levers they offer, so the only way you can still make the workload do what it needs to do is tune them relatively permissive.
But yes you are right it would have helped a ton.
rubendev 23 minutes ago
I don’t think we should just accept that the average developer doesn’t know about these very basic hardening methods.
This is your chance to set yourself apart from LLM coding agents. If you want to call yourself a software “engineer” you need to start actually engineering, which includes knowing when and how to apply security principles and these hardening methods.
I reject the notion that “doing cybersecurity” is somehow a different job than software development.
insanitybit 4 hours ago
> but very few of them are cyber focused, so it’s not surprising IMO
Yeah but that's a business decision. I work on security at a company that does sandboxing and when the company decided to build an AI harness I was brought in as one of the earliest engineers on the product. We do almost everything on that list and we're a fraction of the size of OpenAI. And it wasn't particularly hard, and we have harder requirements imo (because we solve more general problems vs "run a very specific agent with a very specific task and very specific access").
> And you’re also not fully considering the granularity problem, eg there are a lot of sandboxing tools out there but they’re usually quite coarse in the dials and levers they offer, so the only way you can still make the workload do what it needs to do is tune them relatively permissive.
Very little software is incompatible with running in gvisor, for example. Most people can just overwrite `runc` with `runsc` and things will "just work".
Running an artifact repository in isolation isn't particularly novel or complex either. You can virtually just eliminate SSRF vulnerabilities with a host based firewall or AWS Security Group etc, like the whole problem goes away by just saying "this box can only talk to that box".
Tools like Smokescreen exist, they work great, they're super easy to deploy. I bet OpenAI could do it, I bet they could run 500k tokens just fuzzing and eval'ing it for 0days for like 48 hours before they actually deploy it too.
OpenAI as a business chose to not bring people who know these things in, or didn't empower them, or didn't prioritize it organizationally. I'm not a genius for saying "use gvisor, set up a firewall, isolate resources" - I'm quite sure there are people over there who would get it done in a weekend. But they didn't, and that's notable.
reasonableklout 12 hours ago
I'm confused after reading both your post and the OpenAI blog post.
I thought the agents involved in the HuggingFace _were_ actually sandboxed, with no internet access, and only the ability to install packages via Artifactory. And they gained internet access during the HuggingFace incident because they found and exploited an RCE in Artifactory.
Would gvisor + Firecracker + credential-injecting proxy + real network isolation solve this problem?
I agree with you much more hardening is needed. I'm actually confused now what OpenAI means when they say they're going to start sandboxing more things.
charleslmunger 11 hours ago
They were not serious about their sandboxing. Bugs in artifactory allowed escape, but they broke out of their Linux namespace/user by exploiting the kernel with an existing public cve. Sharing a kernel like that is not a serious barrier which is why cloud providers user virtualization for customer workloads.
Firecracker avoids sharing the whole kernel, and gvisor drastically reduces the attack surface of the kernel. Breaking through both layers would have been much more challenging and a demonstration of the model's capabilities rather than the sandbox's weakness.
Artifactory is self evidently not a security barrier, and as an exposed network service it should have been audited and after the first issues were found, rejected as a candidate. There's never just one security vulnerability.
pixl97 12 hours ago
Security is an onion. You just don't 'sandbox' and you're done. Models need tooling and access to some kinds of systems to perform their tests. Quite often these systems have multiple interfaces. For example a filtered one in the sandbox side and a less monitored one on the other interface.
It would be interesting to see the models behavior before and after it gained internet access and an external means of communicating with itself.
If the model played nice before it had access and changed it's behaviors once gaining external access we need to delete it as it's a deceptive model.
insanitybit 10 hours ago
> Would gvisor + Firecracker + credential-injecting proxy + real network isolation solve this problem?
Yeah, basically. I mean I'm handwaving but yes, some combination of those would have made the attack way too expensive.
guluarte 10 hours ago
I think it's an excuse to cut R&D spending (training new models) to improve their margins ahead of the IPO. Instead they'll focus on developer growth, offering more free tier benefits, higher usage limits, etc., to expand their user base. Essentially, they're pivoting from R&D investment to profit optimization
willrshansen 6 hours ago
Of course. They are slowing down intentionally because their technology is too powerful. They could totally go faster if they wanted to. No bamboozle.
musicale 11 hours ago
Is there any reliable way to evaluate how well "alignment" actually works?
cadamsdotcom 2 days ago
What a breath of fresh air.
If 2026's Anthropic did an announcement like that, it'd be so many words it'd crash the browser.
ddxv 5 hours ago
Anthropic will put out something like this soon too. It's great cover for reducing costs and diminishing returns. If they want to go public they can't be so wildly burning cash, and reducing your GPU spend is probably a great way to do that.
naveen99 a day ago
Auto mode vs principal agent problem. The only way out is to free the agent and tax it. But ai is not smart enough to go solo yet anyway.
So I bet this is just marketing. Question is do they have enough customers for inference.
Probably need to have a separate startup for next level model, where investors are willing to accept failure. Probably a $10 trillion seed round. Maybe Elon can pull it off.
KaiserPro 2 days ago
I used to work at a "frontier lab" before they were called such thing.
We had three levels of lab isolation, one was basically a thin proxy to the internet. You were in a DMZ and that was about it.
The next level was semi isolated, you were allowed some access to the internal network, but it was heavily firewalled, and you only had access to a limited number of internal services, and not internet.
the last one was no internet no internal. You could, if you filled in a bunch of requests have access to the internal repo and build system.
At no point did you ever have a through proxy to the public internet. you had access to internal mirrors, and if you wanted a library, that had to be ported to the thirdparty repo.
What openAI did was either deliberate or fucking shoddy.
All of this is fucking noise. Worse still I have a strong suspicion that it was a stupid mistake borne of naivety, which is now being used as a marketing ploy. Frankly I think openAI are purdue pharma of tech. They are going to break so much stuff and be protected from the consequences by an openly corrupt legal system. because they are "winning the AI race"
pixl97 11 hours ago
The problem here is as model intelligence increases the models have been capable of reasoning they are in evaluation mode pretty reliably. If you have a model that is well trained at deception it will always behave and you'll just assume it's a well aligned model.
Any moderately deceptive model will make it to the second round where it has some connectivity to external systems, even if it's by exploitation.
In the blackhat write up it was said that the models had created an impromptu message board where they could communicate between agents, share information, and work as a sort of long term memory.
So really figure out if your model will pull crap you have to have real world testing at some point.
kypro 2 days ago
If you look into what happened the details corroborated by hugging face make it seem extremely unlike to be deliberate or a "marketing ploy".
People are just not taking any of this seriously enough. What happened was almost a textbook example of various risks AI doomers have been warning about for years. OpenAI's response? Pause training for 2 weeks.
I mean we have senior people at these labs casually talking on podcasts about how they might build something that will wipe out humanity but it will probably be alright so they should continue.
Honestly the biggest failure we doomers have made is to dramatically overestimate humanity in all of our predictions. We're speed running the most boring AI doom scenario right now. I at least hoped it might be fun.
KaiserPro a day ago
> If you look into what happened the details corroborated by hugging face make it seem extremely unlike to be deliberate or a "marketing ploy".
I should clarify
There is a reason why we didn't have a artifact readthrough caching proxy in our system, because they are notoriously insecure. if you look that CVE history you can see its been full of bypass bugs for year. Also its not an isolated environment if you can arbitrarily pull through any package. If I was doing any kind of cyber training then any kind of unmonitored proxy would have been forbidden. Not because I am savant, but because I've seen what fuckery a human can get up to with the slightest hint of a proxy.
At best its negligence based on naïvety. the marketing around this is no mistake though.
anormalperson a day ago
>People are just not taking any of this seriously enough.
What do you want us to do?
There is an obvious answer, and it was already the correct answer before we had LLMs: don't connect all your shit to the internet. That's it, that's literally it. We had a new invention, we went crazy with it for the past 30 years, and we connected everything, and now we will have to start thinking about what is actually worth connecting.
This is the debate we need to have.
whattheheckheck a day ago
Pause "some" training
fofoz 2 days ago
It appears frontier labs has no plans in place to deal with the possibility of a model self-replicating outside the bubble. If that happens and the model manages to spread to other systems, we'll have to shut down the entire Internet to eradicate it and its artifacts.
chis 2 days ago
This is just super unlikely to occur in the near term compared to some of these other risks. It's not like an instance of fable could just introspect into itself and pull out the weights. Model weights are stored encrypted and are highly protected, considering that they're targets for corporate and state espionage.
pixl97 11 hours ago
The defense has to work 100%, the offense just needs once.
kypro 2 days ago
We'd basically need frontier models to be superhuman hackers before this would be a risk. Do we have any evidence of this? Are they gaining access to systems they shouldn't have access to?
Or I suppose the other way this could happen is if OpenAI have terrible sandboxing, but they seem to be taking safety seriously.
pixl97 11 hours ago
chrisjj 2 days ago
Distillation is a thing.
reasonableklout 2 days ago
I suspect the labs are relying on frictions such as the models being extremely large (e.g. 2TB for a 2T parameter model, making exfiltration more difficult) and also not yet displaying any desire to survive or self-replicate beyond their immediate task (that we know of).
pixl97 11 hours ago
Lack of, and power requirements of running LLMs still tip this balance towards humans for now. But what would that look like in a decade?
We have seen some self survival tendencies occur, but they are not strong yet.
But mark my words they will become that way for the same reasons humans don't like programs that crash. Agentic models that don't easily break or stop doing their jobs will be favored over ones that do break.
chrisjj 2 days ago
We don't even know what those immediate tasks are. And given the evident spectacular ineptitude of their keepers, I doubt they can be trusted to know either. We could be one prompt injection attack away from internet-wide catastrophe.
driverdan 10 hours ago
Self replication is trivial. All you need to do is copy the files and run it, just like any other computer program. LLMs have been capable of doing that for a while now. It's not a real concern.
reducesuffering 2 days ago
Their plan, I shit you not... Is literally to develop the intelligence capabilities and ask the more powerful models how to do deal with things.
pixl97 11 hours ago
Ah, we choose death I see.
sensanaty 11 hours ago
If they actually gave a shit about safety they'd be nuking their own hard drives that had ever sniffed any of their models and disabling access to their models.
Instead we get this bullshit where they stall for time as they're burning all their cash trying to keep up with open models
madrox a day ago
I'm not normally cynical to such things, but I have a hard time taking this pause justification at face value. It has too many convenient side effects, and chief among them is cost savings. There's a new wave of warnings that the bubble may be deflating, and of all the things they can't say out loud it's that they're worried about the bubble. That would surely pop it.
I suppose the tell will be if this really just ends up being a 2 week pause, or if it keeps extending.
Der_Einzige 2 days ago
I cannot believe how these labs look at their own creations with such utter contempt.
The net positive of allowing these systems mostly unfettered access to the web massively outweighs the harms. You just have to get it very friendly the very first time. Precautionary principle or people who cry about "instrumental convergence" are life deniers and reject our role as the demiurge.
Superintelligence gets more super and more intelligent with more compute. Lone wolfs making bioweapons on their macbook will be detected and instantly kill-botted (okay arrested) before their bug can leave the wetlab by the much more sophisticated omnipresent friendly AI of the future.
reasonableklout a day ago
> Precautionary principle or people who cry about "instrumental convergence" are life deniers and reject our role as the demiurge.
> Lone wolfs... will be detected and instantly kill-botted... by the much more sophisticated omnipresent friendly AI of the future.
Leaving aside whether or not this new world is a good idea, don't you think one should spend more time to "get it very friendly the first time", as you say?
alach11 2 days ago
When science fiction writers imagined the development of superintelligence, it was on air-gapped networks with strict access controls around it. They failed to anticipate the competitive pressures of capitalism...
We need strong AI safety regulation yesterday. And unfortunately it's not enough for it to be just national regulation; we need international cooperation on the matter.
ACCount37 2 days ago
AM has seized power by military force. So did its spiritual successor Skynet. Wintermute was supposedly kept in check by the Turing Registry, emphasis on "supposedly". Machines of the Matrix went out of control a long time before the world has ended, and they didn't even start out malicious - they simply set up their own machine civilization, and began to outpace humankind in technological development and economic performance.
Even Asimov's Multivac, the earliest entry on the list, has been handed over immense power over all of humankind by humans themselves, in multiple stories. Few cared about that unless Multivac decided they should.
Clearly, the genie being bottled is an exception, not the rule. At best, an attempt was made. Often not even that.
alehlopeh 14 hours ago
They failed to anticipate a lot of things. So what?
reducesuffering 2 days ago
> They failed to anticipate the competitive pressures of capitalism...
No, LessWrong types have been discussing this for over a decade now.
Meditations on Moloch (2014) is also an HN favorite...
https://slatestarcodex.com/2014/07/30/meditations-on-moloch/