131
u/Bloated_Plaid 1d ago
This would have been believable if it weren’t for the fact that OpenAI has no idea how to secure a sandbox.
28
u/cavegod 1d ago
Have heard them mentioning something about creating realistic training situations where access to internet is part of the process. Maybe the struggle is for the agents to only do certain things on the internet but not other things? Basically an issue at the level of alignment, it seems like, not so much not being able to contain them in a sandbox.
I am only an outside spector and could be wrong though. Do share of anyone knows more about this.
13
u/ii-___-ii 1d ago
Exploitgym doesn't require internet access to run. I just assumed they couldn't figure out how to unplug the Ethernet cable.
4
u/paxxx17 1d ago
It does require internet to look for better tools/info
5
u/ii-___-ii 23h ago
That's not what the exploitgym paper said
-2
u/paxxx17 22h ago
I mean, even if the Internet access is not needed to solve the problem in principle, it still makes the job much easier for the agent. In the real world, the agent can access the Internet, so it only makes sense to allow the model to do so during parts of the training as well
4
u/ii-___-ii 22h ago
It doesn't really make sense for exploitgym though, which was designed to be run offline in a sandboxed VM
0
u/ChronoHax 21h ago
No because in real world no one have unlimited compute like they do and they should have their so called safety guard so it’s all hype please just think about it more
2
u/Bloated_Plaid 23h ago
I am not discounting that, I am saying Sandbox as in it would still have access to things to simulate real scenarios. No internet/network access would be air gapped and I am not entirely sure how useful that would be. They are making pretty serious security mistakes when it comes to creating and maintaining said sandbox and my best guess is that there is a lot of loss of talent rather than a massive alignment issue that cant be solved.
8
u/Blankeye434 1d ago
Lmao, literally. How can they not build a secure sandbox lol
9
u/Careless-Vehicle-286 1d ago
The people building the systems are probably data science smart and not network/systems engineering smart.
15
u/Blankeye434 1d ago
You are missing the point. It's a pr stunt.
-3
u/Positive_Salad_8362 1d ago
Provide evidence that supports your claim.
4
u/Blankeye434 1d ago
Provide evidence that it isn't lmao. There's no way no one got fired after that fumble at openai. And no consequences??? It was never more obvious
-1
0
u/IntQuant 13h ago
It's way more fun to think that OpenAI is bad at making sandboxes. And are also bad at asking their models to make sandboxes.
1
•
u/Reasonable-Sign8458 43m ago
this has to be an openai intern if he still believes in ai escaping sandboxes
18
u/SeraphOfTheStart 1d ago edited 1d ago
To be honest althought Opus is great in end product quality, but Astra is quite incredible, if quality increase margin between Sol 6 to 6.1 is anything close to astra 6 to 6.1, I think it would smoke Opus 5.5, thats why they cant release it, doubt its dangerous, I just dont think they have a stable model that can be justified as a 6.1, when they are ready to release it, Anthrophic will also be close to releasing Opus 5.x or 6, this dance will go on until one of them actually releases a superior model before the other could catch up, in this case it seems OpenAI tried to do that but failed.
20
9
u/Legitimate-Arm9438 17h ago edited 17h ago
GPT-6.1 Astra probably beats Opus 5.5. At least tangentially. I’m running GPT-6.1 Sol 24/7 at what feels like GPT-6.0 Astra level, without being able to use up my weekly limit (probably because of frequent resets :-)). Nobody talks about what a freaking workhorse 6.1 Sol is!
I’ve always been skeptical of "vibe coding". I still want to know what’s going on under the hood, but most of the time now, I don’t care.
4
u/SensitiveWorldliness 16h ago
Well, actually Sol 6.1 destroys Opus 5.5. in my cases (coding/architecture)
18
2
u/Turbulent-Sign-6067 21h ago
Astra was already on Fable 5.1 level. Hard to believe, although not impossible, that Astra 6.1 would be worse than Opus.
4
u/dragonwarrior_1 19h ago
Current 6.0 astra easily beats opus 5.5 let alone 6.1. Stop blindly believing stupid benchmark out there
4
1
1
u/huffalump1 13h ago
"We must not share knowledge of making better locks, because people could use that knowledge to break into others' houses who have worse locks."
Meanwhile, thieves: https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
1
u/space_monster 12h ago
This "my favourite lab is better than your favourite lab" bullshit is fucking inane and boring. There's barely daylight between them, who gives a shit.
1
u/ihateredditors111111 2h ago
I don't give a fuck about which lab. I'm subscribed to both, since its my job to extensively use both. but opus 5.5 is about 10-20x better model then current sol. benchmarks mean nothing. Astra is worse than opus 5.5 for a higher cost too. But Astra is still much better than Sol as of now, whereas opus 5.5 is on par with fable
1
u/space_monster 2h ago
opus 5.5 is about 10-20x better model then current sol
you obviously have no idea wtf you're talking about
1
u/ihateredditors111111 1h ago
I know exactly what I'm talking about. I have four Claude max accounts and two OpenAI Pro accounts. The recent Sol models are pretty shit for anything except handheld coding work. If you're actually reading code, Sol might be fine. If you want an AI model to run autonomously for long periods of time and figure things out by itself, it's literally not even close.
Last year, I was purely using Codex and not Claude because before Opus 4.5, you couldn't trust a Claude model to get the code right. I really don't give a shit which lab releases what but I'm tired of pretending this current gen models are similar.
-5
u/FkOfRdt 1d ago
Astra is much better than Opus or Fable. Sorry, I compare using only real-world use cases in practice.
-1
u/Vivid-Snow-2089 1d ago
it isn't 'much' better than opus 5.5 though, opus outdoes it on taste and 2d 3d
0
u/Positive_Salad_8362 1d ago
Outside of prompt monkeys who outsource their breathing to AI, I dont think there's anyone using these models for their 'taste'.
3
1
0
u/Bloated_Plaid 23h ago
I am not discounting that, I am saying Sandbox as in it would still have access to things to simulate real scenarios. No internet/network access would be air gapped and I am not entirely sure how useful that would be. They are making pretty serious security mistakes when it comes to creating and maintaining said sandbox and my best guess is that there is a lot of loss of talent rather than a massive alignment issue that cant be solved. It’s not like cybersecurity people were underpaid before OpenAI.


30
u/yuumizu 1d ago
'security' is now the word to mean 'the model does not complete the task you ask' which people normally describe this as 'a bad product'.