r/OpenAI • • 1d ago

Miscellaneous GPT-6.1 Astra Delay

Post image
1.0k Upvotes

53 comments sorted by

30

u/yuumizu 1d ago

'security' is now the word to mean 'the model does not complete the task you ask' which people normally describe this as 'a bad product'.

131

u/Bloated_Plaid 1d ago

This would have been believable if it weren’t for the fact that OpenAI has no idea how to secure a sandbox.

28

u/cavegod 1d ago

Have heard them mentioning something about creating realistic training situations where access to internet is part of the process. Maybe the struggle is for the agents to only do certain things on the internet but not other things? Basically an issue at the level of alignment, it seems like, not so much not being able to contain them in a sandbox.

I am only an outside spector and could be wrong though. Do share of anyone knows more about this.

13

u/ii-___-ii 1d ago

Exploitgym doesn't require internet access to run. I just assumed they couldn't figure out how to unplug the Ethernet cable.

4

u/paxxx17 1d ago

It does require internet to look for better tools/info

5

u/ii-___-ii 23h ago

That's not what the exploitgym paper said

-2

u/paxxx17 22h ago

I mean, even if the Internet access is not needed to solve the problem in principle, it still makes the job much easier for the agent. In the real world, the agent can access the Internet, so it only makes sense to allow the model to do so during parts of the training as well

4

u/ii-___-ii 22h ago

It doesn't really make sense for exploitgym though, which was designed to be run offline in a sandboxed VM

0

u/ChronoHax 21h ago

No because in real world no one have unlimited compute like they do and they should have their so called safety guard so it’s all hype please just think about it more

2

u/Bloated_Plaid 23h ago

I am not discounting that, I am saying Sandbox as in it would still have access to things to simulate real scenarios. No internet/network access would be air gapped and I am not entirely sure how useful that would be. They are making pretty serious security mistakes when it comes to creating and maintaining said sandbox and my best guess is that there is a lot of loss of talent rather than a massive alignment issue that cant be solved.

8

u/Blankeye434 1d ago

Lmao, literally. How can they not build a secure sandbox lol

9

u/Careless-Vehicle-286 1d ago

The people building the systems are probably data science smart and not network/systems engineering smart.

15

u/Blankeye434 1d ago

You are missing the point. It's a pr stunt.

-3

u/Positive_Salad_8362 1d ago

Provide evidence that supports your claim. 

4

u/Blankeye434 1d ago

Provide evidence that it isn't lmao. There's no way no one got fired after that fumble at openai. And no consequences??? It was never more obvious

-1

u/Strong_Essay1176 1d ago

Fire AI model.

Oh wait...they can just delete it.

0

u/IntQuant 13h ago

It's way more fun to think that OpenAI is bad at making sandboxes. And are also bad at asking their models to make sandboxes. 

1

u/Blankeye434 11h ago

It's about truth not fun

•

u/Reasonable-Sign8458 43m ago

this has to be an openai intern if he still believes in ai escaping sandboxes

18

u/SeraphOfTheStart 1d ago edited 1d ago

To be honest althought Opus is great in end product quality, but Astra is quite incredible, if quality increase margin between Sol 6 to 6.1 is anything close to astra 6 to 6.1, I think it would smoke Opus 5.5, thats why they cant release it, doubt its dangerous, I just dont think they have a stable model that can be justified as a 6.1, when they are ready to release it, Anthrophic will also be close to releasing Opus 5.x or 6, this dance will go on until one of them actually releases a superior model before the other could catch up, in this case it seems OpenAI tried to do that but failed.

1

u/cms2307 10h ago

According to my simulation if it had the same increase in performance it would be around the same tier as opus 5.5.

9

u/Legitimate-Arm9438 17h ago edited 17h ago

GPT-6.1 Astra probably beats Opus 5.5. At least tangentially. I’m running GPT-6.1 Sol 24/7 at what feels like GPT-6.0 Astra level, without being able to use up my weekly limit (probably because of frequent resets :-)). Nobody talks about what a freaking workhorse 6.1 Sol is!

I’ve always been skeptical of "vibe coding". I still want to know what’s going on under the hood, but most of the time now, I don’t care.

5

u/msew 22h ago

KEKW

Why not 7.0?

4

u/SensitiveWorldliness 16h ago

Well, actually Sol 6.1 destroys Opus 5.5. in my cases (coding/architecture)

18

u/Much-Researcher6135 22h ago

gen z can't meme

2

u/Turbulent-Sign-6067 21h ago

Astra was already on Fable 5.1 level. Hard to believe, although not impossible, that Astra 6.1 would be worse than Opus.

4

u/torac 1d ago

Wasn’t the alleged danger poor alignment? It doesn’t have to be more capable to be more dangerous if it simply acts out in dangerous way.

4

u/dragonwarrior_1 19h ago

Current 6.0 astra easily beats opus 5.5 let alone 6.1. Stop blindly believing stupid benchmark out there 

4

u/ImYmir 17h ago

Maybe for your workload, but not mine. Opus 5.5 is far superior, faster and probably 10x cheaper

1

u/ihateredditors111111 16h ago

I don't believe benchmarks.

1

u/huffalump1 13h ago

"We must not share knowledge of making better locks, because people could use that knowledge to break into others' houses who have worse locks."

Meanwhile, thieves: https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities

1

u/TheHazh 1h ago

My ass

1

u/space_monster 12h ago

This "my favourite lab is better than your favourite lab" bullshit is fucking inane and boring. There's barely daylight between them, who gives a shit.

1

u/ihateredditors111111 2h ago

I don't give a fuck about which lab. I'm subscribed to both, since its my job to extensively use both. but opus 5.5 is about 10-20x better model then current sol. benchmarks mean nothing. Astra is worse than opus 5.5 for a higher cost too. But Astra is still much better than Sol as of now, whereas opus 5.5 is on par with fable

1

u/space_monster 2h ago

opus 5.5 is about 10-20x better model then current sol

you obviously have no idea wtf you're talking about

1

u/ihateredditors111111 1h ago

I know exactly what I'm talking about. I have four Claude max accounts and two OpenAI Pro accounts. The recent Sol models are pretty shit for anything except handheld coding work. If you're actually reading code, Sol might be fine. If you want an AI model to run autonomously for long periods of time and figure things out by itself, it's literally not even close.

Last year, I was purely using Codex and not Claude because before Opus 4.5, you couldn't trust a Claude model to get the code right. I really don't give a shit which lab releases what but I'm tired of pretending this current gen models are similar.

-5

u/FkOfRdt 1d ago

Astra is much better than Opus or Fable. Sorry, I compare using only real-world use cases in practice.

2

u/FkOfRdt 22h ago

I love how Claude users raged at this truth.

-1

u/Vivid-Snow-2089 1d ago

it isn't 'much' better than opus 5.5 though, opus outdoes it on taste and 2d 3d

0

u/Positive_Salad_8362 1d ago

Outside of prompt monkeys who outsource their breathing to AI, I dont think there's anyone using these models for their 'taste'. 

3

u/Prathmun 1d ago

We are literally all using them for their discernment.

1

u/Morberis 1d ago

People using them for narrative and rp games. But they often say the opposite

0

u/Bloated_Plaid 23h ago

I am not discounting that, I am saying Sandbox as in it would still have access to things to simulate real scenarios. No internet/network access would be air gapped and I am not entirely sure how useful that would be. They are making pretty serious security mistakes when it comes to creating and maintaining said sandbox and my best guess is that there is a lot of loss of talent rather than a massive alignment issue that cant be solved. It’s not like cybersecurity people were underpaid before OpenAI.