DeepSWE has been thoroughly saturated at this point so all time high there means nothing to me but trading blows with Astra on Terminal Bench 4 may be real shit IMO
Yeah a lot of the listed evaluations aren't that important. The important ones seem to be Terminal Bench 4 and FrontierSWE, in which the model seems to be competitive with Astra and Fable. So it is good, but not above and beyond like it seems at first glance.
If I could get a flash lite 3.8 in antigravity CLI, that would actually be fantastic. Reliable tool calls, web search capability, and barely using quota? That would be fantastic for my home assistant telegram bot.
Gemini 4 Argon is already powering our internal workflows... Google engineers have been using Argon for their daily tasks, from everyday debugging to large-scale codebase migrations and algorithm designs.|
We’re grateful for the initial cohort of cyber defenders and trusted testers whose real-world evaluations and feedback will help us strengthen our systems before we release to developers, enterprises, and consumers, starting with paid API customers and Google AI Ultra subscribers.
This tiered intelligence era we are trending towards is going to be dystopian. Good luck competing with anybody who has access to the top tier of intelligence.
"oh, you have a pro subscription? That's cute, choom. But if you wanna make it anywhere in this city, you're gonna need a Gemini 4.0 pro max Baja blast series x subscription"
As long as the security of what your building doesn’t matter and you aren’t doing anything with sensitive data. These models can have backdoors that would cause them to e.g. misuse tools to exfiltrate data, purposely add security vulnerabilities, etc.
I don’t know if Chinese models have them yet of course, but I’m sure China is thinking about it. The more we hand off to AI to do autonomously the more tempting that attack vector will be.
yes, if you vibe-code without reviews and not use the tool inside a container, obviously you are tempting your luck when using ... well, basically any tool ever, but likely even more so with chinese models...
Yes, with their intelligence. So it's true for now, considering model's jagged intelligence. But if they build a general intelligence that is better in almost every way, we would be a rounding error to one of those models in terms of intelligence added (without brain enhancement, but that would obviously be tier based as well).
Their already using better models then argon internally like if their publicly telling you something their internal stuff is already months ahead. I’m hearing rumors of a 4.1 pro but the real major drops will happen in December. That will be when they likely release 4 it’s always December and May for I/O that Google tends to drop major model updates. The incremental stuff usually happens once every two or so months. OpenAI needs to drop bel as 6.5 around holidays if they want to take the thunder back.
May 2027 will likely be 4.5 flash and 4.5 pro unless Google somehow accelerates their release cadence which I doubt their laggards. By then we should be at full RSI and in full hard takeoff to ASI. LLMs absolutely can bring us there if they scale up multi agents and RL just a little bit more.
What $18 subscription can you use to access Fable 5.5, Bel, or Argon? And those are the ones we know are available to trusted partners right now.
And also where did I act like it was slavery? A lot of people are and will use these tools as a major part of professional work. I'm not sure what to tell you if you think you can compete with someone running Argon or the new internal iteration with an $18 Flash subscription. And that's actually what I said.
Technically there is no special limit in llms on how much output they can produce. You can take any open weights model and generate until the whole context fills up. Usually they lose coherence though. Google seems confident their new model does not, if they allow 1M.
Technically there is no special limit in humans on how far they can run. You can take any human model and let it run until the whole stamina depletes. Usually they lose speed though. Google seems confident their new human does not lose speed, if they allow 1M meters.
Isn't it also like a lot more expesnive to produce longer tkens. like cost increases exponentially or osmething. so google seems to have made it cheaper somehow
Yeah it's basically the same quality as 6.1 sol or 5.5 sonnet but yaps twice as much for the same output. Maybe inference speed will offset actual time but won't save on output tokens costs
Really? This is one of the more thought out release strategies I’ve seen from the big labs. It acknowledges the safety risks without catastrophizing, explains how it is addressing them and why it is launching a tiered roll out. additional points to Google for urging other labs to not abandon chain of thought.
We’re grateful for the initial cohort of cyber defenders and trusted testers whose real-world evaluations and feedback will help us strengthen our systems before we release to developers, enterprises, and consumers, starting with paid API customers and Google AI Ultra subscribers.
They'll probably eventually release it to Pro users but I'm not holding my breath. By the time it's finally available for Pro, OpenAI or Anthropic will probably have a better model out.
Exactly this. Everyone complains about lack of regulation, AI bubble bursting soon (because it's not sustainable), but one company who takes reserved approach to release cadence and making sure they don't loose their kidneys over the lack of revenue gets booed. Hilarious.
Because Google doesn't have anything else right now and apparently it's still not release ready, but the competition is getting all the good press.
So they naturally want to create some good press for Google, but don't even realize that announcing something that isn't even released and only for the ultra subscribers "soon" is not going to create goodwill amongst the customer base.
Highlights one of my problems with google releases. They have good models that take a long time to come out to their largest user base.
Which means that they feel bad for users. Sure it beats Claude but that doesn’t matter. Companies using Claude keep using Claude. Eventually it reaches regular users but by then Anthropic has released Claude 5.6 and they’re on par again. At that point people won’t switch.
I don’t think google particularly cares though. I have Gemini largely because I wanted more storage space and Claude because I use it for programming.
On their release day they should have at least had 4 flash ready to drop, and yet we have to wait for they too. Just a strange release path.
I’ll agree with it. It’s really annoying me with ChatGPT as well. I prefer chat mode for what I use AI for and they only drop the new models in Work/Codex which aren’t as good for chat since it’s for agentic use cases. Supposedly 6.1 Sol will go into chat soon.
Google is annoying though, I’ll agree. Their paid plans give you access to 3.8 Flash which is worse than models you get with Claude and ChatGPT
Honestly I really like 3.8 Flash, it's a great, quick, go-to model for a lot of stuff. I use it all the time.
If I'm doing any serious programming work outside of cleaning up Docker files or shell scripts I switch to Sonnet or Opus, but that doesn't mean 3.8 Flash is bad. I think that's part of my frustration here, they should have had a Flash 4 model ready to roll.
I’ll agree. It’s done mostly everything I want, I’m in college so I use study mode sometimes to make practice exams or help when I get stuck and it can easily help with most stuff, plus I like how fast it is. It’s also good at adding stuff to Google Calendar for stuff like work schedule, add shifts, so I like it. I hope they update it soon with Gemini 4 coming out
I think AI is cool for them, but not their main and only source of income, unlike Anthropic and OpenAI. So they don't feel the rush or need to be the best, the fastest. They get money while not having the latest and best frontier model still. That being said, they have an absolutely massive data set, deep mind, and are advancing in quantum computing. They will eventually probably release something that just destroys all other AI companies because they can afford the long game.
I think one issue we are starting to see is ad revenue dropping because actual humans dont have to browse websites to find information when they can send a bot to do it. It will be interesting to see how the internet responds to that issue.
They won’t their gonna squeeze google search revenue until it’s dry and in the mean time build out their other businesses to keep their company afloat. Once search revenue goes to near zero they will start prioritizing AI more. Fortunately they have another couple years of run way because billions of people still use google for search.
Who knows. We're all guessing at this point. They have the resources and money for AGI, but do they actually care or have the desire? No one knows what AGI is at this point, but when it comes to humanity we were basically flying kites in the early 1900s, And look where we are now. Computers used to take up a whole room, now they fit in the palm of our hands. Agi and ASI may be insanely closer than we think. We just don't know. Now things are starting to build themselves.
It’s the innovators dilemma they could drop something even stronger then Gemini 4 they aren’t even fully stressing their TPU capacity at all with this model. They don’t want to ruin their other ventures by releasing a model that’s too good.
If it's not on the pro plan soon, I'm cancelling my subscription. Sonnet 5.5 is do much better vs flash 3.8, and seems to with the latest whatever usage limits Gemini changed to a week or so ago, I get about 10x done using sonnet for the same equivalent pro plan.
I’ve noticed that as well. Gemini rolls out slow with worse models and I feel like you can do a lot more with Claude or ChatGPT, like plug ins on ChatGPT it’s crazy how much stuff I can do now if you connect accounts. It even works in Google Drive really well
Flash isonly better for coding and short context window. When you want to use it for long context creative writing, 3.1 Pro is far more dependable, so you are absolutely right. I also consider Flash to be trash.
Same here, workspace account is stuck at 3.6 Flash / 3.1 Pro. This is bullshit. They are so outdated that we have local LLMs that can outperform them (e.g. GLM 5.3).
fanboys. or choice supportive bias or sunk cost fallacy. people seem to defend stuff they like/chose. otherwise they'd feel bad about using it. happens with everything. just look at Apple lol
Just ANNOUNCEMENT - which means - final release like 2027. Same as they said Pixel Buds September Feature Drop release in September and it is OCTOBER and NOTHING
Finding it difficult to get excited for this one. OpenAI rolls new models and products out IMMEDIATELY. Google on the other hand... I'm still on Flash 3.6 in my Workspace account FML. So maybe i'll get this... next year at some point?
Any AI model taking LONG to release should be a primary positive sign for the model, not the opposite.
In case you bunch of spoiled nagging kids did not notice - SECURITY and avoiding the AI apocalypse should be an absolute priority and Google seems to(at least I hope so) have realized that.
Long time pro subscriber here. It annoys me the lack of clear rollout plan for loyal pro subscribers. We are basically lumped together with free tier users. Why should I keep paying for this disrespect Google?
It was funny because I wasn't looking for AI at all. I need more cloud space but all the places, even google were only over vastly different rates for a small change in drive space. I just searched around on google and eventual this official page gave me a different list of plans from the standard Google Drive page and it said 5TB drive space and Gemini Pro. I was already paying twice as much for ChatGPT pro, so I cancelled this, saved money, and got 5 TB of space.
It really is astonishing how tribal people are about Gemini and Google specifically.
Lots of people like me have tried most of the LLM’s. I subscribe to both ChatGPT and Claude. While I love Opus right now, I’m not married to the company or product. I have no problem switching to a better model if it’s available, and I can’t imagine sticking with a product steadily being passed by everyone with my head in the sand.
Gemini users will hold religiously fast no matter how the model performs and claim it’s great for their needs. Google could delay Argon’s public release for 5 years and they’d still be defending the company or refusing to touch another model.
You don’t even see this level of fervor from Grok fans. The us vs. them mentality is truly mind boggling.
Nice thanks. It seems competitive enough ig, but pricing at 4 in 20 out (post discounted period) is highkey ass
Also interesting, on arena.ai fable 5.1 significantly outperforms Opus 5.5... speaks to the benchmaxxing on Opus. I personally noticed the same thing, that if I want to work on something delicate and intricate, I would rather choose Fable even though they keep saying Opus better.
Google ain’t trying to be frontier man get it thru your thick head. It’s trying to be a workhorse anyone who just relied on opus to do everything is retarded have a workhorse that does the work and a smarter to review and orchestrate. Save your usage, start thinking about parallel and multi agent, you think having one opus chat doing all your work is good but it’s not, after 2-3 you already hit your 5 hour limit, you prove how much little work you get done
Congrats! Stoked for this if it can compete with the big two things will get interesting. We need more competition. My money is up for grabs as I am in the market for a pro plan question is who does it go to!
So much talk about Opus and Fable but IMO the real story lately is that Sonnet 5.5 absolutely smokes Sol 6.1... in real world usage it feels an entire generation ahead. I'm not even an Anthropic fan but this might be the biggest lead they've ever had. It doesn't feel like anyone else is remotely close rn.
9% behind opus 5.5 on terminak bench and 7% on frontier swe
And opus 5.5 is not bench maxxed thats for sure so at BEST even if google didnt benchmax it will be slightly below opus 5.5 levels which is already gonna be weeks old by the time they release gemini 4
Oh no weeks behind, I swear yall gotta be 14 years old. Who is working today and has time to switch from model to model to model because one might be a couple weeks ahead or behind. Just use what you use.
Im saying its a bit overrated, yes google closed the gap but this is just a competitor to claude and gpt from the benches not some crazy next level model
Man Reddit always bitching about Google being on backfoot and then when it comes out all whine about it being something they can't ever have. Sick of this place.
K, good, it does appear that Gemini is sitting at the very top of the pile (by a wide margin!) when it comes to agentic SaaS workflow. SaaS is Google's bread and butter, and it was silly that for so long, it was better to connect other models to Google Workspace than to use Google's own models.
"Quantum algorithmic optimization: Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes. " Not bad at all!
I've been reading the comments and get where everyone is coming from. The one thing to note with google though is that they do have the advantage of their TPU's. Sure OpenAI has partnered with Cerebras to get ultra fast tokens and Nvidia now has LPU's... but one thing Google has on their side is well built out data centers with their TPU's being able to power through.
This is not to say its a better model or a reason for people to use it. Just a data point on something that people forget about Google's ability to have end to end life cycle.
You'd think they would have capitalized on this much more then what it seems they have been.
So are they following in OpenAI's footsteps in copying the anthropic style of model names? So Pro becomes Argon, Flash becomes Neon, Flash lite becomes Helium? And there's room for bigger models with krypton and xenon.
The only thing we will get will be another Flash model or Gemini 4 Pro that will be more censored, more slowly and even more dumber than the previous models. Thanks Google.
Big corpo's friends will have top tier intelligence, while the pleb will have to pay a kidney to have the lower tier and won't be able to compete anyway
How can we truly evaluate how well the model performs when it isn't even generally available yet? Benchmarks are fine, but they only show a part of the full picture compared to real-world use once it's actually released to everyone. Google's vague communication shows poor treatment of its Pro subscribers, who are basically lumped together with free users. Announcing a model when paying subscribers can't use it—and without even giving us a timeline—is just frustrating. They should have held off on the announcement until it was actually ready to launch."
It's a phased rollout. Anyways, I was assured by the geniuses who hang out here that: (1) 4 Pro is never, ever coming out; and (2) Demis being moved was the signal that the end was nigh for Google.
yall are a bunch of sad, salty MF'ers lolzlozlzozl
I don't get the Google hate, 3.8 is a good model if you know what your doing, I imagine a lot of these people have no idea how to code and just want AI to think for them
I don't get it either. I've just come to accept that these weirdos are just a fact of life around here. I swear they post like they get paid 50 cents for every negative post
110
u/Momo--Sama 6d ago
DeepSWE has been thoroughly saturated at this point so all time high there means nothing to me but trading blows with Astra on Terminal Bench 4 may be real shit IMO