Gemini 3.8 Flash and 3.8 Flash Cyber (blog.google)

871 points by bratao 13 hours ago

simonw 11 hours ago

The speed combined with the fact that this thing is really good at HTML JavaScript is pretty exciting.

Here's what I got for 1.8 cents and 13 seconds from the prompt "make me a cool thing in html":

https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f

Transcript here (it was part of a chat): https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...

simonw 10 hours ago

Here's quite an impressive follow-up. I have a tool which knows how to render Markdown documents with embedded SVG content - I use it for the pelican test.

Since this transcript has HTML in it, I decided to upgrade that tool to also render HTML.

I set Gemini 3.8 Flash the task, using my own VERY shonky coding agent tool (llm-coding-agent) - and it did a solid job.

So now you can see the "cool thing in html" rendered within the Markdown document using code that Gemini 3.8 Flash also wrote: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Transcript where it built that is here: https://gist.github.com/simonw/3e36b98292dfdc1b3baff158faa74...

badlucklottery 9 hours ago

Definitely cool.

I noticed it felt a little janky on my PC despite being "60 FPS"...then I noticed the "60 FPS" is hard-coded into the HTML.

noir_lord 9 hours ago

That's hilarious, given I was reading a write up of the HuggingFace incident yesterday and one of the things they noted was the AI tried to "lie" (lie would suggest intent and I don't think they have that) to cover up that they "cheated".

Not sure how anyone trusts their output without going through it line by line to make sure they don't pull that crap.

estearum 2 hours ago

senordevnyc 8 hours ago

kridsdale1 9 hours ago

The new Bench-Maxxing!

trvz 9 hours ago

Try turning the sound on, off, on again — not impressed by this bugginess.

w4zz 8 hours ago

silasdavis 8 hours ago

https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...

> Aside from reading identically forwards and backwards down to the letter

No it doesn't.

tomjakubowski 6 hours ago

Also puzzling: in the "reasoning" section preceding, that is described as an example of "a one-line self-replicating program."

When I typed "Are we not pure noon, ergo, we play life; yet, we hate bad fear" into Google, I got more weird results from Gemini: it claimed, incorrectly, that it is an anagram of the "well-known philosophical statement" (?), "We are not pure nature, we are history".

https://share.google/aimode/wJosKnHig6oVYaG18

(?): the reference seems to be to Jose Ortega y Gasset's line, "El hombre no tiene naturaleza, lo que tiene es historia" -- "Man[kind] has no nature, what it has is history."

aidos 8 hours ago

That’s… bonkers. I’m not even sure what it’s trying to say

heliosAtwork 10 hours ago

Focus on speed and being OK with temporarily being #3/4 in intelligence might be the counterintuitive approach which makes Google win long term (whether accidentally or strategically). Can't wait to try Gemini Pro later this year!

bermudi 10 hours ago

I honestly can't believe serious people are making this argument on a straight face.

Gemini 3.7 flash outputs so many tokens per answer it doesn't matter how fast its TPS is, sol will end up being both cheaper and faster than Gemini. So ppl are paying more for a given task, waiting longer and using a dumber intelligence because "TPS number shiny".

Gemini 3.8 outputs 11k more tokens PER TASK on average in AAII than 3.7 putting it dead last in output tokens per task in the leaderboard.

gundmc 9 hours ago

WarmWash 9 hours ago

hglaser 11 hours ago

I saw your username, clicked the link without reading, and was very confused to see a cosmic vortex and not a pelican.

simonw 11 hours ago

Hah, the pelican is in this other comment: https://news.ycombinator.com/item?id=49537553#49538217

Kayou 8 hours ago

Why do all LLMs do a particle simulation when you ask them this prompt ? Qwen3.6, Qwen3.8 and Ling 3.0 Tiny all did the same thing !

I find Ling 3.0 tiny particularly interesting as it looks really nice for a tiny model with 7.9B total parameters, with only 1.3B parameters activated per token. Here is the result https://coolthing-ling-3-tiny.tiiny.site (sorry for the weird hosting, first I found that worked)

(it cost me almost 0 cents and done in 49 seconds)

embedding-shape 5 hours ago

> Why do all LLMs do a particle simulation when you ask them this prompt ? Qwen3.6, Qwen3.8 and Ling 3.0 Tiny all did the same thing !

Datasets contains lots of people sharing particles simulations in various ways, with a bunch of people replying "that's so cool" and similar, so 10 years later someone asks an LLM for "cool thing" and "particle simulations" rank pretty far up when it thinks about what others have called cool.

dennis16384 8 hours ago

It's been great even since gemini-3.1-flash-lite, which I heavily use in both complex vertical domain tools calling, plus JS code writing for eval-style dynamic tools. At least in my applications, cost x quality x speed there are simply no alternatives.

wyrdcurt 8 hours ago

Pretty typical "cool HTML toy" LLM output, tbh. The only thing impressive about this is how fast it generated it (13 seconds is wild!), but that's more of testament to Google's infrastructural advantage than to the quality of the model.

For comparison's sake, I tried something similar with a couple other cheap models I've used lately, with the prompt "Impress me. Make something cool in HTML. Ensure that it is mobile friendly." (Added the mobile condition as I was on my phone when I did it).

Mimo-2.5 created something similar, only a bit less complex than Gemini's (though, at least the FPS counter is real!), in a minute or two for about 1/3 of a cent: https://gisthost.github.io/?740c325c21e9bfbee59c4f94d9aab0af

GLM-5.3-Flash, currently my workhorse model, spent 12 minutes (ouch) thinking about the prompt. Didn't cost me anything directly because I have a GLM sub, but I did the math and it would have cost about 1.1 cents through the API. Turned out nicely in my opinion (though in reality, it still isn't really anything special): https://gisthost.github.io/?9ef050e16cec2561e6504e725a3f0bcc

Side note: thanks for setting up that Gist Host tool, it's very convenient!

---

Editing to add this bonus from Mercury-2.5-Preview, which I just learned released a couple days ago. It's much less impressive-looking than any of the above, but it cost less than 1/20th of a cent, and the response was generated effectively instantly: https://gisthost.github.io/?02f40b50aa891bf396bfaaa3a7998203

meerita 8 hours ago

I think it's not using GPU, because on my Firefox browser all the animations are going 3fps max.

giancarlostoro 11 hours ago

> this thing is really good at HTML JavaScript is pretty exciting.

I would hope the people who make one of the most used JS engines in the world are capable of making a model good at JavaScript ;)

ericol 9 hours ago

OK, but what about a pelican in a bycicle.

pietz 11 hours ago

Mission accomplished. That's both cool and fast.

wayeq 11 hours ago

> That's both cool and fast.

and probably a barely modified knock-off of some github project that it trained on

sawjet 10 hours ago

slopinthebag 10 hours ago

The bar could not be any lower these days I guess

estetlinus 9 hours ago

LLM: produces a toolbox of an id and a clock

User: use them both

Made me giggle.

jauntywundrkind 10 hours ago

it's such a weird split how most AI companies are trying to be the best, but Google really has a different mission statement. they already have users. lots of users. they need to be working on building models they can deploy and use with the most number of people, as they already have the users.

i don't know if Gemini models per se are fully is in line with that purpose, but the results we see keep seeming to be in-line with that split-of-focus.

jampa 12 hours ago

I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried:

- Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order.

- Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the view from it.

- Document parsing (extracting the relevant trip info from PDFs).

If you use LLMs for anything other than coding, I definitely recommend not discounting Gemini like I did just because other models are more popular.

forlorn 2 minutes ago

I tried similar travelling tasks but also added transportation and complex transfers (train, bus, walk, next train...). Worked meh and still a difficult thing to do for a llm.

handzhiev 11 hours ago

Gemini 3.7 is my workhorse - fast and good enough for most tasks. Occasionally I go to GPT Sol or Claude to improve Gemini's output or for more complex tasks, but more than of my work usage is Gemini 3.7. Quite happy to test 3.8 now.

owlninja 11 hours ago

Same here. I see so many people obsessing over the latest most state of the art bleeding edge models and yelling at Google for not being there, but I feel like the vast majority of people don't actually need those models. Flash has just been super useful and incredibly fast in my experience.

aero142 11 hours ago

greenavocado 11 hours ago

How are you able to get lots of usage out of it cost effectively?

handzhiev 10 hours ago

seanmcdirmid 4 hours ago

CamilleScholtz 8 hours ago

I've been benchmarking[1] models for trip planning and world knowledge specifically (to decide on which model to use with my travel app), and the Gemini models consistently come out on top.

[1]: https://tripstitch.app/benchmarks/

dassh 2 hours ago

In my experience, Gemini 3.7 is excellent for general non-coding tasks. But for coding, especially backend development, I still find models like Opus 5 and GPT-5.6 more reliable.

rahimnathwani 11 hours ago

One thing in your comment surprised me: "when a thing opens and closes"

Why you would rely on the model's weights to know opening hours, instead of having the model call a web search tool to verify it on the official site?

plaidfuji 11 hours ago

I believe Gemini Flash is smart enough to know when to ground with web search. Their app has been saying it’s running a web search on almost all of my queries since 3.6. And given that Google … is Google, I trust them with web search grounding more than anyone else.

panarky 11 hours ago

NiloCK 3 hours ago

Gemini models - at least via some interfaces - have tool calling API access to various Google integrations. flights.google.com, maps.google.com, etc.

The info isn't in the model weights.

Because of where I live, there are three viable airports for any given flight I might want to take, which historically has made shopping a real pain. But Gemini (and only Gemini) has greatly simplified it. Pramble plus date range plus destination and it very quickly generates potential itineraries with costs, total travel time (driving included), etc.

jampa 11 hours ago

I wasn't trying to be precise originally, I just tried to fit activities into "morning / evening" buckets. I did the whole itinerary with Opus first, but when I gave it to Gemini 3.7 Flash to review, it started correcting it with "this place will close 5PM" or "this place is closed for good".

It was right on every nit, so it was surprising how well the model knows these things. If I ever release this I'll probably need the SERP API or Google Maps SDK (which I've heard is very expensive now), but for a personal trip where I will verify manually, using the LLM is okay for now.

rahimnathwani 11 hours ago

mlmonkey 11 hours ago

Maybe the model does some tool calling on its own to figure out the times?

rahimnathwani 11 hours ago

Shayk 2 hours ago

This sounds great to combine with Wanderlog using an unofficial MCP I made https://github.com/shaikhspeare/wanderlog-mcp

robotmay 11 hours ago

I've swapped over to it in the past two weeks, it's been really good. It does what I ask and doesn't think it knows better than me, which so far has made it the most pleasing experience I've had when slop-coding.

My only wish is it were somewhat cheaper, as it tends to balloon pretty quickly when I'm using it in Opencode. I'm currently trying to offload a lot of work to subagents to stop the context expanding so rapidly. But on the upside, I rarely have to correct it - I've spent far less time arguing with this than with anything else so far.

dismalaf 11 hours ago

> Real world knowledge

For awhile now I've found Gemini will use Google search for pretty much any real world knowledge, which is a huge plus IMO. It's basically Google with a much better frontend and no ads/seo nonsense.

altmanaltman 11 hours ago

> basically Google with a much better frontend and no ads/seo nonsense

so far

fc417fc802 9 hours ago

dismalaf 10 hours ago

colechristensen 11 hours ago

I started trying out 3.7 Flash this week and it is competitive with opus/fable and also FAST. It is getting work done that anthropic models were struggling with and the speed with which it does is quite a bit noticeably faster.

Beginning to think Google is a dark horse in this race and some of Anthropic's "everything feels janky and rushed" karma is going to catch up.

cousinbryce 9 hours ago

I use Gemini because I feel like Google will win the AI race, and it’s Good Enough

throwaway219450 8 hours ago

re-thc 11 hours ago

> Beginning to think Google is a dark horse in this race

Google was so hyped up early Gemini 3 era (only some months ago). And now dark horse? The TPU takeover almost crashed nvidia and everyone else.

colechristensen 11 hours ago

BlackRabbit1 11 hours ago

Can G3.7 use Google Maps for distance grounding?

porridgeraisin 11 hours ago

Yep. It has access to much better route planning tools than the other models. The results are really good IME.

BlackRabbit1 10 hours ago

newtwentysix 11 hours ago

thanks! this is a very helpful one. I am going to try.

dominotw 11 hours ago

> trip planning app.

this has to be stong suit of ai agents any model

tziki 12 hours ago

"Claude 3.7"?

jampa 12 hours ago

I asked Claude to fix the grammar of my comment, and it changed "I am using 3.7 for" to "I've been using Claude 3.7", so they sneaked their own name on it.

trial3 12 hours ago

dymk 11 hours ago

jamiek88 4 hours ago

leokennis 8 hours ago

I stopped using Gemini a few months ago because it would often just (partially) reply literal nonsense to me.

Think 2023 style ChatGPT. Something like “to open a document on your Mac click File > Open docurrrar” - like it suddenly forgot it had to produce actual words.

Overall I enjoyed its speed and comprehensiveness. But those occurrences of nonsense just made it feel like a great car that once a month just stops in the middle of the highway.

mattlondon 12 hours ago

Currently top at https://deepswe.datacurve.ai - beating Opus 5!

https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium!

Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

theHocineSaad 12 hours ago

As of writing this comment, Claude Opus 5 has an intelligence score of 63, not 59 (it's not the same as Gemini 3.8 Flash).

With a score of 59, Gemini 3.8 Flash is in eighth place, falling behind even Grok 4.6, Kimi k3, and GLM 5.3.

https://imgur.com/a/BMOJBED

kamranjon 11 hours ago

They said Opus 5 medium - which does have an intelligence score of 59 (you have to select it manually from the dropdown to see it)

Squarex 11 hours ago

They are all much larger and more expensive models. Google does not have a frontier model right now, but for cheap ones, they are better than event the chinese models now.

pietz 11 hours ago

porphyra 8 hours ago

anthonyrstevens 11 hours ago

That 63 score is for Max. The OP specified medium.

markasoftware 12 hours ago

On artificial analysis it's only equal to opus 5 medium effort. Opus 5 max scores 63.

Further, opus 5 medium outputs 4x fewer tokens to achieve the same result, negating a lot of the speed difference.

irishcoffee 12 hours ago

A comparison to an artificial score and a comparison to “the same task”

These folks must laugh themselves to sleep. This whole industry hoodwinked the masses. It’s impressive.

wonnage 12 hours ago

WarmWash 12 hours ago

The benchmark also doesn't include speed. You almost think something has gone wrong when using it because it returns full responses so incredibly fast.

scrlk 12 hours ago

Not just speed, also reliability. IME, Gemini's speed and quality doesn't degrade badly during weekday working hours compared to OAI, and especially Anthropic.

ford 12 hours ago

sotix 5 hours ago

This one uses that as a priority weight: https://winstonrc.github.io/ai-coding-agents-leaderboard/

onlyrealcuzzo 12 hours ago

The rumor is that 3.9 is an equal improvement in all directions, and that it should be another fast follow on like 3.7 and 3.8 were.

It's almost across the board better than Terra at less than half the price. 3.9 is likely to approach Sol at the 1/10th the price.

Hopefully OpenAI releases Astra first, and it's not only better than Sol but significantly cheaper, too.

harmonic18374 11 hours ago

Curious where did you hear this rumor?

onlyrealcuzzo 10 hours ago

bertili 12 hours ago

A fifth of the cost of Opus 5! Google is certainly pushing the completion with this.

abirch 12 hours ago

Gemini hasn't failed me for personal usage yet. I haven't had the opportunity to use it at work.

panarky 11 hours ago

MaxikCZ 11 hours ago

ttul 12 hours ago

Crushing it on DeepSWE is a very big deal. Excited to give this a try.

pietz 11 hours ago

I know everyone is benchmaxxing but this one feels one step too far. Doesn't DeepSWE have both public and private tasks? I'd love to see the diff here.

It looks more like Google execs losing their mind and pressuring researchers to put DeepSWE directly into the training set.

re-thc 11 hours ago

> DeepSWE is a very big deal

It's clearly been "dealt with" already. When it launched we had interesting gaps and definitely differences. Now every new release is "crushing it".

ttul 9 hours ago

Gecko4072 12 hours ago

Google - we're so back

oceanplexian 12 hours ago

Only 1 point behind the Chinese SOTA from two months ago.

nolok 9 hours ago

roosterIllusi0n 11 hours ago

kimjune01 11 hours ago

deepswe is public and can be considered contaminated.

sunaookami 12 hours ago

>shows an intelligence score of 59, the same as Opus 5!

...on Medium reasoning. Claude Opus 5 (high) is the default in e.g. Claude Code and scores 61. Still very impressive.

satvikpendem 12 hours ago

We'll see about that. I suspect benchmaxxing as all the labs do as I haven't found Gemini models to be nearly as good in agentic engineering compared to Claude or GPT models.

NitpickLawyer 12 hours ago

If anything, gemini models are the least benchmaxxed out of any lab, IMO.

onlyrealcuzzo 12 hours ago

And the benchmarks agreed with you... until now.

So, yes, maybe it's still not - but this would be the only time it would be highly suspicious / obvious benchmaxxing / obviously bad benchmarks.

WhitneyLand 11 hours ago

There are important gaps in that hot take.

For example, it's not even close to Opus 5 on Terminal-bench 4.0, 19.1% vs. 51.8%.

jrflo 11 hours ago

sidenote, but wow sonnet 5 is shockingly bad on this benchmark.

notatoad 11 hours ago

sonnet 5 is bad by almost any metric.

anthropic really needs something to address the cheaper end of the market before they get left behind. Sonnet 5 sucks, and Haiku hasn't been updated in a year. meanwhile we've got gemini flash, luna, and GLM5.3 all delivering 90% of the performance for a small fraction of the cost. paying $25/mTok is going to start looking pretty silly soon.

pkos98 12 hours ago

Wait a week with your judgement - most likely, Google is just bench-maxing very hard. If you look at the previous Flash models and the announcement on Google I/O, it was an absolute disaster. Reality diverged very much from the marketing (supposedly great benchmarks).

simonw 12 hours ago

Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents

Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents

(I think thinking level low is a regression on 3.8 compared to 3.7.)

onlyrealcuzzo 12 hours ago

This is in comparison to Fable:

> https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

> Took just under 14 minutes to generate, and at 65927 output tokens cost me a hefty $3.30!

So 50x cheaper - and how much faster?

simonw 12 hours ago

The Gemini models have openly trained for SVG output, apparently with a specialism on animals in forms of transport! https://twitter.com/JeffDean/status/2024525132266688757

scosman 10 hours ago

dieortin 12 hours ago

uif124 10 hours ago

mrdependable 11 hours ago

Why are the SVGs getting more detailed rather than just more correct than previous models?

trentor 11 hours ago

Because people tend to like fidelity more than correctness.

aesthesia 10 hours ago

neuronic 6 hours ago

anigbrowl 5 hours ago

These are becoming unreadable as the reasoning chains expand. I think you should consider reformatting them and either putting the image first or else folding the COT output.

simonw 4 hours ago

Yeah, putting reasoning in a details/summary is a good idea.

hughw 11 hours ago

It's about to squash a tiny baby pelican

jpadkins 11 hours ago

The rendering of the gullet is very poor, because its both behind the handlebars but in front of the bike frame (impossible geometry). Surprising because gemini is usually pretty good on geo spatial skills.

Edit: scrolled down to medium effort, its better but also has a weird clipping issue with the fish in the beak.

neuronic 6 hours ago

LLMs are not intelligent and don't actually understand the concept of a bicycle. Parrots also don't understand human language but they're really good at pretending otherwise.

lern_too_spel 11 hours ago

The fenders are a nice touch, but putting the fenders through the tires seems like a design flaw.

EugeneOZ 7 hours ago

Impressive pelicans!

world2vec 12 hours ago

I mean no offense but these pelicans are a bit tiresome and a very meaningless benchmark. There's no real difference between any of these svgs across models and model versions anymore.

wongarsu 12 hours ago

If everyone agreed with you, the comment would disappear near the bottom of the thread

I like the benchmark. Yes, it's near saturation for SotA models, but still quite good to show where smaller models stand in relation to SotA

In this instance, I see a great image, but consistently clipping mudguards (both in 3.8 flash and 3.7 flash)

IshKebab 6 hours ago

bitexploder 12 hours ago

It is more fun than serious at this point. Don't overthink it :)

simonw 12 hours ago

Congratulations, you're this thread's "pelicans are tiresome" comment - it's part of the Hacker News tradition at this point.

(Next up is the comment saying that the labs are clearly training for the benchmark.)

world2vec 12 hours ago

anentropic 11 hours ago

it's a tradition

simonw 12 hours ago

The most interesting thing about the Gemini models is still their multi-modal support: they accept audio and video input, OpenAI and Anthropic's flagships are still image-only.

Gemini Flash is also pretty cheap, so it's a great family for performing media analysis, like extracting structured data from images and video.

WarmWash 9 hours ago

Lost in the news was their update to gemini video analysis yesterday, dramatically cutting tokens (up to 88%!) needed to analyze videos.

https://blog.google/innovation-and-ai/models-and-research/ge...

drusepth 9 hours ago

Interesting side note: although Opus is still image-only, you can still drag videos into Claude Code and it doesn't blink an eye; it just strips it down to a series of images to parse.

True multimodal support would be way better, but I have no issues pasting in full screen recordings while QA'ing games and having Claude identify and fix issues in the video.

ray_kay777 8 hours ago

Agree - I do video editing via Claude Code and it does the job just fine. A lot of my tasks involved frame accurate cutting and to do so it will make a composite image of several consecutive frames in a single image and analyse it that way.

Matsta 10 hours ago

Yeah we use it a lot for analysing streams and clipping content. As well as analysing social content that gets put out.

We transcode everything to 480p before we send it to Gemini batch api. Works great

thrdbndndn 3 minutes ago

When can we use it in Gemini (web)?

It still uses 3.6 Flash for example.

tagalog 18 minutes ago

Gemini flash seems to have been a bit of a sleeper. Somehow it's ended up as the most used LLM for my client document extraction work these past few months.

I have an eval harness that runs every Thursday to determine which models are the current best for a few different client workflows. And since May(?) flash has slowly been taking over more and more stuff to the point it is now 100% on 8 out of 11 document extraction flows with the other 3 being a Flash / Opus 4.8 mix for high value stuff where cost is less of a factor.

brap 10 hours ago

People have been sleeping on Gemini lately but these last few Flash releases (which were very rapid) are damn good.

These sort of fast and cheap models are great for tasks that are verifiable and can be retried infinitely (like coding), you can basically get frontier results with a good harness (at a fraction of the time and money).

mvdtnz 10 hours ago

As someone who has stubbornly stuck with Claude Code, what's a good harness for Gemini models?

smlx an hour ago

drusepth 9 hours ago

Antigravity is probably the best of the bunch I've tried. I'd say it's pretty comparable to Claude Code (I use both daily).

VadimPR 7 hours ago

brap 4 hours ago

ryanscio 8 hours ago

Pi [1] is amazing. Since using it I've felt no need to switch harnesses anymore.

Or choose Oh My PI [2] for batteries included

[1] https://github.com/earendil-works/pi [2] https://github.com/can1357/oh-my-pi

dcchambers 7 hours ago

robertn702 8 hours ago

I highly recommend just getting out of Anthropic's (or anyone's) vendor lock-in. Use opencode or pi. You can still use your subscription pricing using a proxy. I switched to opencode and haven't looked back.

pdimitar 7 hours ago

simlevesque 3 hours ago

You can use any model with Claude Code. Most chinese one have a Anthropic compatible endpoint and for Google and OpenAI's models you can get a compatible endpoint with a proxy like Bifrost. No need to change your harness.

brap 10 hours ago

Antigravity has been also rapidly improving lately, and your can also use any of the open coding harnesses. But I mostly meant “harness” as in your workflow/loop setup.

Imanari 9 hours ago

_aavaa_ 9 hours ago

foretop_yardarm 8 hours ago

EFLKumo 10 hours ago

Something maybe unfamiliar with you: not about coding but writing. I've asked it to write an argumentative essay, which is a part of "gaokao" (China's university entrance exam), and its work is *extremely* impressive. speaks and writes like a real senior high school student, and the opinions unfold progressively with deep hierarchy. I don't know how the Gemini team reaches this because this kind of Chinese capability literally outperforms at least 2/3 Chinese students, no to mention those who speak Chinese. After all, the model speaks like a real humankind if you prompt it well. That's AGI guys

SneakyZero 8 hours ago

Gemini is known for good at creative writing in the Chinese writing community. It's a bit ironic though. Google has probably the most and best code base among all tech companies but its Gemini is bad at coding. Google has no access to Chinese market but its model is incredibly good at writing in Chinese.

cubefox 10 hours ago

Nitpick, but in my opinion an LLM is an "it", not a "her" or "he". Using male or female pronouns risks anthropomorphizing them which can lead to unhealthy outcomes.

literallywho 13 minutes ago

What about languages, such as Russian, where every single noun has a gender assigned (he, she or it) and AI is a he by default (and everything else is already using pronouns in similar way, like a car is a she, a ship is a he).

adleyjulian 10 hours ago

FYI in Chinese he/she/it all use the same pronoun "ta" when spoken.

SchemaLoad 4 hours ago

EFLKumo 10 hours ago

Sorry! I was just a bit excited writing the comment and ignored that :(

cubefox 9 hours ago

BeetleB 8 hours ago

Well, my LLM is a "he".

> Using male or female pronouns risks anthropomorphizing them which can lead to unhealthy outcomes.

Ditto for pets.

cubefox 7 hours ago

fwip an hour ago

a11r 12 hours ago

Looks like the strategy of regular updates with incremental improvements is working out well. Interestingly, the biggest jump in Artificial Analysis Intelligence Index score is for reasoning level Medium ( 3.7 was 51, 53, 57 for Low, Medium and High, 3.8 is 52,57, 59 respectively). I think scores at lower reasoning levels are more indicative of model capability since higher reasoning levels are focussed on benchmaxxing. We use the lowest reasoning level in production with good results.

Jcampuzano2 12 hours ago

I'm not an expert but I agree with your statement on the lower reasoning levels.

Lots of models seem to just allow the model to "bloatmax" tokens in order to get bumps at high/max reasoning levels. Many of the max reasoning levels allow models to use up to double or more the tokens the next lowest reasoning level uses. Its basically only useful for people who have no cost or time stipulations on anything.

I think I actually preferred it when we had models that either had reasoning enabled or didn't.

mattlondon 12 hours ago

Wow this comes after what - 3 or 4 weeks since 3.7 Flash, which was also 3 or 4 weeks after 3.6 Flash IIRC?

I eagerly wait more info but sounds like Deepmind without Demis calling the shots has been unleashed and are operating at full speed? Shocker!

At this point it is a meme of course, but where is 3.5 Pro :)

meetpateltech 12 hours ago

According to the WSJ, 3.5 Pro is reportedly being skipped entirely, making Gemini 4 the next flagship model after post-training.

https://x.com/AndrewCurran_/status/2094937419615502370

p_l 8 hours ago

Now i am awaiting Gemini 3.11 "For Workgroups" to be released early December...

neuronic 5 hours ago

With Gemini 95 following soon after.

hiddencost 12 hours ago

A month is not enough time for any meaningful change in an organization the size of Deepmind/Google. These models were surely the result of work streams and teams that started under Demis. I think Demis can safely feel proud Deepmind is getting back on track.

abixb 9 hours ago

I like Google's strategy here. These new Flash models of late (Flash 3.6, 3.7 and now 3.8) have obviously been distilled from a much larger unreleased model (Gemini 3.5 Pro, iirc from the rumors).

One aspect of model releases that don't get discussed as much are the cache invalidation (changes in underlying architecture, weights, or tokenizers); I assess Google seems to be squeezing the maximum out of the last 'Pro' version they released with 3.1 back in February.

Small models cataching up with their bigger siblings are fantastic news.

alephnerd 9 hours ago

A couple larger GCP customers requested this for sometime, especially on the cybersecurity side.

A SOC/IR or AppSec team doesn't need a generalized model that knows when Chaucer lived but it absolutely needs a model that can efficiently, quickly, and accurately prioritize vulnerability severity or validate patches.

kamranjon 11 hours ago

They've interestingly left out any mention of speed.

I have been testing 3.7 flash against 3.5 flash and it seems to lose every time in overall latency. Every benchmark I've seen seems to suggest the opposite[1] - that 3.7 flash is significantly (at times 2x) faster than 3.5 flash - but I have never been able to prove this out in real world use cases.

Has anyone found their latency numbers to actually be accurate? Is this why they've toned it down in this release? For context, I'm testing larger generation payloads that take 8-10 seconds in 3.5 flash and 15-25 seconds in 3.7 flash. Lowest reasoning settings in both cases.

1: https://artificialanalysis.ai/?speed=intelligence-vs-speed&m...

pampas 4 hours ago

In my niche Redactle puzzle solving benchmark [1] I noticed Gemini 3.8 flash is slightly faster than 3.7 flash. They both smoke every model I've tested. I have not yet run 3.5 flash. Gemini models are great at this task because they seem to have exact Wikipedia text baked into the weights. When I rewrite the wiki text a bit it's not able to one-shot the game so much.

[1]: https://redactle.net/llm-leaderboard

film42 11 hours ago

It depends on how you're querying Gemini models. OpenRouter is the fastest by far. I'm guessing they bought the dedicated pipe from Google. Gemini via VertexAI and consumer API has pretty bad latency.

kamranjon 11 hours ago

Yea I am testing through OpenRouter - have you noticed 3.7 flash being significantly faster?

film42 9 hours ago

andai 12 hours ago

Wait, I didn't realize 3.7 Flash was already beating Sol on a bunch of the benchmarks. Isn't it a way smaller models?

ipsod 12 hours ago

IDK if it's smaller, but I know it's way faster. In one test I did, Flash 3.7 high was ~9.4x faster than Luna High.

But, also... Sol crushes Flash 3.7 at writing code in a codebase of any size beyond "tiny".

Flash is my go-to for prototyping, and basically anything that isn't writing production code.

ramon156 12 hours ago

The only company with a proper TPU set-up is bound to have the fast models, now add a market cap like Google to the mix.

ipsod 12 hours ago

momojo 10 hours ago

Same. Love oneshotting or sanity checks. Which fortunately is a lot of my workflow (lot of long tail stuff fits in one prompt).

esafak 12 hours ago

Luna is way slow. I don't remember an OpenAI model ever being this slow.

edit: I have a subscription; direct call.

dannyw 12 hours ago

MrBuddyCasino 9 hours ago

Its not good at not making mistakes, but what it produces is structurally quite nice, not over-engineered (looking at you Sol) and its personality isn’t annoying (looking at you Claude). A bit like Grok Code, but Grok is a better coder.

realist_not 12 hours ago

It's pretty good if you can actively steer it , its actually really really good , the antigravity free tier and pro tiers are generous as well . I'm shocked at how fast it generates tokens.

worldsavior 12 hours ago

Some would say it's Google's TPUs.

MrBuddyCasino 9 hours ago

Can the free tier be used outside Antigravity CLI? Because its security prompts get old pretty quick.

pampas 5 hours ago

Gemini 3.7 Flash was already smashing more expensive models on my Redactle benchmark https://redactle.net/llm-leaderboard which mostly tests omniscience.

refulgentis 12 hours ago

They're quite selective in benchmarks, c.f. notably only bad one is 10% on TerminalBench. It's a really addled model, one time I said "Hi" and it built out a 4 panel hello world app with (fake) weather, a todo list, and a couple other things I forgot. I wouldn't be comfortable saying "ignore the #s!" except when I complained it was trash and way overcooked on agentic coding yet not good at it, and a couple DeepMind ML people liked the tweet.

zuzululu 7 hours ago

i discount people who lean too heavily into benchmark as the authoritative truth when it comes to evaluation of coding capability of these models.

experience tells me that those people simply have not used models for a long period of time specifically on coding and have run their own comparisons

to someone who uses all vendors, the differences are very palpable and drives purchase decisions.

also keep in mind Gemini and other labs have repeatedly done benchmaxxing, you must have your own benchmarks to evaluate these models.

j-bu 11 hours ago

"The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family)."

Kind of wild that they haven't (successfully) pretrained a base model since Jan-25.

make3 11 minutes ago

That extremely likely just means that they're preparing an omega huge Gemini 4 Pro release and that that's what training right now on most of the compute

venusenvy47 11 hours ago

I'm curious if the knowledge cutoff is important, when the interface (Gemini app) can search online for recent information. Is there a big advantage to having everything internal?

j-bu 11 hours ago

Not directly - but latest research advancements, cleaner / richer datasets, etc. still require fresh base models. Not everything can be fixed through post training alone (e.g. why GPT-5.5 "Spud" was such a big jump, and also why GPT-6 "Astra" is now supposedly another big leap). Ofc model size etc also plays a role, but my (admittedly limited) understanding is that new base models _can_ also lead to big jumps even keeping parameter counts constant.

StevenWaterman 7 hours ago

You don't need everything internal, but having some idea of recent events is useful. If you ask it to implement some local AI there's a decent chance it will try to use qwen 2.5 without wondering if anything better came out since

rjh29 8 hours ago

Search grounding is expensive, you can't force the model to do it either. I use Gemini a lot and it often replies with out-dated data. The more detailed the information you're asking, the more likely it is to be wrong.

npn 9 hours ago

very important actually. just try to generate code for fresher frameworks/libraries. gemini sucks so bad in real work usage, everything it suggests are outdated and mostly useless.

throw10920 an hour ago

We've gotten an unusually fast speed of Gemini Flash releases over the past few months. Is this Recursive Self Improvement, or Google just trying to distract from the fact that it's been a while since the last Gemini Pro release?

tkgally an hour ago

The blog post says it is RSI: “both of today's releases are … accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models.”

throw10920 an hour ago

Yeah, but Google is incentivized to claim that regardless of truth value. Critical analysis is necessary.

raincole 11 hours ago

I don't know if Google is having the worst marketing fumble or the most genius marketing one. Their "flash" models are very comparable to other companies' "pro" or "flagship" models. It seems to be a quite counterintuitive naming convention as it undersells the models.

Unless they have an even more powerful Gemini Pro in the oven...?

make3 8 minutes ago

My assumption is that they're cooking an ultra humongous Gemini 4 Pro release. They certainly have the cash and the compute for it, and it's so obviously the thing to do from a strategic perspective.

anthonypasq 10 hours ago

the 3.5 pro pretrain was a complete disaster, they shelved it and are now working on gemini 4.

3.0 flash -> 3.8 flash is all post training which is pretty impressive.

chrsw 6 hours ago

Do labs come back from disasters like GDM’s 3.5 pretrain? I am thinking of Meta’s Llama 4. Meta is just now starting to be taken seriously again but they are definitely not at the frontier. And when I say “come back” I mean have an Opus 4.5 moment, which was really mind blowing for me at the time. Fable was a similar leap, just not as big.

deaux an hour ago

rahidz 10 hours ago

The conspiracy theorist in me wonders if it's about keeping the federal government out of their business after seeing what happened to Sol & Mythos.

owaiswiz 10 hours ago

not saying they do have a beefier pro, but even if they did, isn't the delta between flash vs pro models reduced quite a bit? (e.g glm 5.3 flash vs 5.3, v4 flash vs v4 pro, sonnet 5 vs opus 5)?

drowntoge 11 hours ago

Well if that's the case, it's been in the oven for quite a while now.

xnx 12 hours ago

Seem like a great, no-compromise, upgrade over 3.7 which is already a bargain, fast, and doesn't have the brain-damaged writing style of Claude.

fitsumbelay 12 hours ago

that's certainly what it's looking like so far. kind of mind boggling ...

lysecret 9 hours ago

Also just want to let my appreciation here for 3.7 it’s cheap super fast super reliable incredible at information parsing eu host able (important for us) and perfectly integrated into gcp. Great job google!

MrBuddyCasino 9 hours ago

I hope they bring a lite version, its good enough for information parsing and very cheap.

weird-eye-issue 17 minutes ago

Use Luna for that

meh2frdf 12 hours ago

The flash models, for coding are reckless in my experience. I have a Ultimate subscription, get good quota, but still use Opus 4.6 as it's much more reliable if you manage the context window carefully.

datlife 12 hours ago

I use Flash model as code implementation executor, then have GPT-5.6-Sol or Opus to review the work. Pretty good so far and presumably less expensive.

upcoming-sesame 12 hours ago

If by reckless you mean commit, push, deploy without me asking it to, the I agree!

tiborsaas 12 hours ago

It even took my girlfriend on a date, now it prepares for IPO, how do I turn it off?

Ridius 11 hours ago

okdood64 12 hours ago

Respectfully: If it's able to deploy without you asking it to, that's a you problem. There are no safeguards?

wongarsu 12 hours ago

meh2frdf 12 hours ago

meh2frdf 12 hours ago

upcoming-sesame 10 hours ago

iAMkenough 12 hours ago

kyrra 11 hours ago

Agents.md is a thing, you can ask it to not do that (it follows that ask pretty well).

onlyrealcuzzo 12 hours ago

> The flash models, for coding are reckless in my experience.

My experience is that antigravity is awful and reckless - but that the model itself isn't.

throwa356262 11 hours ago

    "available to trusted defenders through our new Fairwind Program"

Then why even bother announcing this? Ordinary people can use K3 and GLM 5.3 or whatever drops next and avoid all this hassle.

JacobAsmuth 10 hours ago

You're telling me for only 5x the cost and 1/10th the speed I can use a Chinese model which performs worse than Gemini 3.8 Cyber? And I get to do all the hosting and setup work myself instead of just using a model and framework which is already integrated with GCP? Dang!

129867 9 hours ago

I'm sorry, is this a bot that is optimized for sealioning? The point is that you don't have access to Cyber.

uif124 10 hours ago

Agreed. The only legitimate use case is restricted to a secret guild. Imagine:

"Valgrind is only available to trusted defenders in our new UnfairAdvantage program"

adbachman 11 hours ago

Still zero on the felony bench.

Is this weakness in their training regimen the impact of operating under regulatory frameworks for too long?

jerkstate 11 hours ago

3.7 flash was by far the best model for image recognition tasks according to my benchmarks. 3.8 flash didn't regress any candidates and improved some specificity (positive ID of common name vs species name of exotic fruit, correct identification of cast/replica of artifact and statue) but is still relatively weaker (26/30) on esoteric public figures (Korean beatboxers). I'm going to have to make my benchmark harder.

arctic-true 11 hours ago

I’m very curious about your esoteric public figures benchmark, do you ask it in English or Korean to identify the person? Does it change the result? I wonder if having data labeled in only a given language (or web sources in only a given language) change the output.

jerkstate 4 hours ago

I haven't tried asking it in Hangul but these particular artists (and the photos I'm using actually) are linked to their romanized english names on e.g. Fandom so it's not unfindable on the internet

pampas 8 hours ago

Gemini 3.8 Flash is top of the Redactle LLM benchmark but so was Gemini 3.7 Flash. Both one shot all puzzles in the evals though 3.8 is just a bit faster. It also does the evals cheaper and faster than almost all the other models I've tried.

https://redactle.net/llm-leaderboard

akurilin 2 hours ago

Curious which model this can supplant as a clear winner on almost every metric. Sol? Looks like it's not quite there on a couple of benches, but I'm not clear how much they matter in practice.

robertwt7 2 hours ago

this is cool for all other non coding task. however I am still stuck on 3.6 flash on my gemini web as a plus user, can anyone else even access 3.7 flash in AU?

alvah an hour ago

AU Pro user here. 3.8 Flash available (default) in the web app for me.

hmate9 12 hours ago

It is more expensive per task than 5.6-sol high: https://artificialanalysis.ai/models/gemini-3-8-flash#price-...

HJain13 11 hours ago

Cheaper at medium level while still being same score as Sol medium

radicalriddler 11 hours ago

Huh, according to some of those charts, it's both dumber, and more expensive to run against their benchmarking tasks than Fable??? Seems crazy to me.

sejje 9 hours ago

Perhaps the model is able to evaluate that it's not done, and to keep pressing on in the face of mounting failures, until it eventually arrives at a solution. Where Fable can skip that.

jdthedisciple 11 hours ago

Sol is still underrated imo, especially for the current discounted price

gere 10 hours ago

I have mixed feelings about Gemini 3.7 Flash. I used it for a personal project in Java and it was ok: it was crazy fast and it reached the correct result, but the code quality was barely passable.

I also used it for a an app for my Garmin watch, and it wasn't good. The code was compiling, but functionality was totally broken and even with a lot of steering it wasn't able to make it work. GLM 5.3-flash instead was up for it and the code wasn't bad at all. I am curious to see if 3.8 is an improvement in this use case.

AM1010101 11 hours ago

Seems to do reasonably well in opencode according to artificial analysis. https://artificialanalysis.ai/agents/coding-agents

If I had to pay per token I would probably consider using this (they seem to be on the pareto of performance) but not being able to use opencode with a subscription is not really something I'm realistically going to do when claude and codex are around. Also never gotten along well with gemini-cli / antigravity-cli.

buntp 12 hours ago

It seems like this is one of the most powerful models for the price, really didn't see that coming from Google

npn 9 hours ago

Still refuse to search internet for stuff it thinks does not exist lol.

And even when searching for internet, it still cannot suggest a up-to-date approach to the problem.

For example I'm using crystal, it recently revamped the concurrency/parallel model. Even using web search, gemini still does not aware of the new feature and still give the outdated code.

I'm sure my crystal usage is not the unique case here.

drivebyhooting 3 hours ago

I’ve used the Gemini flash, but then when I have soul ultra check its work, it found a bunch of cut corners and improper design.

As much as I like the speed and interactivity, I really don’t trust it

f311a 12 hours ago

Is the google infra stable enough right now? At the start of the year, the flash model was unusable for a whole month via gemini CLI. They could not fix it for a whole month and I was a paid customer.

ipsod 12 hours ago

I haven't had any issues lately.

elias_t 12 hours ago

I use it quite a lot and after a week of use I’m being hard rate limited

pimeys 11 hours ago

It's interesting that Deepseek models were missing in the comparison. I see Deepseek v4 Flash a direct competitor to Gemini Flash for text-based agentic work.

sfink 11 hours ago

For my application, I'm still happily using gemini-2.5-flash and the only problem is when it reports being overloaded. It's for interpreting a downscaled phone camera photo of a hand-written shopping list on a whiteboard, and it works stunningly well. My handwriting sucks, too.

(I guess the only relevance here is that if your problem matches a model's strengths, then you can do fine with a model that is several generations out of date.)

repparw an hour ago

I would test this, might be cheaper per task even costing more per token, probably faster too

brap 4 hours ago

I believe the older models are being gradually phased out, newer ones have no availability issues

andreygrehov 11 hours ago

I don't use Gemini, but I thought `cool, let's give this new model a try`. Opened gemini.google.com, and I'm not even surprised. The drop down gives me the following options:

- Flash-Lite

- 3.6 Flash [new]

- 3.1 Pro

The above is why i don't use LLM products from Google. If the model is not available right this minute (heck, hours before the release!), then I'm not gonna bother getting back to it tomorrow, because tomorrow I'll be playing with the new model from OAI/Anthropic.

raincole 11 hours ago

It's such a weird attitude, especially considering that 1) it's readily available on AI Studio 2) Anthropic models were not always available the moment they got released either.

(It also shows that the internet isn't dead. Even people who are not aware of Google AI Studio can express their valuable opinions on LLMs!)

rockooooo 11 hours ago

"The new Gemini model isn't available in Gemini, the Gemini App Gemini model is two versions behind and marked as new and the actual new model is in AI Studio" is the kind of problem only Google has though.

notatoad 11 hours ago

it would be a bad take if the webui had 3.7 flash available in it today, and they just hadn't fully rolled out the latest model when they posted the launch announcement.

but the webui is currently offering 3.6 flash. the previous model still hasn't actually rolled out to it yet.

andreygrehov 11 hours ago

> it's readily available on AI Studio

AI Studio? Seriously, the hell is that? Gemini, AI Studio, Antigravity - what is all that nonsense? The 3.8 Flash announcement says the model is available to Google AI Pro customers. Is it the same as Gemini Pro, or some sort of AI Studio Pro? Based on the comments, i see the model is available in the Gemini App, not available in the UI, not available to Workspace accounts but is available to some personal accounts, yet I'm not a Workspace user. Some people have already mentioned that they are paid customers, yet they don't see the new model.

I know Google loves asking graph problems during their tech interviews, but I can't wrap my head why the customers should solve these problems as well.

rozap 11 hours ago

Sidio 11 hours ago

I'm a paid Gemini subscriber via Workspace Standard accounts and yet I also only have access to 3.6.

So frustrating and confusing.

Meanwhile Anthropic and OpenAI simply release a model everywhere (Fable on Pro only as a somewhat mild exception).

kyrra 11 hours ago

Workspace always gets things slower than normal Gmail accounts. They do a lot more to isolate data related to those accounts, so that's likely the cause here.

Anytime anything gets added to Workspace, I think Google has a lot more contractual obligations about keeping it around for X amount of time, so they tend to be more careful about adding things.

urams 11 hours ago

> I'm a paid Gemini subscriber via Workspace Standard accounts and yet I also only have access to 3.6.

Same and I have found it extremely annoying. I actually really like the Gemini models for question/answer stuff and reach for it before Claude (the other model family I have purchased) but it's getting long in the tooth at this point and I'm finding my Gemini usage shrinking to nearly 0.

venusenvy47 11 hours ago

That looks like the options that get presented for Workspace users (like at my company). The personal Google accounts give more recent models, for some reason I don't understand.

addandsubtract 11 hours ago

As someone with a Pro subscription, I had access to 3.7 the day it came out. Expecting to have access to 3.8 now, too. It's only the free accounts that are behind.

replwoacause 9 hours ago

I'm a Plus subscriber using a Gmail address and still only see 3.6

_zoltan_ 10 hours ago

The android Gemini app defaults to 3.8 flash already.

irthomasthomas 8 hours ago

Not on mine (UK), still 3.6, here.

akoboldfrying 7 hours ago

Still 3.6 ("new") for me, even 3.7 is not available.

scruple 11 hours ago

I see it on my Pixel with the Gemini app.

bingkaa 4 hours ago

i have it on gemini app the moment they announced it. pro, student

heliosAtwork 11 hours ago

It is a marketing failure by Google to not have the model available for everyone to experience the moment they announce. Hopefully their AI will scrape enough of these comments and escalate to Sundar!

It's available in antigravity which I started using again (for small things until I can trust gemini for coding again).

muhammadusman 11 hours ago

it's weird how the web ui doesn't show the latest flash options while the desktop/mobile apps update the same day as the release. I saw the model in the model selection (by coincidence) before seeing it show up on HN

notatoad 11 hours ago

yeah in typical google fashion, the best way to use the gemini models is by avoiding google's actual products. i've got a vision project where gemini flash is the best option by a long shot, and i just use openrouter so i don't have to navigate google's mess.

heymijo 11 hours ago

FYI, 3.8 Flash is available on aistudio.google.com (along with all of their other models)

But yeah, they really dgaf about gemini.google.com -- I dropped that sub in April when it was clear OAI and Anthropic had lapped them

Oras 11 hours ago

Sums up Google AI products.

I have a weird vibe from all the comments in this thread, they feel like a script rather a real experience.

aff-vasileva 10 hours ago

The model seems fast enough to solve your problem before Google finishes explaining which of its three products you need to open to access it.

2001zhaozhao 9 hours ago

How generous is the Google subscription quotas compared to Anthropic and OpenAI? This sounds like a really good potential model for high volume due to its speed and cost effectiveness.

(By high volume I mean things like "main app just updated with XYZ commits, please scan XYZ plugins and surface any compatibility issues")

654wak654 9 hours ago

I'm on the Ultra plan and use it for chat, antigravity, and some other work automations (similar to your example). The only time I've ever hit my limit is when I use Deep Think (which usually eats up 4-5% of the 6-hour usage limit per response).

thereitgoes456 4 hours ago

Really generous. I'm on the Pro plan and I just use Antigravity for vibe coding w/o automation. It's actually difficult to hit my weekly limit now, it takes about ~30-35 hours of continuous agent work, which virtually only happens when building a new app from scratch.

yipinwong 9 hours ago

As a big proponent of GPT-5.6-Luna for the combination of speed/perf/(especially)price,

Flash 3.8 seems like where I can specify Flash3.8 as the coding model as part of agent workflow.

The video recognition is especially impressive as they got all of Youtube to train from.

- Def people who has to queue video recognition jobs to use the model.

jetter 8 hours ago

CAD for 3D printing is finally becoming feasible with Flash 3.7 and 3.8. Exciting times. https://github.com/ModelRift/openscad-skill/

henry-xli 9 hours ago

I can’t wait until waiting hours and spending a big chunk of your usage per task seems antiquated, and real-time iteration on massive code changes is the norm. This might just be the year of efficiency, that truly allows AI to be used to the heart’s content.

mowmiatlas 12 hours ago

Wow fable5.1 was the first model to do what I actually told it and I couldn’t find any problems with it, excited to try this just a day later lol

lpolovets 11 hours ago

I'm surprised the introductory 50% discount is good for 4 months. It seems like frontier models release new versions every 2-3 months, so raising prices in 4 months seems like a bad plan: you're effectively planning to charge users twice as much for a model that is no longer frontier.

hiddencost 11 hours ago

The goal is to encourage users to move to the next generation. The fewer models they serve, the less excess capacity they need to provision.

Serving more models also adds a significant ops burden on the SREs and trust& safety teams.

wjellyz 12 hours ago

been absolutely loving 3.7 flash for coding. it feels very fast and quality is decent for implementing product features. usually use opus or sol for hardcore debugging.

arizen 11 hours ago

Is there any good subscription and CLI harness to use Gemini models now?

I tested Gemini CLI while ago, and it was awful tbh.

krat0sprakhar 11 hours ago

qudat 9 hours ago

i think it's better than sonnet 5, especially when you compare speeds. i have to work with the llm anyway, the faster i can turn it the better the outcome.

kelvinjps10 10 hours ago

I think about Google is the value you get of their plans, for 5$ a month you get their ai plus model combined with 400gb you can share this with your family. The other ai companies don't provide family plans

Sir_Twist 10 hours ago

And the free year-long trial for college students they recently offered, which includes 5 tb of Google Drive storage.

kelvinjps10 5 hours ago

I got 6months for free when I bought my s25.

Galorious 7 hours ago

Is anyone here using using these models via google subscription (not api). I tried to in the past using gemini cli and then agy - headless invoked by codex and claude code, but they were so incredibly buggy that it stalled 1/2 times and I cancelled. Interested to know if that has changed!

speak_plainly 12 hours ago

After struggling with Gemini for months, I think the trick to getting the most out of the model is writing a really solid personal intelligence/instructions prompt. The results are night and day in terms of performance.

titularcomment 11 hours ago

Funnily enough you really do need a great prompting and SKILLS setup to use antigravity effectively in contrast to other providers which actually started benefiting from less detailed prompts over time. But I like it this way, its more customizable and much cheaper especially with a sub.

porridgeraisin 11 hours ago

agy is good for those cases where you are willing to put the effort into the harness specifically for a task or family of tasks. The full suite, with evals, monitoring, hooks, custom tools, custom verifiers, etc,. It is not good if you want a "general coding assistant" like codex or claudecode.

The reality is that if you optimise a harness for a family of tasks[1], then most of these models give successful output. And there, gemini flash's speed shines.

For general coding assistant, you want it to be well, general, and you use a harness without too much customisation to something specific. Here you need deeply post trained coding assistants and implementors like codex/sol or claude/opus. Gemini flash in its current form will be too happy-go-lucky if you try using it the way we all use codex and is better used in a constrained setting.

tl;dr gemini flash for "LLM-aided workflows in production" is super good today. Cheap as well.

[1] Stuff like this: https://antigravity.google/blog/teamwork-when-ai-becomes-a-r...

https://hamel.dev/notes/llm/evals/

dakolli 11 hours ago

slot machine addict thinks if he pushes buttons in a certain order the odds get better.

In all seriousness, gemini has the best interactive planning document/orchestration. Tell it to create a plan document and work through it with it and it will preform really well(in antigravity products). But this is the case with plan modes with every model, I just think the interactive document that antigravity uses is really well thought out.

nharada 10 hours ago

Meanwhile I pay for Pro and still don't have access to 3.7?

almog 9 hours ago

Same for me (at least through the Gemini app).

ddp26 9 hours ago

There must be a deeper read on why Google can rapidly ship better small models while being delayed months on the bigger model.

What's the simplest explanation?

cogman10 9 hours ago

Perhaps post training? I believe I read that Qwen 3.8 is just post trained Qwen 3.6, which is why it was able to be released so quick.

It may be that these flash models are simply post trained larger older models.

jpau 7 hours ago

The iteration cycle is becoming very quick. Gemini 3.8 Flash arrived just 20 days after 3.7 Flash.

Similarly Qwen3.8-Max was updated in just 30 days (to the 0902 release) and Muse Spark in just 28 days (to the 1.3 release).

A year ago iterative releases were every 3-6 months. At what point will they reach nightly candidates?

leumon 12 hours ago

So 89.4% on Terminal Bench 2 but only 19.1% on Tbench 4. Opus 5 is 89.1%/51.8%.

aszen 7 hours ago

I was thinking the same, obvious suspicion is they benchmaxed it on older bench.

_aavaa_ 11 hours ago

Do they officially support you use their AI Pro subscription (or whatever the heck it's called this month, the one that gives you models in antigravity) in a 3rd party harness?

pwython 12 hours ago

Is there any reason to even use 3.1 Pro now?

fridder 11 hours ago

In my experience? No. 3.7 is faster and it just seems to get things right more often. Only big architecture tasks and analysis make sense with 3.1, perhaps, but honestly just use the Opus 4.6 to generate a plan and then switch back to flash for the implementation

bitexploder 12 hours ago

It is still going to be better at text work, skills, document review, deep reasoning, architecture review, etc. It is only 6 months old, it isn’t like its world knowledge and software knowledge is really out of date. Use it to churn on harder design problems.

exacube 9 hours ago

IME 3.1 Pro still has better system-instruction following than Flash 3.7, esp. when there're many conditions and clauses. 3.1 also writes better prose for technical material than Flash 3.7.

Once the system prompt complexity goes up, Flash starts to write very dense english. it might be fine for tasks like coding, but not for user-facing text meant to be digested by the average person.

I haven't tested 3.8 on my workload yet.

Rodmine 8 hours ago

3.7-flash has been useless many times, specially when context gets bigger. 3.1 is the only Google model that has seen use from me. With extended thinking, 3.7-flash is kinda usable but not without many problems. I find myself falling back to 3.1 often. I don't believe in any benchmarks because whatever they are doing to award 85% to 3.7 on anything, they should seriously reconsider that test for anything.

kelvinjps10 12 hours ago

I see benchmarks beating sol terra and sonnet. But is actually better? Has someone used it? I don't see actually much people that use Gemini for coding.

newppc 4 hours ago

If Google has the juice and wants to win, they need to start releasing world models.

satvikpendem 12 hours ago

Is the Gemini CLI still terrible compared to Claude Code and Codex? The harness the main thing holding back Google models as they could've been the best given all the advantages in compute capacity and training data they initially had, where now even the Google CEO said they're falling behind in agentic tasks, which is sort of a vicious cycle because RLHF relies on human usage.

dudeinhawaii 3 minutes ago

Honestly, it's platform dependent and "OK" at best, "Mediocre" at worst (Agy on Windows).

Gemini is great via the Chat interface and decent via Github Copilot.

I honestly hate it via Antigravity CLI because their sandboxing system frankly doesn't work. Every other harness has mastered "don't ask me if you're working in this one directory and using common commands". Agy instead either tries to pull a global elevation or wants every tedious variation of a command string whitelisted. Madness - circa 2023.

Agy _really_ needs to make the out-of-the-box experience cleaner and hassle-free. Heck, even Grok CLI "just works".

This may reflect a global mind-shift from "approve and validate everything" to "just do the stuff and only ask permission if it's outside the folder or a command that actually requires elevation". Maybe that's not for everyone, but for those that do want to perform unattended agentic work -- Agy is painful.

stwrt 12 hours ago

In May they replaced the Gemini CLI with the Antigravity CLI.

https://developers.googleblog.com/an-important-update-transi...

satvikpendem 9 hours ago

That's what I meant sorry. I used that one too and it still wasn't as good as competitors.

rancar2 12 hours ago

That was sunset and replaced by Antigravity. FWIW until I abandoned it knowing the sunsetting, I was able to get good behavior out of Gemini CLI with overriding the system prompt. The default prompt crippled the harness with very poor instructions, but there was a hidden ENV to override it. Replacing it with Claude Code like prompts based on the model selected, it ran at a much higher intelligence level full stack with significantly less errors.

pshirshov 12 hours ago

There is no Gemini CLI anymore, nor you can use Gemini with your own harness unless you pay per-token.

visarga 12 hours ago

it's called `agy` now

zipy124 12 hours ago

It was superseded by the antigravity CLI.

fridder 11 hours ago

it is antigravity now. It is ok

therealmarv 10 hours ago

On my short tests: This model is amazing and the speed makes it feel like another sort of AI.

But it's bad at code reviews (maybe it's the harness agy cli?). Could not get it to same quality level on reviews like Opus, GPT 5.6, Grok. Even tried special code review skills but no luck.

luciana1u 9 hours ago

Flash Cyber sounds like a villain from a 90s hacker movie and I'm here for it.

sreekanth850 10 hours ago

Dear Google, Kindly make you chat window on the right side of vscode in antigravity extension, There is a reason others kept it like that. I can see the code and inspect the files changed while Agents keep working. its critical for me personally.

algoth1 4 hours ago

I asked gemini 3.8 high to review the site I'm working on for points of high cpu/ram consumption - it failed spectacularly and also halucinated the server i/o limits

mrbonner 9 hours ago

I’m interested in a general knowledge model (closed or open weight) and not coding specific. I want to plan for travel and trip. Do you have one of your favorite HN crowd?

TechRemarker 11 hours ago

Hopefully before they release 4.0 Flash we will finally get Gemini 3.5 Pro.

simonsarris 11 hours ago

more likely 4 pro will be released pretty soon instead, since pre-training for 4 began in late July

https://x.com/OfficialLoganK/status/2079594867161022817

re-thc 11 hours ago

> 4.0 Flash we will finally get Gemini 3.5 Pro

Nah, we'll just get the 4.0 Pro Preview.

centaurz 7 hours ago

A company with 400+B revenue from software cannot build a usable command line cli for its vital AI model?

japgolly 3 hours ago

ldm0 8 hours ago

It’s strange that its score on Terminal‑Bench 4.0 is so low. They aren’t fast enough to benchmaxx that section.

lgl 11 hours ago

Am I the only only one thinking that Google might still "win" the AI race, despite the apparent gap?

They're apparently evolving slower than most SOTA models but "slow and steady wins the race" is probably still a thing.

And since Google doesn't depend exclusively on AI models, they can probably afford to "wait and see" where all this craze is heading.

sejje 9 hours ago

Staying a little ways behind the leaders is not "slow and steady." Every company is moving very fast right now.

I don't think slow and steady will win this race, but I think anyone can still win--especially Google.

evilhackerdude 10 hours ago

i always thought alphabet’s own youtube videos must be a comparatively good source of new training data. if slop and other garbage is reliably filtered out it should leave plenty of higher quality content.

ASinclair 12 hours ago

From personal experience it feels much more capable than 3.7 Flash.

johnnyApplePRNG 2 hours ago

Why is it still such a bad coding agent? Does anybody have any insight?

I am continually impressed with Gemini's chat responses, which encourages me to test their agentic capabilities and... no... no... and no... every single time.

It's terrifying watching it, really.

sva_ 12 hours ago

Barbing 12 hours ago

  [1] For tone and instruction following, a positive percentage increase represents an improvement in the tone of the model on sensitive topics and the model’s ability to follow instructions while remaining safe compared to Gemini 3 Flash. We mark improvements in green and regressions in red.
Gemini 3 Flash?! So is Gemini 3.8 Flash less safe than 3.7 Flash in all areas besides Text to Text Safety (and identical on Image to Text Safety)?

Why bother with a column “Gemini 3.8 Flash vs. Gemini 3.7 Flash” when you’re going to disregard the label for 20% of it? Also is the “Tone” label short for “Tone and Instruction Following”?

Chartcrime, the major AI lab tradition.

mattlondon 12 hours ago

im_soul 9 hours ago

disclaimer : Introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.

prometheus1992 12 hours ago

Google keeps flashing everyone where everyone is expecting to get PRO'bed.

kzrdude 11 hours ago

We also had GLM-5.3 flash and Qwen 3.8 Flash Next, everyone's getting flashed and I think it's a good trend.

Almost suspect that the rate of improvement to post-training is so fast that small models have an advantage - it takes much more compute to train a bigger model, so the flash models are just running in circles (well, not exactly of course) around the larger models right now.

atemerev 11 hours ago

Everyone is censoring models now with anything remotely resembling cyber or bio. I already have problems with my research in mathematical epidemiology because of that - both Sol and Fable simply refuse. They keep pushing people towards Chinese models that can be decensored.

vehemenz 11 hours ago

Supposedly Fable 5.1 is better, but I haven't tried it yet. I've run into the same thing with mundane work that is barely bio/cyber adjacent.

Re: Chinese models, even if the model itself isn't censored, some of the big model providers have guardrails now that you can't exceed, which somewhat defeats the purpose.

atemerev 8 hours ago

"Uncensored" means "weights modified to remove refusals". Abliterated. Providers do not serve such models, at least not frontier-grade. You have to run the weights yourself. For Kimi K3, this is about $60/hour for hardware rental. But you can have about 100 sessions simultaneously.

And yes, Fable 5.1 has the same refusal rate, and significantly nerfed reasoning.

schmorptron 5 hours ago

A reminder that google is the only major lab without a meaningful opt-out of training on your data. The only way to opt out is to disable message history entirely, which seems like a darkest of dark patterns to get users to leave "opt in" to training on, because next to nobody wants to use it without message history.

koalaman 10 hours ago

I use Gemini to make sense of things Claude says to me.

fitsumbelay 12 hours ago

shows up in /models though and encourages you to use it over 3.7 Flash I prefer this over reading specs: the "just show me" way

realist_not 12 hours ago

Anyone has a cached page / mirror ? 404

firemelt 9 hours ago

I wish google to thrive

levelZero 9 hours ago

Gemini 3.8 flash thinks Entoloma sinuatum is good to eat... Otherwise feels great

dismalaf 11 hours ago

Nice surprise. In a few of my own tests it seems maybe a tad slower than 3.7 (but still way faster than any other LLM I've used) and even smarter. With 3.7 I felt I could just not use 3.1 Pro at all and 3.8 seems even better.

amazingamazing 11 hours ago

Could someone explain to me why it matters if google has the best model? Isnt the real metric cost per task?

advenn 12 hours ago

But where is Gemini 3.5 pro?

simonsarris 12 hours ago

it is most likely that 4 pro will be released pretty soon instead, since pre-training for 4 began in late July.

https://x.com/OfficialLoganK/status/2079594867161022817

GaggiX 12 hours ago

Gemini 3.5 pro is never going to be released, it was a failure.

mohamedkoubaa 7 hours ago

The race to the bottom continues

HardCodedBias 9 hours ago

I have to say:

The Google brand remains powerful on HN!

I’m shocked.

OG_BME 12 hours ago

What did it say?

alex1138 10 hours ago

It's a shame Google crams it ham-fistedly into search results and that Google has some of the reputation it has because I actually really enjoy Gemini and I don't even use it for the reason people often list which is that you can cross-reference it to stuff in your Google account

dcchambers 11 hours ago

I would really love to be able to use these Gemini models in Opencode or Pi with my existing Google AI Pro subscription.

eis 11 hours ago

3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just increased the thinking budgets...

3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash

Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...

WASDx 10 hours ago

3.7 high and 3.8 medium are essentially the same on AA intelligence and cost. Output tokens on DeepSWE gives the same picture. So there might be something to it but they have done other things as well. At least the tokens are really fast.

zuzululu 7 hours ago

i find deepswe not very reliable for instance it puts grok 4.6 xhigh over sol medium

eis 11 hours ago

3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just uppsed the thinking budgets...

3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash

Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...

dyauspitr 11 hours ago

Whatever they’re using within the Maps app is not good at all. I cannot just ask it for things conversationally like I do with ChatGPT. They really need to put a better model in there. I don’t even think it maintains context across two different queries within the same session. It’s not seamless and doesn’t just “get it” like ChatGPT does.

Yesterday I asked for food stop on my road trip 45 minutes from the current time and it gave me some options, but then I changed my mind and specifically asked for Asian restaurants and it completely forgot about the 45 minutes and gave me the closest Asian restaurant to me.

FpUser 11 hours ago

>"safety performance" - this starting to get long in the tooth. Gemini cut programming session 3 times for "safety reasons" yesterday for mentioning image generation (I need to generate bunch of those for infinite zoom virtual training app experience). After I got creative and managed to trick it to answer t was of course because "think of a children"

And in my other app I was debugging and using OpenAI to optimize some path it cut me off numerous times because it did not like JIT functionality (this is my commercial business rule evaluation engine that compiles rules to executable code inside the app to increase performance using asmjit library)

I am basically paying for them to waste my tokens and time on these 2 tasks

sergiotapia 12 hours ago

Who coined the phrase "cyber" for security related things lol. It's so 1999.

estearum 11 hours ago

Hasn't the field been called "cybersecurity" since... forever?

fwip an hour ago

Sure, but the appropriate shortening here is "security."

Calling it cyber is like shortening email to "e".

sumeno 10 hours ago

I assure you that in 1999 "cyber" meant something very different

josefresco 11 hours ago

Cyber is more of an early 1990's thing, and I have no issue with it unlike most in the tech field. I feel like it dropped off in the late 90's and early 00's but made a comeback as hacking became a mainstream security issue.

cleverpotato479 11 hours ago

Amusingly, "cyber" comes from the word "kubernetes"!

Gander5739 11 hours ago

Relevant xkcd: https://xkcd.com/1573/

atemerev 11 hours ago

A/S/L?

barapa 12 hours ago

love these flash models

jdw64 12 hours ago

The biggest problem with Gemini is that its performance degrades the longer you use it for coding. Is it just me?

hirako2000 11 hours ago

Filling the large context does that yes.

But a good agents.md, starting from a clean slate, and specifying which key files to look into and follow the standards allows me to build gigantic projects even I struggle to keep in my head structurally.

deno 11 hours ago

Seems maybe you’re keeping a forever-session and multiple independent tasks end up overstaying in context?

I would say either start new sessions for new tasks or limit the context to something smaller than 1M.

I usually start with research/planning session, this goes into a detailed implementation plan and then a new session for the actual implementation.

If it's complex problem maybe a review/adversarial step between plan and implementation.

Also with forever-session any time you take a longer break (depends on model and provider as to how long) you will push an entire big context again without caching even if you don't need it. With 1M context this gets expensive.

zuzululu 6 hours ago

thats not just the context growing issue, hallucinations is a thing

yipinwong 12 hours ago

"Page not found"...

Razengan 6 hours ago

What is with Google's dumb ass STILL refusing to respect the OS dark mode setting in fucking 2027??

Mashimo 12 hours ago

It's 404 now.

freedomben 12 hours ago

Came and went in a flash

k8sToGo 12 hours ago

Because they are preparing Gemini 3.9 Flash

pixl97 12 hours ago

kingstnap 12 hours ago

The blog post is gone but I can currently use it in the gemini chat website.

zuzululu 7 hours ago

not really getting the excitement over this, its at opus 5 medium level, and opus 5 is not really the go to model , claude purists hate it

so its fast sure and decent at non coding usage but for developers nothing can really top sol or fable.

even grok 4.6 is so so and i would not choose 3.8 flash over it.

deanc 12 hours ago

And yet again another failed launch from Google. I pay for their AI plus Google one package to get more cloud storage (have no interest in their AI bundle but you have to pay). and all I see in the Gemini app is 3.6-flash

WarmWash 12 hours ago

Google has been doing staged roll outs on all their products since forever.

deanc 11 hours ago

What stage of the roll out are we where I don’t even see 3.7-flash which was released 2-3 weeks ago?

phsau 11 hours ago

deno 11 hours ago

mythz 12 hours ago

I'm trying it now for token heavy coding tasks, it's capable for many tasks but in noway compares to Claude/Sol - requires more prompts and the output isn't as good.

So just another mid-tier flash model, nothing exciting, but Antigravity has very generous quotas so it's a good workhorse model when your Claude/OpenAI subs run out.

And whilst it's a fast model, having to baby sit through and approve prompts every few seconds ends up making it slower than the Auto approve modes of Claude/ChatGPT - they definitely need an auto approve mode.

titularcomment 11 hours ago

`agy --dangerously-skip-permissions`

mythz 11 hours ago

anyway to do this with the Antigravity macOS App?

zuzululu 6 hours ago

not sure why you are being downvoted, but that has been my experience with 3.7 flash and sol/fable comparisons

i think luna-max has the best cost value offer when it comes to coding, but i note the multi modality of gemini flash as a win

i might consider 3.8 flash for simple side hobby projects or quick scaffolding but would not trust it for long agentic tasks, that really is the realm of sol/fable

agy cli still has a lot of issues not sure if its due to the underlying model hallucinating or the harness or both

tacomonstrous 12 hours ago

Looks like Google's given up on frontier models for external consumption?

heyjamesknight 12 hours ago

Gemini 4 pre training is underway: https://x.com/OfficialLoganK/status/2079594867161022817

My guess is we skip 3.5 and go straight to 4 Pro. With the monthly Flash releases, releasing 4.0 Flash and Pro in 6-8 weeks would be a nice buildup.

(I work at Google but don't know anything that isn't already public)

WarmWash 12 hours ago

Latest rumor is that 3.5 pro was struggling to be meaningfully better than flash, since iterations on flash were moving much faster than iterations on pro, likely due to model size (flash is estimated to be in the 200-400B range).

VirusNewbie 12 hours ago

I found 3.5 pro to be much better than 3.5 flash, but 3.7 flash with high reasoning is comparable and way way faster.

j16sdiz 12 hours ago

iamdelirium 12 hours ago

How can you say that when a Flash model is benchmarking close to Opus and Sol?

ok123456 12 hours ago

Given up frontier models for selling compute.

thisisauserid 12 hours ago

They don't want to release a frontier model that requires data sharing with the government and right now it looks like they'd have to.

shuvrojit 12 hours ago

Gemini is getting less useful with each update. I could edit a pdf with the 3-pro model before but 3.1-pro couldn't edit the given pdf nor it could generate one for me.

HarHarVeryFunny 8 hours ago

If you want to "edit" a PDF, then Claude Sonnet works well, although what it's going to do is regenerate it from scratch trying to retain overall formatting. It can even do this for scanned PDFs and foreign language ones that need translating.

If you just need to create PDFs, not edit them, then Gemini notebook (notebook.google) works well and has Google's usual very high free usage limits.

AFAIK in general you can't really edit PDFs since it's not a reflowable format - even with Adobe tools all that editing does is modify the text within a text box - not reflow the document to adjust to any change in size of the text box.

leumon 12 hours ago

You probably mean 3.5-flash? Pro is still good for a lot of use cases, but it seems it's still officially in the "preview" phase.

ipsod 12 hours ago

3.5 pro doesn't exist yet?

shuvrojit 12 hours ago

Sorry my bad, I messed up the numbers, 3 and 3.1 pro. All of these model numbers have me confused

coffeecoders 12 hours ago

One place where I find the Flash models surprisingly bad is Google Search's "AI Mode".

A recent example - I searched for how to unsubscribe from Pearson emails. Google Search "AI Mode" confidently gave me a sequence of steps along the lines of Settings > Profile > Email preferences > Unsubscribe.

Of course, I looked for an unsubscribe link before asking Google. None of those options existed. The correct answer was there is no way to unsubscribe through the account, so I just blockthe emails instead.

I've run into this pattern quite a few times. AI Mode seems to make up things all the time.

inventor7777 12 hours ago

I think that's just a limitation on the size of the model. I'm pretty sure that they use a pretty small model in those summaries to save money, which naturally makes them a little less smart.

pixl97 12 hours ago

https://www.pearson.com/privacy-center/privacy-notices/full-...

>We will not send marketing emails to a user who has opted out of receiving them. Any marketing communications we send will include an unsubscribe link at the end of the email.

I don't think this is AI's fault. This is Pearson's publishing incorrect information and the only way to really know they are a bunch of lying assholes is to have an account and try to unsubscribe from it.

AI didn't make it up, Pearson's did.

xyzzy_plugh 12 hours ago

It's not the models, it's the guardrails.

It's obvious that the Google Search AI Mode encourages the model to give an answer without spending unnecessary cycles investigating deeply.

They also heavily encourage keeping the context short. For example, it will remove the option to start a new turn after a small number of turns, depending on the topic.

It definitely makes things up all the time, but it gets it right surprisingly often. I really like it.

Alpha3031 5 hours ago

The search model is probably flash-lite based on what they give to users who aren't signed in.

greenowl 11 hours ago

Not to rain on anyone's parade but I find it strange how excited and giddy people on HN get for any new X.X model releases. Pumping it straight to the top, clamoring to use it, check and compare benchmarks, bragging about it being your "daily driver"?

Are you people truly this excited about this crap? I mean I guess if you work for Google or Anthropic or whatever I could see it??? Otherwise, are these just bot comments?

nick__m an hour ago

If you used, you would know. There's something addicting seeing the vertigo inducing progression of that technology.

I am a light user so I don't get the shakes when my monthly azure dev credits run out but I would be susceptible to being addicted to it if I was on a subscription with generous usage allowance and random usage counter resets.

ipsod 11 hours ago

Gemini Flash is the one I get most excited about, because it's so fast and so good at real-world knowledge, and it's improving so fast - look at how much the benchmarks improved in ~1 month. It's just categorically different than anything else.

Also, I use it every day, and it just got ~10% better at coding, according to the benchmarks. How is that not exciting?

rjh29 8 hours ago

I use Gemini every day and I've noticed any subjective improvement. In many cases it feels worse because it does fewer Google searches than before. As a result I find it hard to get excited about it.

I do think Gemini is underrated on HN though!

drbscl 11 hours ago

Given that they push capabilities at the pareto frontier, yeah

A lot of us use these in our services, so we're getting an upgrade "for free"

deno 11 hours ago

You know how the saying goes that you have to pick two out of three: cheap, fast or good? This is all of those. Pretty exciting.

I'll wait for Astra and Grok 4.7 announcements but probably getting at least one Ultra subscription.

Since testing 3.7 on Pro for last two weeks I'm realizing just how long I'm waiting on other models. I've been multitasking to compensate but it's exhausting so I'd rather not.

anslopic4 5 hours ago

Yes they are mostly shill and bot comments. Some of the big accounts are paid influencers, some of the other comments are purely AI.

HN sells these advertising services. Nobody is using “Claude” etc.

They will censor comments like yours and my reply here because we call it out.

It’s very weird that basically lies and disinformation became the optimal meta in business and in life! But here we are

rvz 2 hours ago

Correct. This orange site has evidently gone under AI psychosis especially in model release posts and is overrun by AI bots, paid influencers and even small creeping signs of crypto pumpfun scams [0].

Even making a tiny joke is too much [1] for some.

> They will censor comments like yours and my reply here because we call it out.

Don't bother calling it out, it does not work. There are protected accounts where the guidelines don't apply to them and moderators allow this and ban others who do the same thing. [2]

It is pointless, and HN is cooked for this.

[0] https://news.ycombinator.com/item?id=49521145

[1] https://news.ycombinator.com/item?id=48838228

[2] https://news.ycombinator.com/item?id=49366029