Cerebras CS-4 (cerebras.ai)

299 points by sunils34 11 hours ago

KronisLV 3 hours ago

God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing

Guess they don't care about regular devs atm and are focused only on hardware sales.

scosman 2 hours ago

I don’t think they will until they change the architecture.

They don’t have a prefix cache like other providers, or at least don’t have a discount in their billing structure. Each message charges for the whole context window. It’s wildly more expensive for long multi turn scenarios with lots of tool calls (coding). It’s better for short few turn tasks.

Edit: I don’t know if they actually have a proper cache. This could just be a billing artifact.

preommr 2 hours ago

> GPT-OSS 120B which is nigh useless nowadays:

I still think that was a really great model that got overlooked. It was really great in terms of latency/throughput while still being fairly intelligent.

I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.

KronisLV 2 hours ago

> I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.

Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs and you have to fix a non-insignificant amount of it all manually: https://blog.kronis.dev/blog/i-blew-through-24-million-token...

Admittedly that post was before agentic development truly took off and that 3k EUR figure when paying per API tokens would nowadays be closer to like 6k EUR for the volume of work I do, but still.

It's the same how Qwen 2.5 was pretty problematic for anything remotely serious, same with Qwen 3 Coder Next (80B), and at least the most recent versions are getting better but still not quite good enough in real world use cases outside of benchmarks. They've come a long way, regardless!

wongarsu 2 hours ago

As MoE with 5B active parameters it's pretty fast. But you still need a lot of vRAM, or have to run small quantitations. Qwen models just gave you more bang for your buck, and the gap became worse with every qwen release

LoganDark an hour ago

gpt-oss-120b is absolutely unusable over Cerebras. It fails to call tools half the time and just continues to think about what tool it'll call repeatedly. Like it says it'll call a tool and then it doesn't, and then it says it'll call the tool again and then it doesn't, and it just does that in a loop forever. It's awful. Also forgets to end the thinking block too. Even if the model itself was just-okay for its time, even at 1000t/s+ it's not worth it

xmorse 19 minutes ago

Cerebras is very fast but you can basically never use it because of its scarcity

syntaxing 9 hours ago

I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.

yorwba 3 hours ago

You cannot infer this because they only show the tokens per second per user. One way to get a higher number is to have fewer users per chip.

I'm pretty sure Cerebras has a confidentiality agreement with OpenAI, and this press release was carefully constructed to avoid leaking details about the model weights. For example, the graph of tokens per second vs. tokens per second per user doesn't have any numbers that would allow you to translate between the two. (And in any case the relationship depends on the model.)

Copenjin 2 hours ago

Didn't they say that they can support bigger models now?

logicallee 9 hours ago

(Where did you see that?)

This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use?

I think the frontier providers keep the size of their models carefully hidden.

redox99 6 hours ago

You don't really need to train a 10T model to test cerebras against a 10T model. You can feed it an untrained (randomly initialized) model and benchmark it. Result will be gibberish but performance the same.

nl 6 hours ago

Mythos/Fable are around 10T:

> According to FT, industry estimates say Anthropic's most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion

https://www.reuters.com/technology/bytedance-targets-mega-ai...

I believe this report has confused Opus (which is known to be around 5T) and Fable.

Other reports say 10T. See for example https://eu.36kr.com/en/p/3760679047267075?ref=explainx where Musk talks about the models being trained on Colossus2

YmiYugy 5 hours ago

zozbot234 5 hours ago

alightsoul 6 hours ago

pretty sure 10 trillion parameters is now the norm among closed ai labs, given that nvidia also references the same 10 trillion number for their nvl72 racks

KronisLV 2 hours ago

ewild 9 hours ago

It's rumored fable is around that 10T number

walrus01 9 hours ago

johnnyApplePRNG 7 hours ago

sreekanth850 9 hours ago

AMD along with cerebras may probably compete with NVIDIA monopoly in near future. Also, NVIDIA will have competition form multiple companies. Just my prediction.

anonzzzies 5 hours ago

Like NVIDIA bought Groq, AMD might do well buying Cerebras.

FranGro78 an hour ago

AMD did enter into an agreement to buy Taalas, which is speculated [1] will be used to augment their Helios offering.

1. https://www.youtube.com/watch?v=3MKRjt59hh4&pp=0gcJCRMMAYcqI...

sreekanth850 5 hours ago

They are Ex AMD employees. Nvidia tried back, but they rejected.

vezycash 3 hours ago

eitally 8 hours ago

Maybe, but GPU is just one aspect of NVIDIA's dominance. If you are buying Vera Rubin GPUs, you're getting an NVL72 rack, which is only one of several racks that you're probably buying. You'll also need your NVIDIA racks with NVIDIA networking & storage gear, too. At the end of the day, they're "vertically integrated" for your accelerated computing data center (e.g. the "AI Factory"). This doesn't even count the software layer, where CUDA + CUDA-X (not to mention the software for all the sysadmin pieces) has a huge first mover advantage over anyone else.

zarzavat 7 hours ago

Is CUDA still a moat? Are we not at the point where frontier models can reimplement software stacks, given you throw enough tokens at the problem.

fooker 4 hours ago

someothherguyy an hour ago

aurareturn 4 hours ago

alightsoul 6 hours ago

vatsachak 7 hours ago

_zoltan_ 4 hours ago

Press doubt. Single GPU? Maybe. MultiGPU behemoths like NVL144 and NVL576? I don't think so.

NVLink is at gen9. they had a lot of teething problems and can codesign the hardware and software.

in the name of openness (AMD's only """weapon"""), the UALink spec is a hodgepodge of corporate opinions with very different implementations (looking at you, Broadcom). at spec version 1 (in hardware).

I wish them good luck as I really like AMD, but they compete no more on this than Lambo vs Bugatti.

epolanski 3 hours ago

It's not really a far fetched prediction: high margins and huge market attract competition, that's just the law of economics.

rajnathani an hour ago

Interestingly they’re still on the WSE-3 (5nm TSMC) wafer chip and slightly bumped up the specs there (overlocking mostly it seems), for why it’s called WSE-3 Turbo now. I think people were also expecting WSE-4, as it’s been 2 years now since WSE-3 was launched.

reilly3000 9 hours ago

> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters

Oops did they just out GPT-5.6 sol’s parameter count?

sho 9 hours ago

Sol is supposed to be 5T according to rumour. The imminent Astra is allegedly 10

nozzlegear 8 hours ago

Rumors and allegations aren't worth much. Why don't they just tell us mere mortals?

WinstonSmith84 3 hours ago

brookst 7 hours ago

kanwisher 4 hours ago

cerebras model are different size then the original models

whatever1 9 hours ago

I mean we kinda know the frontier models are multi trillion parameter models. The only open weights that are close to the frontier are that size too

verdverm 6 hours ago

save qwen3.8 27B which is outclassing much larger models and is in spitting distance of the top 10 in https://artificialanalysis.ai/models#intelligence

Vax- 6 hours ago

ethanzhang1024 9 hours ago

If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?

walrus01 9 hours ago

Without having any inside information, one possible theory:

All or a vast majority of of the cerebras manufacturing capacity was going to a few companies that aren't publicly available inference providers on openrouter, for their own internal use.

or

The asking price of the S-3, no matter how speedy it might be, for small/medium size customers made it economically prohibitive to purchase and use to sell public inference vs. buying more common nvidia b200 or whatever.

wmf 7 hours ago

Cerebras provides high-speed inference at high cost. It's never going to be the cheapest and thus it will probably remain niche.

zurfer 3 hours ago

But that's supply and demand, not technology. Right now a lot more people want their inference than they can supply. as supply catches up in the next 5-10 years, the underlying tech at scale is probably cheaper than GPUs per token produced.

aurareturn 4 hours ago

Probably the same reason why there are more people who takes buses, subways, trains than drive Ferraris.

epolanski an hour ago

I love how instead of comparing a Ferrari (fast and expensive) to some average car (not fast, not expensive) to make your point..you went for public transport where your comparison cracks from multiple angles.

smallerize 9 hours ago

If you're willing to pay a significant premium for latency, why use openrouter? And anyway Cerebras only supported a few specific models.

gampleman 2 hours ago

Cerebras capacity was pretty much entirely bought out at some point. We needed it and couldn't get it.

aseipp 8 hours ago

The WSE is very expensive to build, and they have a waiting list of customers who are already willing to pay a lot of money for the available supply.

doctorpangloss 9 hours ago

it only takes ~445 GB300 NVL72 (about $22b) to run ALL of openrouter demand for a year. Microsoft rolled out $32b of DC 2026Q1.

imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)

aurareturn 4 hours ago

I thought your numbers must be wrong.

So I plugged 288 trillion tokens/month (OpenRouter's current rate), 500 billion MoE model average, and the math comes out to be around 620 B200 GPUs minimum.

So basically, OpenRouter's volume must be absolutely tiny compared to the volume hyperscalers are getting.

HDBaseT 8 hours ago

It is worth mentioning, the OpenRouter demand isn't static though. It has increased week on week since early 2026.

senordevnyc 8 hours ago

I was curious so I looked it up: looks like a GB300 NVL72 is about $4M. So $22B would buy you 5500 such racks, no?

selectodude 4 hours ago

petesergeant an hour ago

I use Cerebras via OpenRouter. It’s every bit as fast and reliable for my needs as claimed. I suspect the reason is that they can either be making peanuts selling inference to plebs like me via OpenRouter, or making bank selling the more expensive models to businesses directly. In short: I would be very surprised if they have die capacity, and are at this point maximising revenue per chip.

aneryu 9 hours ago

It would be even better if a version available to individual users were released soon.

gpm 8 hours ago

They do offer API services to individual users... though with a set of models that makes it unlikely that you want to use it. They are promising Qwen 3.8 27B any day now though*.

if you have the money as an "individual user" to purchase one of their racks... save your money and retire.

* Actually they sent out an email claiming they already have it, but I don't seem to have access, they're promising to release it to the "shared tier" any day now.

drcode 7 hours ago

not only that, but I was so happy with their GLM 4.8 that they got rid of yesterday :(

apitman 7 hours ago

fragmede 8 hours ago

> save your money and retire.

Now that this hypothetical person has retired, what are they gonna do all day? Just sit on the beach and drink Mai Tais? If that's what they wanna do, sure, but nerds gonna nerd, and if I had that kind of money to retire on, I'd totally buy some ridiculously expensive AI box for fun.

gpm 8 hours ago

dnautics 7 hours ago

needs an sla that says power will never ever ever go out or else you will have a useless shattered plate of silicon.

gpm 6 hours ago

Huh, why would it shatter if the power goes out?

ttul 8 hours ago

I’ll get that 250kW home power service dropped in next week!

z2 8 hours ago

For now you can rig an adapter to your nearest DC EV charging station, but make sure it's near a body of water for the cooling.

johntash 5 hours ago

I'd like to see a consumer version too, I don't need a whole rack of them. I probably can't even afford one gpu-sized one

anonymous_user9 10 hours ago

Conspicuously missing: power consumption figures

wmf 9 hours ago

162 kW

walrus01 9 hours ago

I guess we know why there's a fair bit of investment money going into small modular nuclear reactor startups now.

fragmede 8 hours ago

dakolli 3 hours ago

roughly 9 hours ago

God, I was going to ask if this could be deployed in a standard existing datacenter, but I guess that answers that question.

KeplerBoy 6 hours ago

xattt 9 hours ago

I presume per rack?

Can you imagine something radiating that much energy into a space in your home?

walrus01 9 hours ago

wmf 9 hours ago

undefined 9 hours ago

logicallee 9 hours ago

"10x more throughput per watt than CS-3"

9cb14c1ec0 10 hours ago

Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 per month.

> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters

Wow!

SwellJoe 10 hours ago

And, the software side isn't finished being optimized, either. We've seen with Qwen 3.8 27B and DeepSeek V4 Flash 0731 and GLM 5.3 that quite small models can pack a punch. Intelligence density will improve, efficiency of kernels will improve, efficiency of KV caching and MTP will improve, algorithms for splitting workloads across compute units will improve.

It'll all be as cheap as DeepSeek was before the price hike. And, it'll become more and more realistic to run near-frontier intelligence on personal devices.

dgellow 3 hours ago

> Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 per month.

We can have that discussion now: sounds like that would kill OpenAI and Anthropic

moralestapia 8 hours ago

Hence why taalas was one of the best strategic acquisitions of the year.

I'm honestly baffled they were not acquired by somebody else (sorry AMD).

adventured 7 hours ago

Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down.

The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years.

Cerebras will win in terms of approach.

It's 1998: hey, I can drastically speed up your web service, let's etch it right to silicon.

Iolaum 5 hours ago

walrus01 an hour ago

NitpickLawyer 6 hours ago

WithinReason 4 hours ago

pimeys 6 hours ago

alightsoul 6 hours ago

conception 5 hours ago

moralestapia 7 hours ago

api 10 hours ago

This is part of why I think the data center build-out is a bubble. We've barely scratched the surface when it comes to hardware optimization. We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear.

GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new designs. Basically every chip engineer on the planet is working on this right now.

mindwok 9 hours ago

Whether it's a bubble or not depends on how much the demand for compute and the type of workload keeps growing, though.

If AI tends to be something used mainly in ideation and development, which is how a lot of people use it today, then once consumer hardware gets good enough you could see a bunch of the current data centre workloads move onto consumer devices.

But if AI starts being used more in repeatable, operational workloads I think it makes sense to have significant cloud infrastructure for it. TBH I haven't seen much of this, and I've been skeptical about people using agents for much of anything when it can be done with just software. But we are starting to see more of this kind of workload, like the taggable Claude in your slack etc that people seem to really love.

aurareturn 3 hours ago

By the way, this is the same argument that Michael Burry used to short Nvidia.

He claims that GPU depreciation/obsoletion is much faster than hyperscalers are assuming because new chips will be much better. He's being proved wrong right now because H200 rental prices have been claiming for the last 8 month despite B200 having 10-20x better inference efficiency.[0]

The logic is fundamentally flawed in my opinion. Let's use future Nvidia chips being much better optimized for LLMs for example.

New Nvidia chips 10x better than H200 --> data centers buy a lot --> Nvidia profits a lot.

New Nvidia chips 10x better than H200 --> data centers don't buy --> no faster than expected obsoletion.

In other words, the very act of buying many new Nvidia GPUs would be the event that causes faster than expected obsoletion. Yet, if you don't buy those new Nvidia GPUs, then there is no faster than expected obsoletion.

We also live in a world where there is competition. If Amazon doesn't buy but Microsoft does, suddenly Microsoft can offer better $/token prices.

[0]https://inferencex.semianalysis.com/inference

haldujai 2 hours ago

petra 8 hours ago

I wonder: in world where inference is cheap, how many engineering agents that use simulation as their feedback we will use?

In the scenario, engineering everything becomes so easy - so why not optimize everything? every component, every product, every system?

And maybe llm's could invent. So even more to simulate. And simulation is inherently compute-heavy.

So unless there are some other bottlenecks, we'll use a lot of simulation servers.

RachelF 8 hours ago

True, I have to agree with you. The AI giants might be investing a huge amount of money in generation 1 technology. There might be a much better way to do it just around the corner. They might know this and thus the hurry to IPO.

A rough analogy would be if the first generation of ISP's spent billions on dial-up exchanges, when fibre could be invented next year.

winrid 10 hours ago

On the plus side, lots of cheap servers to swoop up :)

sroussey 9 hours ago

dgellow 3 hours ago

jeffybefffy519 4 hours ago

Exactly right, and nVidia is protecting their moat through business practices rather than genuine product innovation.

undefined 7 hours ago

[deleted]

__turbobrew__ 8 hours ago

By the time these gigawatt datacenters are done being built the hardware will be so far behind state of the art they may be mostly useless.

rvz 9 hours ago

Congratulations! You have just realized that the AI data center build out is a total scam, built on both the insurmountable trillions of debt, and the assumption that only GPUs are all we need to continue scaling.

There exist other AI accelerators (TPUs, ASICs) that perfectly exceed the throughput that LLMs need to scale as well. But the true solution is more software optimizations. There's a tiny handful of them but more needs to be discovered so that we can reduce building hundreds of more data centers as the alternatives mature.

As better software becomes more useful for the alternative AI hardware for developers with LLMs running efficiently you then would have more choices of hardware to run your LLMs on rather than just only GPUs.

blovescoffee 9 hours ago

TPUs and ASICs run in data centers too. Your argument only holds true if there's some satisfied limit to demand for inference. If not, data centers will continue to spring up to host more and more agents. Even if agents were running on hardware and software as efficient as the human brain, its conceivable we want trillions of them running at any given time which would require data center scale.

georgeecollins 9 hours ago

skyberrys 9 hours ago

I wonder what this looks like in 5 years... Will there be a massive push to repurpose these giant boxes into housing? Will they get turned back into the farm land from where they came? When a data center goes bust, what happens to the parts left behind?

gpm 9 hours ago

jryle70 9 hours ago

Huh, why I'm not surprised that HN is full of opinions confidently stated without any numbers or resources to back up?

> built on both the insurmountable trillions of debt, and the assumption that only GPUs are all we need to continue scaling.

Insurmountable according to whom? And who assume that only GPUs are all we need to continue scaling? Google, Amazon, Microsoft, Meta and OpenAI, all have or plan custom non-GPU AI chips. Do they plan to use them not for scaling?

kobe_bryant 7 hours ago

can these vibe coded sites please set a max width and overflow so their sites work fine on mobile

dgellow 3 hours ago

Good news, future models will have your comment in their training set, making them slightly more likely to fix that problem!

denizay 7 hours ago

The comparison seems incomplete. CS‑4 is a full rack-scale system with three wafer-scale processors, but the exact GPU models, GPU count, power consumption, price information are not disclosed. We still don't know if buying a multi-GPU rack (or racks) is cheaper and/or more efficient in power. The fact that they didn't disclose these numbers makes me believe that the numbers are not in their favor. And personally, makes me see them as disingenuous.

lostmsu 9 hours ago

KV caching status?

What's the point of 1000tok/s if you have to do prefill on every agentic turn which at 100k depth would make it 1.5 min latency every turn?

walrus01 9 hours ago

Information about RAM type/size and connection topology of the RAM to be used for context cache seems to be conspicuously absent from the slick looking marketing materials.

gpm 9 hours ago

There's a few more details at the bottom of this page: https://www.cerebras.ai/blog/introducing-cerebras-cs-4

44GB on-chip-sram * 3 chips. Per chip: 43.2 PB/s memory access + 53.5 PB/s on-chip fabric bandwidth + 2.4 Tbits/s "IO" bandwidth (I think that means their RoCE v2 RDMA over Ethernet interface).

I suspect there might be a certain amount of customization for how much RAM they attach when you order it.

porridgeraisin 7 hours ago

selimonder 3 hours ago

That "GPU" comparison is the vaguest i seen so far

epolanski 3 hours ago

True, it's also "per user", somehow, but I think it's a misleading metric. Cerebras chips take the whole wafer?

A single TSMC wafer contains 60 to 65 B200s, assuming 70% yields that's 40ish wafers per die.

Cerebras cannot redefine wafer economics.

tjoff 3 hours ago

Depends on what you mean, they have more redundancy which means that the yield can be much higher.

arthurcolle 5 hours ago

What's the sticker price? If I have 20 million in the bank can I just like buy one or what

runeks 3 hours ago

Surely it would depend on your model, since they need to etch this into silicon.

yvdriess 2 hours ago

You're thinking of Taalas.

sva_ 8 hours ago

> enabling massive clusters and models with more than 50 trillion parameters

4k0hz 9 hours ago

> Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers upto 30x faster inference compared to GPUs, enhanced economics, and a simple path todeploy [sic] hyperscale capacity.

Did nobody proofread this?

algoth1 9 hours ago

If they had ask Claude it would probably look like this: Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers up to 30x faster inference compared to GPUs, enhanced economics, and a simple path to load-bearing hyper scale capacity.

jm4 9 hours ago

That's unusually honest and the sharpest thing in this thread.

wren6991 9 hours ago

SoMomentary 9 hours ago

Sometimes I wonder if mistakes are now used to indicate the possibility that a human actually wrote it.

ceejayoz 9 hours ago

There’s been a spate of Reddit AI bots using all lower case in hopes of evading detection.

It’s still incredibly obvious.

VladVladikoff 8 hours ago

helloplanets 5 hours ago

Well at lest it's written by a human.

dgellow 3 hours ago

« Make it look like human written »

geodel 9 hours ago

Maybe it is just part of their "compact design".

dpkirchner 9 hours ago

An error no frontier LLM would make, eh

avantnyc 7 hours ago

Cerebras should slowly also move to dgx/ryzen market for a desktop version for masses at affordable price yet providing substantial tokens/second on desktop

wmf 7 hours ago

Desktop SRAM isn't really viable because it could cost $100K just to load the model.

aenis 5 hours ago

...so, in the same ballapark as ddr5? :-)

tamimio 9 hours ago

I wonder what are the benchmarks of hashcat on different hashes.

OutOfHere 10 hours ago

Five years from now, I don't know why anyone will still be using Nvidia for inference. Note that Cerebras is for inference only, not for training.

I understand that Cerebras has competition, but this bodes even more poorly for Nvidia for inference. Nvidia may still have a role to play for training, however.

dgellow 3 hours ago

NVIDIA has the best supply chain in the entire game. They are the only ones who can produce at their scale. You really shouldn’t underestimate their position

kcb 7 hours ago

Nvidia is at this time a pretty well run company tech wise. They are going to keep iterating on the inferencing hardware stack over the next five years too.

OutOfHere 38 minutes ago

The only way I see in which Nvidia can catch up is by buying Cerebras.

wmf 9 hours ago

Cerebras is only claiming ~2x the performance of Groqvidia which usually isn't enough for people to switch.

adventured 6 hours ago

OpenAI needs to immediately move to acquire Cerebras.

Nvidia's extreme margin is the opportunity for OpenAI's cost reduction. Buying Cerebras would pay for itself and they should take all of its future production (after filling required contracts).

Right now China's models have no silicon moat. Cerebras as a drastic speed-up / cost-reduction potential, can assist in building a competitive moat. And every time a Cerebras pops up, OpenAI or Anthropic should eat them if at all possible.

There's no stand-alone frontier AI company of great scale in the near future that doesn't have a large silicon advantage in-house. Apple knew it in smartphones, Google figured it out a long time ago as well.

undefined 3 hours ago

[deleted]

thefounder 6 hours ago

With what? More debt? What will nvidia say?

wmf 6 hours ago

Do you know about Jalapeno?

mmmeff 6 hours ago

^

OpenAI is partnering with Cerebras while simultaneously investing in their own silicon play. Hedged bets.

After sitting thru their keynote today, it makes sense. The main throughput speedups they tout are an obvious evolution of the GPU that all companies will be building in the next year. Wafer-scale interconnected memory and compute is just going to beat out mountains of network cabling any day on both cost and performance metrics.

gpm 9 hours ago

Is it just me or is it bizarre that they're advertising old open-weight models.

GLM 4.7 (December 2025) not 5 (Feb) 5.1 (April) or 5.2 (June). 5.3 (4 days ago) is, to be fair, not open weights yet... but there's a lot since 4.7.

Kimi K2.7 (April) not K2.7-code (June) or K3 (July).

Gemma 4 (April), Llama (April), and gpt-oss (August 2025) are up to date, but old (for models).

Meanwhile the closed source GPT 5.6 sol is up to date (June)...

Should potential purchasers take away from this that they're not going to be able to run recent models unless they front the cost of developing software or something?

eli 9 hours ago

I think they run whatever models they get paid to run. But mostly from enterprise. They are clearly not interested in consumer dollars.

gpm 9 hours ago

I mean the product is a server rack and while there's no advertised price I would assume it's six figures. So yes, an enterprise product.

But even an enterprise is going to care about the difference between "we can run the model we want with support from the manufacturer" and "we have to purchase the product, and then spend another 6 figure sum having developers port a recent model to the product to use it".

kube-system 8 hours ago

WarmWash 9 hours ago