AMD acquires Taalas to boost inference performance by etching models in silicon (theregister.com)

464 points by itvision 8 hours ago

LarsDu88 7 hours ago

I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition.

Baking models onto silicon would've been the next logical move to get a moat.

Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

anthonypasq 6 hours ago

Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.

Melatonic 6 hours ago

The Taalas chips are not physically small. And part of their secret (if you look at the design) is just locating a bunch of memory soldered on the edges ( I belive higher amounts of SRAM ? )

chorizo 4 hours ago

selcuka 3 hours ago

adgjlsfhk1 6 hours ago

I don't think this works out from a cost/silicon perspective. Small models already run pretty well in software (since the weights fit in cache) and big models require silicon area proportional to the size of weights. On a mobile device putting a chip like this is competing directly in BOM and power against a whole lot more l3 cache, and the l3 cache makes everything faster

trebligdivad 3 hours ago

teaearlgraycold 6 hours ago

dboreham 4 hours ago

bastawhiz 6 hours ago

bsaul 6 hours ago

That's actually a really good point... There's currently zero incentive to buying more hardware, and that's one very good reason do have a new one.

sebular 5 hours ago

superb_dev 6 hours ago

From what I remember, these chips are not mobile size yet

bradfa 6 hours ago

makeitdouble 5 hours ago

Slightly besides your point, but it's interesting how many here naturally ponder about how the current winner could or "should" keep winning, instead of how another company could become a competitor by doing the more clever thing the incumbent isn't thinking about.

krisoft 4 hours ago

whatsThisBtn4 6 hours ago

Apple is somewhere between fashion company and second rate tech company.

They could have 9 year old AI and still post profits.

Not sure if it's my pixel or android, but I made a randos jaw drop with what the crappy AI on android can do.

When are we getting android OpenClaw?

moshun 6 hours ago

Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.

breuleux 6 hours ago

If you’re only running models for frontier capabilities, yeah. For tasks where current models are smart enough, running them 100x faster is the most impactful improvement you can make. Consider all the things you could use a model for, but don’t, because the latency is just a bit too high.

nowittyusername 2 hours ago

Depends on how much it costs the consumer. If I could buy a "cartridge" of Kimi K3 for 300 bucks I 100% would buy that shit asap. Even if it's "no good" after lets say 4 months still would be worth it IMO.

desmaraisp 2 hours ago

zxspectrum1982 5 hours ago

I'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.

subroutine 4 hours ago

andix 4 hours ago

NiloCK an hour ago

Gigachad 5 hours ago

topspin 6 hours ago

"seems like baking models into silicon is speed-running obsolescence"

Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.

mdp2021 6 hours ago

heywoods 4 hours ago

ray_v 6 hours ago

I could see this making sense when model development start to settle down ... it's going to settle down, right? ...

amelius 6 hours ago

Not sure. You can fix the transistors but leave the connections between them open for flexibility, so you only need to change the manufacturing process for the upper masks for every new model.

tsujamin 6 hours ago

sroussey 6 hours ago

mdp2021 6 hours ago

Compute the cost of producing n of them devices, imagine a fair price based on that, and see if that local, blazing fast card* can be an asset that could be replaced periodically.

*(It's local: private files managing firm oriented. It's blazing fast: it can be placed into recursive, intensive local workflows.)

alightsoul 6 hours ago

Which is exactly what companies and shareholders want to increase sales.

flyinglizard 6 hours ago

Look at it the other way: compared to the cost of training a model, the cost of making a custom ASIC is trivial.

try-working 5 hours ago

obsolescence is the whole point. apple gets to sell a new phone very 6-12 months because of it.

i have written about this:

"For device makers

Packaging models with laptops and smartphones will let application access near free, low latency inference and potentially offer users a better experience with the option of preserving data on-device. This is viable under the condition that tasks that do require larger expert models that run in the cloud can be routed to external models. A side-effect of local models and what will let Apple cut upgrade cycles from ~4 years (?) down to 12-18 months is specialized hardware to run them. For almost a decade, smartphones have been trying to compete on better cameras. This coming decade will see them selling better GPUs, NPUs, ASICs and whatever other things they'll be calling the inference chips, to drive re-purchase. Every six months will see a better model on new hardware, which will enable better performance in certain applications."

https://try.works/role-model-the-case-for-a-model-routing-pr...

nomel 5 hours ago

wraptile an hour ago

This seems like a very bad and dangerous direction for our society.

throwaway27448 5 hours ago

You need to find customers for several-generations-ago models before this makes any sense. AMD is a lot more incentivized to look than mr vanilla llm is

giancarlostoro 6 hours ago

ASICs is what took over Bitcoin mining, cheaper in all ways, and lasts longer than Nvidia GPUs for inference.

SR2Z 5 hours ago

> cheaper in all ways,

Bitcoin mining doesn't have large memory requirements, but does have huge compute requirements. ASICs work great there because it's very straightforward to add some circuits for computing hashes. If you _also_ have to add many GB of memory, then suddenly ASICs will cost as much or more than comparable off-the-shelf hardware and they won't be faster unless you've also invested in huge memory bandwidth.

giancarlostoro 5 hours ago

mrtksn 6 hours ago

Isn’t that kind of useless for the stock? It sounds complicated, unlike having number of CPUs go up.

It’s like talking about anything else than Megapixels when everyone was convinced that megapixels must go up in certain periods of the smartphone boom.

LPisGood 6 hours ago

I’m surprised Nvidia hasn’t partnered to make a Claude chip yet. It’s a win/win you can license them out, sell them when they become obsolete, etc.

UncleOxidant 3 hours ago

I guess I'm not understanding why this makes sense for AMD to buy Taalas unless they plan to get into hosting. It doesn't seem like a great fit.

CircuitSeuss 6 hours ago

mdp2021 5 hours ago

Not necessarily: it is relevant to Taalas only if it is a compute-in-memory architecture.

The Jalapeño mentioned («Anthropic is not alone in walking this path») in the article is still a classical Von Neumann architecture.

And Taalas' idea makes sense in a perspective of scale - producing a large number of cards; "for internal use" (a lower order of items) means a high production cost.

stingraycharles 2 hours ago

Didn’t Anthropic acquire Cerebras? Seems like a move into the same direction.

I also think that etching models into ASICs may be a bit too inflexible for what OpenAI and Anthropic want.

la6479 5 hours ago

Just to see how fast it is try chatjimmy.ai

tasty_freeze 2 hours ago

It is really fast and ... really hallucinates. I asked "Does the Wang corporation still exist? If not, what happened to it?" and it replied (in part):

"Yes, the Wang Corporation, the company that originally developed and marketed the Wang 2200 computer, still exists as a rebranded company under the name PPL (Precision Pencil and Label), but it has undergone significant changes and challenges over the years.

Here's a brief overview of what happened:

    Founding and Growth: The Wang Corporation was founded by An Wang in 1969."
In fact, Wang labs was founded in 1951. PPL seems to be a made up entity. But it did generate those "facts" in 0.033 seconds. If people value speed over accuracy then I can write an LLM that is 100x faster than chatjimmy.ai and make big bucks by responding one of N canned responses to any question.

mickaelkerjean 2 hours ago

mr_mph 5 hours ago

Pretty incredible to see. It reminds me of when I first used the Groq chatbot, except in this case it's a full response instantly.

alightsoul 6 hours ago

Because Openai and anthropic are not hardware companies. They outsource that to Broadcom and AWS' Annapurna labs.

wmf 6 hours ago

OpenAI and Anthropic are both designing ASICs.

alightsoul 5 hours ago

karmasimida 6 hours ago

A model can't be updated, and a chip that is only relevant for 6 months at max?

anigbrowl 6 hours ago

Depends what you mean by relevant. If you use AI primarily as a search/knowledge engine, it makes no sense. If it's your capable assistant that has a lot of general knowledge, can do tool calls, and has a big context window, very doable.

Indeed, for some kinds of applications involving secure/legal data etc. I can see the consistency of silicon winning out, because it combines performance with immutability and guardrails in hardware. Some chips have write-once PROMs to store password hashes and similar, you could do the same thing with prompt hashing to absolutely force or forbid certain behaviors. A model that can't be updated is also a model that can't be hacked.

askvictor 6 hours ago

People already buy new phones every year, this just creates even more reason to do so

Gigachad 5 hours ago

throwaway240403 5 hours ago

simpsond 3 hours ago

Base model sure, but the stack will be hybrid. It’s still early days here. Too bad FPGAs have such large feature size.

hamdingers 5 hours ago

One of these chips smart enough to take orders at a drive-thru would be relevant for a decade, minimum.

wolttam 6 hours ago

It's a terrible moat. You etch the silicon then nobody wants to run it in 6 months because models have advanced that much further.

nine_k 6 hours ago

Not so if it's embedded in something smart enough for its intended purpose.

Think vision, spatial reasoning, speech synthesis, even some speech analysis. Think self-driving cars (and drones) that need 10x less power for the brain, and can think at 10x situation per second.

anigbrowl 6 hours ago

This is only true for people who are solely focused on performance. There is absolutely a market for acceptable performance combined with predictability.

teraflop 5 hours ago

speed_spread 6 hours ago

If a model is good enough today, it's still gonna be good enough in a year. Except you'll be able to serve it 1/100 of the price. Or 100x the speed.

twobitshifter 5 hours ago

OTOH, people get a new iPhone every year and they are ok with it.

nomel 5 hours ago

bamboozled 6 hours ago

It googles models suck

linzhangrun 3 hours ago

Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.

dyzone 7 minutes ago

It tells me that they have some kind of insider knowledge that the models have hit their limits and won't be getting much better, and it makes sense economically speaking to just bake the current models and use them for the next 5-10 years. Looks like we're near the top of the S curve.

ilaksh 3 hours ago

Cerebras already runs large models like Kimi 2.6 or GLM at like 30x speed. 100 times is next year, not six years.

You can actually test it out on their website, just imagine 3 x faster and maybe 15% smarter.

kllrnohj 14 minutes ago

Cerebras is literally the entire wafer, so it can't get bigger. So where is the jump from 30x to 100x coming from? Node improvements only yield like 10-20% gains these days...

keepupnow 3 hours ago

This.

1saadcodes an hour ago

Feels both unreal and dystopian. The speed at which these models are developing is very scary

DiscourseFan 3 hours ago

It will be cool but also violent and terrible.

pizzaiolo 3 hours ago

So, like the present

barbazoo 2 hours ago

yassa9 8 minutes ago

Can anyone imagine if a video generation model with the speed of ASICs baked into silicon ? real Sci-fi

mNovak 4 hours ago

What I like about this, is that it significantly increases the probability of a sci-fi scenario where you're picking up a hot chip on the black market; rumor has it, Mythos 9 weights baked in...

bigyabai 2 hours ago

Plug it in, and it's a old prototype with Gemma 5 weights baked onboard. Dammit, fucked by Craigslist again!

NitpickLawyer an hour ago

Back in the kazaa and limewire days, you'd sometimes try to get a movie / episode from a series, wait hours / days for it to download, and when it was done you had a ~50/50 chance to actually watch what you wanted or an old german porn movie :/

ratsbane 13 minutes ago

Smart move by AMD. Chatjimmy is very fast and not very good, but I think it might become very fast AND very good.

msteffen 6 hours ago

This is neat but IMO a little crazy.

Something I personally haven’t seen much of, in all the discussions of model benchmarks and AI breakthroughs, is a distinction between “peak performance” and “reliable performance”. The “peak performance” of frontier models is very high: they’re solving open math problems, analyzing large codebases, etc. But my subjective impression is that “reliable performance” is mid at best: out of 100 random questions I might think to ask, it’s likely to say something wrong or stupid a handful of times at least.

I think there’s inherent tension between the two: the more a model reaches or outright hallucinates, the more likely it is to come up with tricky, subtle solutions to problems (I think people are somewhat like this too: Terry Tao’s brother is nonverbal, Jim Watson’s son has severe schizophrenia, etc). But then the less likely it is to generate a sensible email reply.

I use models all the time for coding, but I would not let one take over my daily correspondence. If the idea here is to run frontier models at high speed in data centers, that could be useful (the speed would be cool), but I’d be surprised if the cost of that hardware churn is worth it to frontier labs. But if the idea is to turn this into a chip that goes in your phone as some kind of routine, low-power inference thing…taking something too kooky to be relied on and baking it into your phone’s hardware like that doesn’t make sense to me.

dumberquestions 5 hours ago

I think you're underestimating both their reliability for standard problems and the usefulness of that level of reliability.

tyre 16 minutes ago

This is a good point. Opus does some silly shenanigans sometimes but then catches it later. It’s still an order of magnitude faster at getting to a working system than I am, for ones I don’t know.

It’s really a dream for setting up a homelab

daishi55 5 hours ago

> out of 100 random questions I might think to ask, it’s likely to say something wrong or stupid a handful of times at least.

What are some examples?

wmf 5 hours ago

There's a benchmark for this and a lot of models get negative scores because they're so unreliable: https://artificialanalysis.ai/evaluations/omniscience

daishi55 3 hours ago

yumraj 6 hours ago

Given the fast churn of the models, how does it work out?

Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out.

Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.

sigmoid10 6 hours ago

I find speed alone would be a game changer for current models. I hardly find any task anymore that the current frontier models can't do with max reasoning after several rounds of feedback (provided sufficient instruction and the right harness). But waiting an hour or more for reasoning to finish is getting really cumbersome. If they could do the same in seconds (and for cheap of course), I'm pretty sure we'd pretty soon see major software companies pop up that are run by a single human.

deadbabe 5 hours ago

Can you give some examples of these tasks that require an hour or more of reasoning?

xyzsparetimexyz 5 hours ago

craftkiller 4 hours ago

I think the real value here is not as a customer-facing agent/chatbot but for for automated processes. Think of all the companies out there that have LLMs doing simple tasks like categorizing customer feedback emails. For such tasks, you don't gain much from better models, so if you could run it 10x cheaper on a slightly older model, it would absolutely be worth it. Pretty much any place people are currently running a flash model could benefit from this since they're already deciding that speed+price is worth using a less capable model.

XCSme 5 hours ago

I think this would make sense for consumer hardware, not for AI companies.

AI companies constantly update/change stuff, new models come out, new requirements, etc.

But if you ship an "ai-powered" dishwasher, it can come with the chip built-in to do computer vision and precisely target each spot, and will be sold as-is with no updates.

yumraj 4 hours ago

Makes sense. Actually to expand, I believe this can make a lot of sense for industrial robots and such which have a more or less fixed job and latency matters more, so a well tested model may be more valuable than need to keep updating them

throwaway173738 4 hours ago

You don’t need this chip to do that. Computer vision has used machine learning for decades. The task you’re describing is pretty rudimentary and an off the shelf model with a control system would do it way cheaper.

tyre 13 minutes ago

XCSme 4 hours ago

m463 4 hours ago

subscription "ai-powered" dishwasher with personalized user ads, most of the chip dedicated to "personalized" not spots.

XCSme 4 hours ago

christina97 5 hours ago

There’s some kind of tradeoff between speed, cost, and quality for every application. I would be perfectly happy with a model 6 months old that was 50x faster for many uses. Right now I use either Opus (for smart stuff) or Flash without thinking (for fast stuff). I would take an even dumber model for more speed (lower latency in particular).

nullbio 3 hours ago

Perfect for consumers. You buy it and then you need to buy a new one in a couple of years. If they can make them affordable they'll sell like hotcakes.

etoxin 2 hours ago

And the second hand market. I'd love to see this integrated into motherboards like RAM. Someone could have a motherboard with 4 sticks of different AI with various models. Swap, change and trade.

chorizo 3 hours ago

That’s not going to be true forever. As models mature, we will hit diminishing returns. Major improvements will come annually rather monthly - matching the roughly annual release of new processors. Model ROM’s will likely get integrated into die packages just like DRAM now.

pennomi 2 hours ago

I’m hoping for SNES style cartridges

mrheosuper 2 hours ago

I'm still using Opus for most daily task because Fable is too expensive.

If they begin etching Fable into silicon now and release it 2-3 years later, i can see the market for it

noosphr 2 hours ago

This is a feature for most local use cases. You don't want all your work flows to start failing because of a model update.

brokencode 5 hours ago

Already models have gotten really good at a lot of things.

A lot of people would probably be happy to stick with the same model for a year or two if it’s 10x faster and cheaper.

And perhaps older models can become cheaper over time as newer models come out on new silicon for a higher price. That incentivizes people to stick with older models.

prinny_ 5 hours ago

They expect a sort of breakpoint at which each subsequent model version will only be marginally better than the previous ones, thus allowing them to retain their value for some time. Their business doesn’t work if each year the new model demolishes the previous one in terms of performance.

laweijfmvo 5 hours ago

pretty much everything is “1 or more versions behind” by the time it comes out. the question is whether or not it’s still useful? at some point, presumably not every application will need the latest cutting edge huge model.

casey2 2 hours ago

There isn't a fast churn in the underlying pretrained model, nor RL. It's mostly orchestration around the model. Said another way you could just pretrain and RL for longer.

Also I believe there is both a market for extremely fast local inference with current model performance and that such fast inference would unlock unforeseen usecases. Especially as TPS approaches early computer clock cycles and data rates.

deadbabe 5 hours ago

You could take your silicon chip and have it re-etched only with model diffs for an upgraded version.

yumraj 4 hours ago

How does that work, as in re-etching of silicon? Any pointers to read?

kristianp 2 hours ago

I've been eagerly awaiting their 2nd gen HC2, which uses multiple chips to host a "mid sized reasoning" [1] model. Its due in summer according to the article, I wonder if it will ever be released in that form now.

[1] https://www.forbes.com/sites/karlfreund/2026/02/19/taalas-la...

NitpickLawyer 12 minutes ago

> I wonder if it will ever be released in that form now.

Yeah, I had the same thought. The key thing for them was the price point at which they could deliver a ~30B model. I would buy one today if it was ~1000$ and could run whatever the best 30B model is today, at those speeds advertised. Even if the model becomes superseded by model.5 in a few months, there's still a lot of things you can do with a "good enough" model for some tasks. And things like maj@x or generate 10 times and choose "at a glance" what you like (think frontend stuff) would be worth it.

No idea if them selling to AMD is good or bad.

analog31 23 minutes ago

Wow, we're heading back to mask-programmed ROMs. I'm feeling young again.

hliyan 2 hours ago

Question: we currently emulate neural networks by performing matrix math in synchronous clock CPU architectures. Would it not be better to abandon synchronization and etch neuron synapses directly in silicon, keeping only the weights variable? I think some researchers are pursuing this, but I forget what the approach is called.

freakynit 36 minutes ago

"Neuromorphic chips" .... and I have the exact same question in mind.

zkmon an hour ago

I guess the idea is, gains from inference speed could offset the cost of upgrading the chips to a new model when really required. I think general purpose models would consolidate and release frequency might flatten out, favoring this strategy.

est 2 hours ago

Waiting for intelligence on a stick, plugin an USB, characters in, characters out.

100% local and no leaks.

matheusmoreira an hour ago

> Once the chips are deployed you’re stuck with that model.

At least we can be sure that's the model we wanted. Service providers could be serving modified versions and nobody would ever know.

redox99 6 hours ago

Is there any LLM from exactly one year ago that would be worth running?

In Aug 2025 you had

- OpenAI o3

- Opus 4.1

- Gemini 2.5 Pro

- Grok 4

Even if those were almost free to run, you'd be way better off with Deepseek flash 0731 or GPT 5.6 Luna, which already are almost free.

Other than for things where the t/s are critical, it seems like a bad idea to etch a model into silicon.

mdp2021 5 hours ago

> Is there any LLM from exactly one year ago that would be worth running?

Bad perspective: consider the correction: "when are thresholds of sought quality reached"? Hence: not "is there a 10yo from last year that could compete with the current 13yo", but "will there be a 30(?)yo from last year that could compete with the current 33(?)yo" ('(?)': the scale of yearly growth in the future is uncertain).

redox99 5 hours ago

It's not just about it "being smart enough". It's about there being actual user demand when it needs to compete with the shiny new model.

A 10 year old iPhone is probably good enough, but is there demand for it? In a vacuum a 10 year old iPhone is good, but why would you pick it if you can have a current one for a reasonable price?

singingtoday 2 hours ago

We still run GPT 4.1 for some of our use cases. We want to replace it but are having trouble finding models that are as fast with similar or better intelligence.

daishi55 5 hours ago

That is fkin wild. o3 was just a year ago? The progress is truly insane.

redox99 5 hours ago

Yeah I had to double check, o3 feels like it was ages ago. But GPT 5 came out Aug 7, so it's only one day off from my 1 year ago cutoff!

num42 2 hours ago

I have used chatjimmy before, it is incredibly fast, waiting for latest SOTA model on the chips in future. Great!

mikeayles 7 hours ago

AMD could have saved their money and used their own hardware! I've got a language model doing 60k tok/s on AMD hardware already, a Xilinx Kria K26 SOM, with the weights baked into URAM/BRAM with zero DRAM in the token loop. Same thesis as Taalas: single-stream decode is bandwidth bound, so stop fetching weights from far away.

Caveats stacked high, obviously. It's 3.16M parameters (tinystories, and I also have a kevin-speak lemmatised version), the tokens are characters, and the 60k record is 16 streams that each remember exactly one token of context, so it's blisteringly fast at saying nothing. The honest build with full context and KV caching still does ~19k tok/s on one stream though.

I keep messing with the blogpost with the live demo, but I'm planning on flipping it to live in the next day or two

Melatonic 6 hours ago

Yeah Im surprised nobody is talking about this. When everyone first saw Taalas I looked at the design and it had a big legup in physical cache availale compared to most chips. Makes you wonder how much of a benefit there is to the actual "baking" of the model vs just having a large chip with a ton of SRAM (or whatever) soldered close to the edge physically.

I feel like what we really need is the ability to solder computer cache on all sides of the chip Meaning above and below as well. If you can only attach it to the edges you will be inherently physically limited on the amount you can put (and maybe even have latency benefits as well)

Legend2440 5 hours ago

What you're describing is what Cerberas does.

Talaas is different, it's a true compute-in-memory architecture where the weights are stored in the connections between the transistors that perform the matrix multiply, rather than in seperate memory cells.

Most of the benefit comes from this architecture; hardwiring the weights into the silicon is just the easiest way to implement it. SRAM requires too many transistors, DRAM requires an incompatible manufacturing process, and exotic phase-change memories aren't readily available.

Melatonic 5 hours ago

wmf 5 hours ago

Taalas does not have cache so...

I agree that Groq with multilayer hybrid bonding could be a good idea.

tandr 7 hours ago

Well, technically it is their hardware now...

questionableans 7 hours ago

And their team, if they treat them well.

zxspectrum1982 5 hours ago

1. How come you didn't make your implementation public? You could be a millionaire now. 2. Especially if AMD has the technology to do what Taalas does, it makes a ton of sense for AMD to acquire Taalas: remove them from the market. Make sure nobody else (Intel, Huawei, Alibaba, NVIDIA, etc) acquires them. It could have been a great acquisition for a rebirth of BlackBerry btw.

ggm 6 hours ago

Field reprogrammable, it's an FPGA on steroids. Field upgradable.

Burnt in, it needs a zif socket and easy access in every car, aircraft, a pull out slot in a phone, or it's new era planned obselescence.

XCSme 5 hours ago

Why not have some a device/hardware that programs itself on-boot.

Sort of a FPGA, that (electrically) arranges the connections on-boot, and then it's like a static inference chip.

wmf 4 hours ago

FPGAs already configure themselves on boot.

XCSme 12 minutes ago

mdp2021 6 hours ago

Can that be done when the whole idea is to store a multiplier into a handful of transistors?

ggm 5 hours ago

I have no idea. It makes my comment a statement posted as a proxy for a question, a question you correctly pose explicitly.

If it can, then deployment in a sea of gates can make a chip viable across model generations as weights change, inside some scale factor.

If not, unless the part is under a pinout and address model which can scale on the bus, and can be easily replaced, it makes the entire dependency a replacement, not just this part. So embedded use has consequences.

xyzsparetimexyz 5 hours ago

It can just be pcie

andix 4 hours ago

It would be quite ironic if this technology would render all those AI data centers practically useless. If the next step are just a much smaller amount of expensive chips, and the bottleneck becomes manufacturing those chips fast. Not building huge data centers and fighting for electrical power.

yunnpp 2 hours ago

I would've hoped the company stayed independent instead of being engulfed into a behemoth. I'd like to see more diversity in the hardware ecosystem, but I guess the economics of hardware manufacturing aren't there.

3836293648 2 hours ago

They moved from HBM to dedicated silicon and only got a 48x speed up? That is so, so, so much less than I would've expected. Any numbers on how it scales?

preommr 5 hours ago

People are missing the point if they think this is useless because frontier models keep changing every few months.

We really, really need better secondary models that can do things fast and do them cheaply for lots of dumb tasks. Not only because it can be used as sub agents by frontier models, but also because it can be like a universal grease for all kinds of software.

I've got an app I am building and I don't want to tie myself with frontier models because I'll never be able to beat openai/anthropic. I just want a simple, cheap, instantaneous model that can just go through my documentation and tell the user what to do next and how to integrate with whatever ai subscription they have.

MarkWayneNewton 8 hours ago

While this design is self-limiting I think its a good approach. It doesn't take an entirely new architecture or infinite memory to produce significant performance improvement.

Legend2440 5 hours ago

This is a new architecture. It's a non-vonn neumann device.

bhouston 7 hours ago

Toronto Canada startup btw.

cmrdporcupine 7 hours ago

Seems to be somehow some kind of offshoot from or connected to Tenstorrent, which is just down the road. Founder looks like he was/is maybe at Tenstorrent and previously associated with Keller?

Always fantasize about applying at Tenstorrent, but wrong side of Toronto. 2 hour commute.

kridsdale1 7 hours ago

Works well, I remember driving by the ATI building as a kid.

ford 2 hours ago

I've been showing people chatjimmy for months - it's incredible. Both reasoning and tool use generation scale with TPS. Imagine 100x more reasoning on a model, or 100x parallel tool uses.

sgc 4 hours ago

What does it take to go from here to a model on a pcie card or an m.2 card, so I can plug one into my workstation / laptop? Will 'intelligence' become much like a gpu, where most people just live with the performance of whatever they have installed, outside large companies that must have cutting edge, or prosumers that have a incrementally better version than the masses?

Are we a couple years away, a decade away, or something else?

mdp2021 4 hours ago

> What does it take to go from here to a model on a pcie card or an m.2 card

It is already that.

> Will "intelligence" become much like a gpu

As an option among the implementations.

> Are we a couple years away

They could mass produce now, but it makes no sense at this rate of improvements in the models.

nojs 7 hours ago

Can anyone comment on the economics and likely turnaround times of this process, when it’s more mature?

Would it be realistic for a frontier lab to deploy this or would the turnaround time mean the model is always too out of date?

Assuming the weights and architecture are eventually stable, how much cheaper would this end up being?

2001zhaozhao 6 hours ago

There are always uses for outdated models.

Claude Code is still using haiku 4.5 from ages ago for explore subagents for instance. Not to mention production uses like customer service that only need to be "good enough"

edot 6 hours ago

Just looked this up, no longer true. Explore subagents inherit whatever model the parent is. And you can of course make other subagent configs.

samtheprogram 6 hours ago

AussieWog93 6 hours ago

alightsoul 6 hours ago

Customer service has really degraded huh. 4 years ago they expected opus performance out of human call center agents

I guess losing some customers due to poor customer service is ok if the price of customer service is right.

cogman10 6 hours ago

2 to 3 months optimistically assuming everything goes smoothly and is fully automated.

6 months or even a year if something goes wrong in the fabrication process and you need to update things.

If they do more standard asic design, it could be a lot longer as the design needs to be validated on an FPGA cluster, which would necessarily need to be very big for something like a LLM. Easily up to 2 years.

There's a reason chatjimmy isn't demonstrating newer models and why they only show of an 8B model.

shangofox 6 hours ago

I mean even if it take a few months, it'll still be out of date. But there was a hypothetical when it came up in Feb, would you want Qwen 3.5 at like 10k tokens per second.

At the time people were no doubt saying yes but now 3.8 is out, is that still desirable?

xienze 6 hours ago

There's soooo much stuff that such a model is still capable of doing in the pursuit of getting a better overall answer. Imagine a powerful research agent that blasts out dozens of the small, cheap models to fetch and summarize one page each. Then the beefy researcher model performs the final analysis.

syntaxing 7 hours ago

Honestly, this is starting to make more and more sense. SOTA models are starting to converge to certain architecture and capabilities. I wouldn’t be surprised we end up with a base model ASIC + “fine tune” card where it’s a physical LoRA style adapter.

encyclopedism 7 hours ago

Imagine a multi-modal model with 1000's of tokens per second. Realtime inference for a host of applications. This is a BIG deal and will change the landscape in unfathomable ways.

The https://chatjimmy.ai demo was impressive.

Once models settle down this makes sense. Imagine a cartridge with a physical model on it. You purchase a cartridge and stick it in your computer/phone/server. Want to upgrade? By a new 'cartridge'.

This should bring inference cost down dramatically, I wonder how OpenAI/Anthropic feel about that.

2001zhaozhao 6 hours ago

i'm looking forward to Qwen3.8 27B launch to see how much models have peaked at a given size.

it might already be time to start burning the best small models onto hardware since it's possible they can't get much better at many tasks like knowledge recall due to the inherent information density limits for models at a given size.

Grosvenor 7 hours ago

> Imagine a cartridge with a physical model on it.

I can finally have my own Dixie flatline. Cool.

mdp2021 6 hours ago

anthonypasq 6 hours ago

very interesting idea. i didnt think of that. i was just assuming youd have an additional one of these in your phone for actual lightning fast local inference

VladVladikoff 7 hours ago

Wouldn't this mean someone with sufficient hardware could lift the SOTA model weights off the chip? Or are you saying that these chips would only be used internally by these companies and not sold to the public?

dumberquestions 7 hours ago

I wouldn't expect companies not sharing their weights today to be any more likely to share them if they're on hardware, this doesn't sufficiently hide weights from a local user.

snek_case 7 hours ago

The weights are very unlikely to be on the chip itself. That wouldn't work for SOTA models that are terabyte scale, even quantized. This is probably an accelerator for specific kernels in the model, but the weights are likely loaded from memory. The chip may have SRAM to store some of the weights temporarily during inference.

foltik 6 hours ago

syntaxing 7 hours ago

I don’t get why this is an issue? You can run Claude/OpenAI SOTA models through Amazon bedrock. These weights have to live somewhere to run on Bedrock.

wmf 6 hours ago

amazingamazing 7 hours ago

One idea would be to use an open model.

kevin_thibedeau 7 hours ago

Then we can have machine psychologists pull cards when they run amok.

all2 5 hours ago

You have a robot. You need it to be smarter. You buy a new model cartridge (probably a PCIE 9.x). Now you need some domain specific skills. You'd like it to be able to cook, and you'd like it to not dent your walls anymore. You buy 'improved spatial reasoning LORA' card and 'Gordon Ramsey's Chef ULTRA9000' card.

Now your robot can respond sarcastically when you ask for chicken nuggets. Again. It also doesn't dent your walls anymore.

walrus01 7 hours ago

Having a base model ASIC as a physical piece of hardware makes me think of the early days of microcomputer desktop stuff where having a socketed ROM or PROM was a key piece of hardware, and people actually knew/cared what ROM was on their system's motherboard.

Imagine if like instead of having a specific Mac Plus ROM, you had a thing that looks like a fat ASIC that can hold models sitting on a slotted daughtercard directly next to the CPU and RAM.

breadislove 6 hours ago

we have not converged at all, if you look at how different the chinese models in terms of architecture you can guess that the labs are experimenting a lot as well. we are seeing all different types of hybrid architectures, different attention methods and so on. Of course on a high level its still a transformer but if you take a proper look we are seeing more divergence then a convergence.

smokel 7 hours ago

The technical aspects of SOTA models are not publicly documented. How do you know if something is converging?

syntaxing 7 hours ago

SOTA American models are not. SOTA Chinese models are. From a physics aspect, closed source models cannot be too far from open source ones in terms of size. There’s only so much you can squeeze out a B100 style cluster even with fancy Dflash style diffusion model for the speculative model.

_aavaa_ 7 hours ago

If we had deepseek v4 flash 0731 etched on a chip it would be more than capable enough and fast enough for so many people's needs, even hardcore engineer.

nurumaik 7 hours ago

cyanydeez 7 hours ago

if they were still exponentially increasing, they wouldn't be preparing for an IPO. IPO is where companies go to die and founders escape.

cyanydeez 7 hours ago

I don't think there'll be a fine tune card; you'll have the base model vintage whatever year, and then your GPU will do whatever LoRA layers you want it to do; the LoRA will wrangle older dated models into the current of whatever your looking at.

But yeah, for things like programming, if it can do linux and python and some go and sql and javascript, larger domains can be threaded with LORA

galaxyLogic 3 hours ago

I think the big news is that AMD is getting into memory-business so they won't be so dependent on Hynix and what have you. Memory is the bottleneck currently.

redmoonx 6 hours ago

It obviously won’t be continuous delivery but could make sense if the lifecycle of a model (train, deploy, iterate (meaningfully) is about 1-2 years. In that case it fits nicely in the “this year’s model” already established with cars, phones, etc.

laweijfmvo 5 hours ago

I’ve been using Gemma as my default (via Kagi) because it’s served on Cerebas hardware. The speed is honestly a game changer for day to day queries.

jackdoe 5 hours ago

Can you imagine in few years getting Fable level intelligence at 20k tokens per second?

"You are not prepared" --Illidan Stormrage

drob518 5 hours ago

So, Kimi K3 in silicon sometime soon?

roughly 5 hours ago

How's that jive with the fact that they're introducing a new model every other week?

drchickensalad 5 hours ago

The new model every week is not necessary at this point really. What if you could run opus 5 for the next couple years at 1/20 the cost?

roughly 4 hours ago

What's interesting about this is that I as a user would find this useful, but I think the AI industry as a whole would find it an absolute goddamn disaster. Opus 5 is a very good tool, but it is not a human-replacement-level intelligence, which means the entire revenue stream the industry's built on - labor replacement - is not met by this, and the only slightly charitable read of the industry's finances is that they're gonna bootstrap their way to creating the labor replacement hypothesis by getting people to spend money on Opus/etc, whereas if the actual product is a 1/20th the cost Opus-on-a-chip, the entire business and financing model that's tying up $N Trillion dollars of investment money goes out the window.

Great for us, looks like a recession as far as the Market is concerned.

jaggederest 5 hours ago

Pipeline the burn into silicon, lower the latency as much as you can, for the 10-100x operation cost it's worth it. Imagine if frontier models cost $5/mtok and the 2nd or 3rd tier models cost $5/billion tokens for 3-month-old models.

yousif_123123 4 hours ago

If things like this get traction, will we need all the datacenters?

downrightmike 3 hours ago

You are mistaken about what the datacenters are for

tecoholic 6 hours ago

With web search and tool call a decent current generation model at the speed of the chatjimmy could do a lot. People saying it would be out of date are missing the point. It’s not going to make much sense for frontier companies that’s chasing the SOTA. But for a lot of business use cases if someone can put GLM 5.2 and sell it as a box, it would make so much sense.

My partner has been asking for a “completely private” model for doing research and shifting through volumes of data that can’t leave the office and $$$ for the current hardware makes no sense. It would be an easy sell if someone walks in with a black box that contains “ChatGPT”.

5555watch 5 hours ago

In my understanding the first Deep Think / Pro models were already very good as they were doing some kind of parallel repeated reasoning, thus were slow and expensive. So if chatjimmy speeds enables a fast deep think level performance, I think that would be great.

equinumerous 6 hours ago

100% agree - you don't need the most up-to-date model to have something that's useful in agentic contexts. They could even produce chips with weights that make all the decision making/logical reasoning and have it delegate to other specialized agents. If it becomes cheap enough to print a run of custom chips, releasing a batch for each major advancement does not seem unreasonable for SOTA companies.

cephei 6 hours ago

There are so many use cases for supremely fast offline models. The first thing that comes to my mind is for real-time video processing or other non-textual content in real time.

anigbrowl 3 hours ago

I wouldn't call it supremely fast but zippy and versatile, yes: https://shop.m5stack.com/products/ai-pyramid-computing-box-p...

jauntywundrkind 5 hours ago

Core rope memory is back baby!

Enjoying the Ian Cutress / TechTechPotato video on Taalas. Some ok good technical details on the tech, and some good insider baseball, whose who stuff. (What a treasure having tech discussions like this about.) https://youtu.be/3MKRjt59hh4

OddMerlin 3 hours ago

Congrats to the Taalas gang.

tech234a 2 hours ago

See also: Twitter statement from Taalas https://x.com/taalas_inc/status/2085458427757937097

concraper 3 hours ago

A massive L for Canada

galaxyLogic 3 hours ago

"... the chip serve Meta’s Llama 3.1 8B at a blistering 16,960 tokens a second — when announced last February, that was 48x faster than Nvidia's GPUs and 8.5x faster than Cerebras' accelerators. "

bob1029 7 hours ago

I feel like NAND process tech could become useful at solving some of these problems. A GPU where you can update the weights a few thousand times may be sufficient.

mdp2021 7 hours ago

The basis of Taalas is "compute in memory" electronics - past Von Neumann's separation of processor and memory.

You need to be able to add|mul where the data (the weights) are stored.

addaon 7 hours ago

NAND hasn't been scaling great lately. It seems like PCM or MRAM would both be better fits.

kridsdale1 7 hours ago

FPGA model storage?

jijji 2 hours ago

taalas is great for llama 3.x 8B models, really bad for one board serving Kimi K3, it seems like you would bottleneck at a few hundred tokens no matter what you do.... spreading the big model against multiple cards seems the only way to get into the 1k+ tok/sec range. Another thing taalas is doing is masking the model weights into the silicon itself, not a flashable firmware, which would increase latency....

fellowniusmonk 7 hours ago

Token quantity will have a quality all its own.

ycui7 7 hours ago

so qwen3.x-27b on hardware? or better deepseek-v4-flash on hardware .

ilaksh 7 hours ago

I wrote them an email asking for PrismML Bonsai 27b Ternary which is like 6b or something crazy small and would be a lot easier for them to do initially.

mdp2021 6 hours ago

They were specializing their forthcoming system on 4-bit FP - which I understand is a structural decision.

Bonsai Ternary (1.7bits/weight) is a compromise, compromise that has to make sense in the context - efficient when translated into transistors.

andrewvl 6 hours ago

It must be a “super model”. What will be if new model released? New chips?

downrightmike 2 hours ago

Chip pops out like a gameboy cartridge. AI not working? Blow on it and jam it back in

api 5 hours ago

I've had an endgame idea in mind for a while.

Models, probably first open weight ones like Kimi K3 class, are etched into silicon like this and sold as cartridges almost like old school game cartridges.

You buy a USB-C dongle that the cartridge goes into, or for data centers you have PCI cards that take these in slots.

wmf 5 hours ago

Each cartridge costs $1,000. Do you still want it?

anigbrowl 3 hours ago

For fast Kimi K3? You're damn right I do

wmf 2 hours ago

singingtoday 2 hours ago

Yeah. I have 3 max20 plans.

api 5 hours ago

Me? Probably not. A business or a hoster, sure. There'd probably end up being an aftermarket in used cartridges with slightly older but still good models on them.

ur-whale 5 hours ago

Yeah, so https://chatjimmy.ai/ ... the model is crap, but the speed is amazing. Worth checking out.

walrus01 7 hours ago

Imagine the size of chip needed to 'etch' something like Qwen 3.6 27B in size.

golem14 6 hours ago

Interesting thought, because it's a yield question. How tolerant are models today to a few broken weights.

If tolerant, they could churn out many cheaper chips, some perhaps with slight abnormal tendencies ;)

thepasch 6 hours ago

> How tolerant are models today to a few broken weights.

Extremely! You can remove entire layers and the model will still work just fine, with barely perceptible capability losses.

I've cut/bypassed ~15% of total parameters out of Gemma 4 31B on a pod once. Still got perfectly coherent responses out of it. Certain layers are a lot more important than others, particularly early and late ones; but it's honestly astonishing how much can be cut out from the middle without destroying the model's coherence.

I didn't run any meaningful benchmarks, so I have no idea what the capability loss looks like exactly. But "produce coherent and sensible English in response to a wide variety of prompts" was definitely not among the things the model unlearned.

walrus01 6 hours ago

walrus01 6 hours ago

I wonder if you had a few percent of problems in the yield, if it would be functionally equivalent to the difference between a unsloth-published Q6 standard size GGUF vs. the nearly perfect precision of an unsloth Q8-K-XL. Or more like Q4 vs Q8 where a lot is lost.

mdp2021 6 hours ago

Not too dissimilar to the first HC1 (6nm 815mm² 53B Transistors embedding an 8b LLM):

> Our second model, still based on Taalas’ first-generation silicon platform (HC1), will be a mid-sized reasoning LLM

flog 6 hours ago

If someone has that sort of knowledge; how big a chip would be required? Is it possible?

mdp2021 6 hours ago

Well, given the data above, roughly a 220b transistors chip for the HC1 tech.

cubefox 6 hours ago

> At 20 billion parameters per chip, you’d need just 50 accelerators to support a trillion-parameter model

I don't see any evidence that this is possible. From my understanding, the whole model needs to be on a single chip. Which rules out any popular frontier models with several trillions of parameters. Even smaller sub-frontier models have hundreds of millions of parameters, so these would be ruled out as well.

octoberfranklin 43 minutes ago

They pipeline-parallelize across multiple chips. DeepSeek v4 Pro will be 30 chips.

wmf 5 hours ago

The methods for splitting weights across multiple chips are well established. Groq/Cerebras can't hold a model on one chip either.

pyrolistical 3 hours ago

Umm I have an extra 35, do you have layer 6?

IsTom 6 hours ago

I think it's enough that a single layer fits on each chip if you can daisy-chain them with good interconnects.

moralestapia 4 hours ago

Taalas is just a phenomenal startup from Toronto. My dearest congratulations to the founders.

Edit: Lol, downvotes? Stay jelly, meanwhile Talas goes brrr.