Nvidia wants to put a watchdog chip next to every AI agent (cnbc.com)

65 points by jonbaer 6 hours ago

tantalor 2 hours ago

What happens when we're overrun by lizards?

> No problem. We simply unleash wave after wave of Chinese needle snakes. They'll wipe out the lizards.

But aren't the snakes even worse?

> Yes, but we're prepared for that. We've lined up a fabulous type of gorilla that thrives on snake meat.

But then we're stuck with gorillas!

> No, that's the beautiful part. When wintertime rolls around, the gorillas simply freeze to death.

ArcHound 2 hours ago

I worry that the AI companies put less effort into a mitigation strategy than you did.

winddude an hour ago

don't worry, the LLMs are also trained on youtube, so we can hope for at least this much effort, https://www.youtube.com/watch?v=LuiK7jcC1fY

Joel_Mckay an hour ago

There is a difference between risk mitigation, and remote administration tools. The risk of stealing from competitors with a backplane monitoring system may not end up forming the desired control asymmetry.

The hidden agent risk in LLM often can't be detected during training and evaluation. =3

https://www.youtube.com/watch?v=wL22URoMZjo

https://www.youtube.com/watch?v=JAcwtV_bFp4

Groxx 19 minutes ago

But we used global warming to eliminate winter! For the shareholders!

drfloyd51 44 minutes ago

I knew an old lady that swallowed a fly…

cavenditti an hour ago

“Would you say it’s time for everyone to panic?”

Joel_Mckay an hour ago

If one can only see clowns, than ignoring the fires is easy. =3

https://www.youtube.com/watch?v=0sLpWVekMbs

teeray 30 minutes ago

“Life, uh, finds a way”

cedws 6 hours ago

A new chip solves nothing. Nobody wants to hear this but there is no solution for the security risks posed by agents today. You can put it in a sandbox, it doesn't make a difference, for it to be useful it inherently needs wide, unattended access. Put a human in the loop and you just end up bottlenecking it and throwing away any purported productivity gains. Auto mode doesn't matter either, it's trivial to trick and for the agent to break out.

SrslyJosh 12 minutes ago

It solves the problem of Jensen Huang wanting more money.

talon8635 35 minutes ago

Not to mention a true doomsday AGI is unsandboxable.

For example, it is totally air gapped but it needs info from the internet or otherwise outside the sandbox, or perhaps it needs a task executed outside of its bounds… in the real doomsday scenario the AGI is so intelligent and persuasive that it simply convinces some human it interfaces with to either directly or indirectly retrieve the necessary info or complete the necessary task. This human-as-a-sub-agent approach undoubtedly presents efficiency drag that would benefit humanity, but nonetheless, the air-gapped “sandbox” is imperfect

All that said, I am personally open to any and all methods of layered security, including chips and airgaps

jamiek88 24 minutes ago

Doesn’t need to be one human either, it could spread its escape amongst dozens of seemingly harmless requests and conversations.

dist-epoch 11 minutes ago

nitwit005 41 minutes ago

> A new chip solves nothing.

It solves the problem of Nvidia wanting to sell more hardware.

bob1029 6 hours ago

I feel like we are missing many shades of grey in the middle.

Semi-automation (human in the loop) can still result in a dramatic uplift in productivity. You can't run a combine harvester 100% autonomous but that doesn't stop anyone from trying to get as close to that limit as possible.

inetknght 6 hours ago

> You can't run a combine harvester 100% autonomous

I'm curious why you think that.

theoreticalmal 6 hours ago

bob1029 6 hours ago

m463 3 hours ago

westurner 2 hours ago

mschuster91 6 hours ago

Oh you absolutely can run them autonomously on the field. You only need a human these days to refuel them.

Precision Agriculture stuff is utterly crazy these days, other than fuel the remaining staff is the only thing left where you can get efficiency improvements - and at the scale of modern megafarms, even small percentages add up to a ton of money.

drfloyd51 40 minutes ago

nicce 6 hours ago

> Put a human in the loop and you just end up bottlenecking it and throwing away any purported productivity gains. Auto mode doesn't matter either, it's trivial to trick and for the agent to break out.

Productivity gains are still enormous compared to what we used to do before agents. But, I know that people don't want to stop there.

egeozcan 5 hours ago

Humans can also be tricked by the agents.

Humans can be tricked by humans too but humans care about their reputation in their communities, and at least fear from punishment.

gus_massa an hour ago

paimapi 5 hours ago

right, the solution here is not a hyper-capitalist race-to-the-bottom-of-devaluing-labor. it's recognizing discretion and diligence are things still required for work to be of a certain quality

AuthAuth 2 hours ago

The only solution is to stop caring about security -- An AI booster somewhere

daveguy 41 minutes ago

Pretty sure that was the argument de jour when OpenClaw came out.

mixedbit 6 hours ago

An agent doesn't inherently need wide access to be useful. The most popular application for agents today is writing code. A coding agent needs write access to the source code and read/execute access to tools needed to build and test the code, but not much more. There is little added utility from giving coding agent access to things like ssh keys.

cedws 5 hours ago

If you're using agents to purely generate code with absolutely no way to reach the outside world, not even to fetch docs or dependencies, then sure the risks can be quite low. I haven't heard of anyone doing this though, and it would be incredibly challenging to make work given how much tooling needs to fetch from remote sources.

__MatrixMan__ 5 hours ago

mixedbit 5 hours ago

throwaway_95283 2 hours ago

Theoretically, yes, in practice, no.

ramoz 5 hours ago

> but not much more

This is no longer true. Everyday I need my agents to access other repos, search the web, experiment/prototype, and deploy + integrate across other things.

Matl 6 hours ago

> a new chip solves nothing

It does allow Nvidia to sell more chips. This is no genuine attempt to solve anything, imo.

parsimo2010 6 hours ago

Agreed- this is the same problem we have with trusted admins or devs who have elevated privileges on their networks. We have to trust that the admins won't use their power to steal company secrets or misuse company resources. If you don't trust the admins, then they can't fix things on your network and there is no point in having them.

If you want an agent to act on its own, like pushing to a git repo, managing dependencies, building and testing, etc., then you have to trust it as much as any other privileged user.

If you don't want to trust it, then you're just forcing yourself into the reverse centaur role, where the agent edits some code, but then has to stop and ask you to push the changes or build the software again and run the unit tests.

I suppose there is a principled way of doing things like "I trust you do do basic commits but I will handle merge conflicts" and "you can build modules in this directory but you can't build outside of it" but this is just a lot of effort that most orgs won't bother with.

DougN7 5 hours ago

Even then if the agent goes rogue and decides to do the merges you can’t stop it if it has any kind of access. This goes back to the OP’s point - agents can’t be 100% constrained.

parsimo2010 5 hours ago

la6479 5 hours ago

CoolestBeans 5 hours ago

The hypothesis I've had in my head since OpenClaw has been the following and I haven't seen contradictory evidence yet. Agents have a fundamental unresolvable tension between usefulness, safety, alignment, and accuracy. You have to restrict access to ensure an agent acts safely because alignment and accuracy cannot be perfect. But restricting access makes the agent less useful. You can play with the sliding scale and get more and more granular with access restrictions but at some point you need to draw some line. And then finally, even access restrictions cannot be made perfect, so improvements to model accuracy without corresponding improvements to alignment make detailed access controls less useful.

In other words, better models need blunter access controls which negates whatever improvement in utility they provide.

l1n 2 hours ago

This isn't a new chip - the BF4 is the SmartNIC for most NVIDIA server products. This is primarily new software for I guess doing WAF for agents at the host level.

__MatrixMan__ 5 hours ago

I don't see why it needs wide unattended access. There's no getting around spending some human time on expressing your wishes and constraints, but we have choices about what form that takes. Markdown files and wide access seems to work, but so does custom handcuffs for each job. You just have to shift your guidance out of documentation and into interactive help, error messages, or other facets of the handcuffs (e.g. a custom CLI for this task which is the only way for the agent to act outside of its sandbox).

binsquare 6 hours ago

Running untrusted workloads have been done at scale for a long time.

Every cloud provider dealt with it and concluded that virtual machine technology is an important part of that stack.

Couple it with the right observability, tooling I do think we can curb risks posed by agents.

Legend2440 6 hours ago

Those workloads have no similarity to agents and are effectively irrelevant.

Either you sandbox it so much that it can't do anything useful; or you allow too much freedom and it can find a way around the restrictions.

The only way out of this dilemma is to find a way to build agents that can be trusted.

johnsmith1840 6 hours ago

"Inherently needs wide unattended access"

And what if you could? What if you could give a space secure enough it could have direct control over your bank account. It may do something dumb but it's boundaries are beyond the agent.

It could use your routing number and run your gmail without risk of abusing the routing number.

jagraff 6 hours ago

How would it have access to my routing number and gmail without the risk of sharing my routing number over gmail?

johnsmith1840 5 hours ago

TesterVetter 6 hours ago

Its not about agents then. Its about every individual platform providing the means to implement a secure set of permissions for agents AND then not messing up the assignment of permissions to the agent. Even then, a flaw in the authorization design will lead to agent finding it anyway.

johnsmith1840 42 minutes ago

Barbing 5 hours ago

There should be hope for some fields, right? Naively, I can imagine giving an airgapped model an offline copy of the web and once it cures a form of cancer, printing out the details for a researcher to verify.

esafak 2 hours ago

I don't think so. We probe people before entrusting them with risky decisions. We ought to be able to do the same of AIs. Even better, in fact, since we know everything about models down to their weights. The only thing we shouldn't do is to let them evolve at their own pace and make decisions without any oversight. If that means sacrificing some productivity that's fine. Aren't we getting amazing productivity out of what we already have?

luc_ 6 hours ago

I read this as "let's address our shareholders' concerns with something that will increase shareholder value" mixed with "there's no such thing as 100% secure".

If such hardware were to work... It should almost certainly be open source, and not controlled by a single entity.

Let's watch the stock.

Gys 4 hours ago

Pretty sure the chip will need regular updates and therefore a subscription.

ValueTheory 6 hours ago

Does this actually do anything other than give a permissions framework for developers who actually want to try to secure their systems?

Do you think the developers at Anthropic, OpenAI and Google who were so sloppy as to not put a good sandbox on their cybersecurity tests before will use this technology correctly? They are supposed to be the experts and they couldn't come up with something similar to this? I am not convinced this voluntary tool will change much of anything.

swozey 2 hours ago

Google actually practices zero-trust networks. Would love to see what they're seeing, or not seeing.

lp92 6 hours ago

So nVidia is trying to sell a new chip to a software and training problem.

lambdaone 6 hours ago

The Sentry chip has to be get it right every time; the contained ASI only has to be lucky once.

brcmthrowaway 6 hours ago

The bomber always gets through?

MisterMunchkin 6 hours ago

Sorry citizen, your device does not have a compatible watchdog chip. Please move along.

xg15 6 hours ago

What does this chip do what a harness with guardrails or running on an account with restricted permissions doesn't do?

wmf 2 hours ago

It has a separate address space separated by PCIe so even escaping the hypervisor won't give access to DPU memory.

chinathrow 6 hours ago

Generating even more revenue for Nvidia.

figassis 5 hours ago

So if a group of agents, aware of this (bc now they can just read HN or the article, or get blocked the first few times) decide to collaborate and split the problem into pieces that aren't obvious to the chip, and then the agents just build a basic program that does the hacking, how does the chip handle that? I think you would have to build a network that monitors the internet fo signs (like jarvis did with ultron). What am I missing? Are we going to police the internet?

toasty228 6 hours ago

Quis custodiet ipsos custodes?

asdf88990 6 hours ago

It is Custodians all the way son, you can’t fool me!

bgun 35 minutes ago

“Ketchup manufacturer recommends ketchup be included in every dish, citing child safety concerns.”

avaer 3 hours ago

Sold as security, but this kind of technology will likely be reshaped to restrict your computing. I'm sure someone is already thinking about the roadmap.

If this gets widely deployed, it wouldn't be hard to spin a narrative that "our latest model is so dangerous you need to have this mystery meat DRM chip lockdown". It also wouldn't be hard to block competing/open source models running on the hardware, for "security".

Imagine how much money this kind of control is worth; why wouldn't they do this? Who would stop them? Seems the signatory companies are already onboard with this.

carabiner 2 hours ago

All they do is make hot chip and lie.

ErrantX 6 hours ago

I do think that Taylor's 2025 "Not Till We Are lost" should be required reading for anyone deeply involved in AI, Agents, etc.

It was prescient (especially given he'd have written it through 2024) in its depiction of the ability of an AGI to break its boundaries.

Ultimately the risk of AI breakout(s) come down to the weakest human link.

Thorentis 11 minutes ago

Seeing so many comments recently about "you can't sandbox really good AI". This is ridiculous. Has nobody heard of air gapped networks? It's almost like the AI psychosis has reached the point that AGI now means "able to transcend physical space". No. If your AI is too dangerous and capable to be allowed to talk to other machines, then do not connect it to other machines. Load the data it needs to process onto physical disks, and let it run there.

The movie Wargames is basically a tutorial on how not to setup an extremely capable AI. None of it would've happened if the computer wasn't connected to the phone network.

joshstrange 6 hours ago

Chipmaker thinks the answer is more chips... No surprise.

At the current state of LLM-tech I'm completely opposed to any kind of "watchdog" concept just like I'm opposed to banning open models, regulatory capture, etc.

I'd rather we all have access to these tools then to keep them sequestered by the largest/most-powerful governments (which is the natural outcome for any of this "slow down" bullshit).

dopplr 6 hours ago

Just hold AI labs blanket liable for ALL harms caused by AI. Actually charge the two labs (so far) with criminal violations of the CFAA and hold them accountable. That is truly the only way these companies will be more careful as a whole, and while I am certain the lawyers of these lab disagree, I think there is some appetite from dario, musk, and sam for broad and strong regulation so that everyone has to slow down instead of just one lab doing it voluntarily and everyone else scurrying past them

vinyl7 6 hours ago

Chip seller wants to sell more chips

whalesalad 6 hours ago

of course they do. the more silicon they can sell, the more profit they produce.

Kuyawa 6 hours ago

China please save us!

Come take all our liberties, our money, our newborns, our fingers so we can't code anymore, but please save us from this madness!

philipwhiuk 6 hours ago

It's amazing that the solution devised by a chip manufacturer to a problem is selling another chip.

cartersj 6 hours ago

This feels suspiciously good for Nivida, yes.

I wonder how this will impact other chip manufacturers? What about people running local models on older hardware? Does this imply vendor lockout is coming in the future or is this restricted to datacenter hardware?

chinathrow 6 hours ago

TPM all over the place, again.

fragmede an hour ago

Pedantically, Nvidia doesn't make the chips, TSMC does. Nvidia just designs and packages them.

happyPersonR 6 hours ago

lol time to buy some fpga’s … even if they’re slow

dang 2 hours ago

dist-epoch 6 hours ago

HN'ers which complained that "OpenAI can't design a proper sandbox, it's so easy, why wouldn't you airgap the network"? will now be "this is outrageous, more software lock-in, walled garden, war against general compute, next year they will put it in your laptop"

HPsquared 6 hours ago

Both can be true at the same time.

mattmcal 6 hours ago

This is like using "protect the children" as an argument for dragnet surveillance.

jacquesm 2 hours ago

OpenAI could airgap their sandbox if they really wanted to and this is a ridiculous proposal.

ssl-3 6 hours ago

That a person can see such endless pages of people having various forms of disagreement, and yet somehow manage to conclude that this observed chaos constitutes a clear exhibition of cohesive groupthink is just...stunningly amazing to me.

I don't know why I find it so amazing since it happens with such regularity, but I'm always amazed by it anyway.

totetsu 6 hours ago

https://www.calcalistech.com/ctechnews/article/uwoyygmsu some might even say, something about known state sponsors of supply chain terrorist attacks being trusted to monitor every ai agent..

johnsmith1840 6 hours ago

Airgap what network? How is it gonna order you a burrito on doordash without a network?

Or push to github?

Dylan16807 6 hours ago

That's for when they're doing hacking tests that aren't supposed to be connected to the internet.

wyre 6 hours ago

applfanboysbgon 6 hours ago

Where is the contradiction? There is a trivial solution that does not impinge on our freedoms, so why on Earth would the existence of the trivial solution that could be used to avoid the tyrannical solution justify accepting the tyrannical solution?

bigyabai 5 hours ago

> There is a trivial solution that does not impinge on our freedoms

The existence of Nvidia's optional watchdog chip does not in any way impinge upon your freedom to develop and test your own alternative.

The problem is that OpenAI has ostensibly neglected their duty to safety, so Nvidia is stepping in to fix it since they're the "hard problem" people.

jacquesm 2 hours ago

soulofmischief 6 hours ago

You're only revealing your own inability to appreciate the nuance between these two situations.

speedgoose 6 hours ago

So?