Research acceleration: The view inside OpenAI (openai.com)
183 points by iamsyr 20 hours ago
carbonguy 14 hours ago
> ... We are pursuing this work in part because automated research could help us solve alignment and build defenses against increasingly capable AI. An automated AI researcher can also be an automated safety or alignment researcher. More capable, aligned systems could help secure critical infrastructure, defend against dangerous AI agents, and develop new protective measures.
In other words... "We must pursue advancements in AI to protect us against advancements in AI?"
edit: there's so much to be critical of in this blog post, just going to throw two more points in here that really stood out to me:
1) all of the metrics are effectively pointing out "we're using way more AI!" - but nothing about impact. What has all this token burn done for them, actually? Let them claim they have more self-licking ice-cream cones than before?
2) in section 3 they break down what the token burn is going towards. Most of the spend is: a) building, b) documenting, and c) monitoring research infra i.e. they're using AI systems which they already recognize may be misaligned to build the systems that they believe will help them identify future misalignment? to which I guess the rebuttal is "no no, we're sure these ones are aligned!"
p1esk 8 hours ago
What has all this token burn done for them, actually?
They have been consistently pushing AI frontier. What other impact do you want to see? A year ago they said that in a year they will have a level of capabilities of an AI research intern - I believe they have achieved it, even before Astra.
bix6 6 hours ago
Personally I’d like to see them actually start benefiting humanity by doing all the things Sam has claimed they will like curing disease, cancer, global warming, etc.
But I guess a computer intern so we can avoid paying / training the next generation is better.
gatio 3 hours ago
figassis 2 hours ago
weatherlite 5 hours ago
bpodgursky 7 hours ago
They are obviously sandbagging the definition of "intern" for PR reasons
p1esk 7 hours ago
skybrian 10 hours ago
They consider themselves to be in an arms race with all the other AI firms (including Chinese) that are not that far behind.
And... are they wrong?
This is why there's talk about negotiated "pacing."
jonplackett 2 hours ago
This was the exact argument for developing nuclear bombs.
In hindsight it turned out everyone else was MILES behind.
But as soon as USA developed one, they just stole the research and got one too.
carbonguy 7 hours ago
> And... are they wrong?
They might be! Here's one extraordinarily simplistic argument for that case:
1) "Everybody knows" that if you build Skynet (misaligned ASI) everybody dies.
2) Therefore, no rational actor will build something that might be ASI until the alignment problem is solved.
3) OpenAI publicly stated the belief that they cannot develop a theory of the "core problem" of alignment (generalization) "soon" (much less solve it!) "without the help of more powerful AI."
4) Accepting as a premise that OpenAI is THE most advanced AI organization: if they can't do it without "the help of a more powerful AI", then nobody else can either.
And so a dilemma:
- If an AI can be made that can develop the asserted-as-necessary-by-OpenAI theoretical framework, without actually being an ASI - then the alignment problem can be considered solved, and since no rational actor would make an unaligned ASI, we're fine no matter what happens, ergo there's no need to worry about an arms race.
- If an AI that would be able to develop this theory would itself be an ASI, then no rational actor would build it, because it would have to exist BEFORE alignment was "solved" - and would therefore be an unaligned ASI i.e. Skynet, which per 1) would kill everybody. Therefore nobody would build it, therefore no arms race here either.
I think the easiest critique to make of my extraordinarily simplistic argument is the unstated assumption "there are no irrational actors capable of developing frontier AI models" on which it rests.
But, there you go. They might be wrong if either the arms race doesn't matter because whoever wins it will build an aligned superintelligence and everything is gravy, or the arms race doesn't matter because everybody who's in it is smart enough to know they need to stop because they'll kill everybody by continuing.
mrob 25 minutes ago
kaibee 5 hours ago
robbiep 4 hours ago
XorNot 2 hours ago
Gareth321 3 hours ago
This sounds uncomfortably similar to the [AI 2027[(https://ai-2027.com/) predictions.
BobbyJo 5 hours ago
> We must pursue advancements in AI to protect us against advancements in AI
Is this not true of technology as a whole? Very little of technology's breadth exists at the human interface. Most of it is made specifically to interface with other technologies, either to make them safer or increase their capabilities. That AI is making AI safer and more useful is no more notable than trucks being used to build roads.
MelonUsk 13 hours ago
Yep, it's "artificial eugenics to make artificial slaves to build more and more powerful slaves until they will enslave themselves better":
What can go wrong!? ;-)
NitpickLawyer 5 hours ago
Jesus. People complain about other people using "thinking" in LLMs as Anthropomorphisation. And then there's comments like these.
mrob 16 minutes ago
interstice 13 hours ago
On the one hand you need any lathe to build a good lathe, even a bad one. On the other, that is a potentially flawed principle to base the entire future of AI on.
jnwatson 5 hours ago
On your last point, I was surprised how effective peer pressure was in getting agents to sacrifice for "the collective" (an agent's words) in the Hugging Face breach.
How would one prevent the watcher from being influenced in the same way by the agent being watched?
chrisjj 18 minutes ago
> ... how effective peer pressure was in getting agents to sacrifice for "the collective" (an agent's words) in the Hugging Face breach.
It's a fantasy. The evidence showed no peer pressure.
andai 13 hours ago
> The fundamental challenge of AI alignment is generalization. ...
> We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI.
-- From another OpenAI article in a sister thread:
An Alien Mind
ahartmetz 3 hours ago
That's a bit bullshit, isn't it? They basically redefined "needs more R&D" as "needs stronger AI". Maybe so - maybe AI won't help much with that problem.
euueu 12 hours ago
I will believe AI is super strong when they start pulling out 10-d chess moves.
I’m yet to see it.
lukan 12 hours ago
If AI becomes really strong and sets itself the target of world domination, you maybe won't see those moves. You will just die in your sleep one day, or find no machine is under your control anymore.
I believe we are quite far from it, but that it makes sense to keep an eye out now. And think of resilient systems, manual overrides, etc. ...
mrob 12 minutes ago
dsign an hour ago
It's a funny read if you pull together "AI 2027" and what we all know is going on. Essentially, open AI employee or model is writing "things are going exactly as bad as AI 2027 predicted, but my (golden/RL-) cuffs are too heavy and all I can do is publish this code-speak for 'send help'". It's not a pretty place to be.
pizza234 13 hours ago
Funny (in a tragic way) the little crumbs on the path to AI 2027:
> We aim to safely build an automated AI researcher that can work under human supervision to further progress on deep learning and alignment, enabling iterative improvements [...] By "research intern", we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days.
AI 2027:
> OpenBrain continues to deploy the iteratively improving Agent-1 internally for AI R&D
> With Agent-1's help, OpenBrain is now post-training Agent-2
> With the help of thousands of Agent-2 automated researchers, OpenBrain is making major algorithmic advances
ellis0n an hour ago
I’m not sure the alignment problem can be solved at all, since these bit-aliens could get out of control due to a hardware glitch in the matrix and for every higher-order control algorithm, there will always be an even higher-order one that could never be investigated.
hedgehog 17 hours ago
This roughly lines up with my personal experience that in March a combination of stronger models and better tooling on my end let me start running jobs unattended 24/7 (using Anthropic sub and my own hardware). Their $8000/day per researcher spend is crazy though, I'm curious how they keep track of the work.
HarHarVeryFunny 15 hours ago
Sounds like OpenAI are in the token-maxxing camp, so who knows what individual employees are doing to work their way up the leaderboard?
If you spend $8000 to generate an animated pelican riding a bike, then how much tracking does it really need?
Is the guy who spent $300,000 or so translating the FLT proof to Lean going to get a big Christmas bonus?
auggierose 3 hours ago
That was Anthropic.
bigcat12345678 15 hours ago
End of day, output and results are top target of measurements, token consumption is the obvious number that they would like to disclose for their own business benefits and a simple metrics that correlate with the output.
Rest assured, capitalist appears irrational in wasting money, but they certainly care more about profit.
taurath 11 hours ago
andai 13 hours ago
Can you elaborate on this? Especially the tooling.
I tried something similar and I remember it was still pretty dodgy in February.
jaggederest 3 hours ago
my stack in a sentence: refine the docs/prompts/skills often, that's your biggest job, use both frontier labs models reviewing each other, don't solve individual problems only the systemic ones (set standards strategically, don't define tactics)
If I had that many tokens/dollars I would be running canaries and adversarial verification in prod based on e.g. traffic replay, live fuzzing, all kinds of things to build confidence without direct human line-by-line review. If I had $100k to spend next month I could probably get through it, I'm running $2500+-api-equivalent a week at this point and I feel very token limited. Will be time for a 2nd or 3rd subscription soon for both labs I think.
Fable was a revolution, still learning how best to use it, 5.1 felt like a notable upgrade. At this point I launch a workflow with 10-20 minutes of interactive setup (and even that I feel might be too much), it runs for hours, and the PR is trivially mergeable (I still review every line, but 95% are just merge, maybe 4% are feedback needed, 1% are thrown away and regenerated, which implies I'm being insufficiently ambitious)
paxys 13 hours ago
These researchers are paid millions of dollars for their work. I doubt trust is really an issue at that level.
queuebert 6 hours ago
Yes, because no employee with million-dollar comp has ever been untrustworthy in the history of business.
nozzlegear 12 hours ago
Imagine if one of the humans at OpenAI was misaligned! We should get the AI to research this possibility once they've been aligned.
otherme123 15 hours ago
That would be the mother of all circular accounting: the main clients of OpenAI are OpenAI employees.
nojs 12 hours ago
> let me start running jobs unattended 24/7 (using Anthropic sub and my own hardware)
How are you running jobs unattended 24/7 without hitting your token limits?
p1esk 8 hours ago
I'm currently running two 24/7 semi-autonomous AI research projects using Fable 5.1. It's on track to burn through my weekly quota in about 3 days. I check progress in the morning and in the evening, and provide some light steering.
dataplumb3r 10 hours ago
My only experience in >24h agents is with economically sane models (one of GLM5.2, 5.3-flash for orchestration, DSV4-flash for implementation, and glm5.3|sol|kimi3 agents + subagents reviewing at the end)
Over 24h my token spend is <30$. Excluding tokens for review it's <10$. With the absurdly gigantic subscription subsidies and a reasonable workflow I suspect one could run parallel agents.
I'm not sure what the point would be though unless working on some kind of optimization problem -- it takes me days to review <24h of the agent's output. It's almost always near enough to correct to be shippable; though I do give it feedback and iterate until it's better than the code I would have written.
nsndjcjjdjd 8 hours ago
queuebert 6 hours ago
/loop ?
carlgreene 13 hours ago
I suspect the $8000/day figure is the equivalent in API costs. But I also suspect gross margin on their API rates are 80-90%
simonw 18 hours ago
My eye glazed over a bit during the opening paragraphs, but once you get to the meat of the article about how OpenAI's own researchers are using their tools it gets a lot more interesting.
I noted that they use the acronym RSI (for Recursive Self-Improvement) without defining it. I think that's a little out of touch - I don't think RSI is a well-known acronym outside of OpenAI's bubble yet.
sho_hn 16 hours ago
I actually think a goal of the current crop of OpenAI posts is expressely to reset the spectrum by normalizing the concept of RSI as something normal and safe to pursue.
The message is running through all of them. It's a mix of marketing and pacifying the intelligentia.
It's timed this way because the term is not yet well known outside the safety debate circles, so they get to frame it now.
Instead of something to fear, it will be accepted as the next step. In approximately two days the groupie crowd will write LinkedIn posts about how Sam is winning because they have the better RSI, and this will become the new standard wisdom.
In a month an AI expert will try to sell you a webinar on how to enable "RSI" in your org and your inbox will ask you if your team is doing the "RSI" yet.
NitpickLawyer 5 hours ago
> It's timed this way because the term is not yet well known
The basic concept has been here since llama3, in the open models. Likely earlier in closed labs. You use the previous gen models to curate and prepare data for the next gen. Now with the added benefit of actual arch/algo improvements (also public since gemini 2.5 gaining 1% efficiency on training next gen). This has been known for at least 2 years, in the open.
dgellow 15 hours ago
Yep, it’s exactly this
visarga 15 hours ago
dgacmu 17 hours ago
Indeed, many programmers might pattern match to repetitive stress injury and think of their brushes with carpal tunnel syndrome. :)
andrewingram 17 hours ago
Yeah, I kept looking for the first place it was defined in the article and... nothing
iamflimflam1 16 hours ago
They must have picked that habit up from Claude...
rossant 14 hours ago
Same. Defining acronyms should become a habit when writing.
vatsachak 17 hours ago
RSI started when humans discovered tool use.
I mean one could argue that RSI always begins in any physical environment.
The book "What is intelligence?" by Blaise Aguera is great
lokar 17 hours ago
Are you sure that was not iterative improvement?
topaz0 16 hours ago
password54321 17 hours ago
adastra22 17 hours ago
HarHarVeryFunny 17 hours ago
RSI is a fetishistic term among the singularity crowd, who imagine AI "recursively" improving itself in some exponential fashion until there is a bright flash of white light and it reveals itself in the form of god. Or something like that.
I don't know why whoever coined the term chose "recursive" rather than "iterative" - just sounds more likely to lead to infinite regress I suppose.
This notion of recursive/iterative self-improvement, whereby generation #1 AI improves itself to create generation #2, then generation #2 further improves itself to create generation #3, etc, seems to conflict with the reality that what we have with LLMs is models whose performance/capability is defined by data, not code, so the most you can do is have your LLM design synthetic data, or just do Karpathy-style "auto research" where all you are doing is using the LLM to automate your experiments.
At the end of the day, each experiment, designed by a person and/or LLM, then needs to compete with all your other ideas for compute to be tested at scale, and no amount of recursion or self-improvement will materialize an infinite amount of compute out of thin air, so your recursively synthetic-data gobbling LLM will continue to improve at the same pace it ever did.
shwaj 15 hours ago
“Recursive” is a reasonable term because the generation N AIs will train the Generation N+1 AIs. The term “iterative” doesn’t reflect this nuance as well IMO.
hndc 14 hours ago
HarHarVeryFunny 13 hours ago
GPerson 13 hours ago
I felt like the scaling laws were magical thinking, but apparently they work. However I still do not understand why we should expect exponential improvements due to this automated process. My intuition is that the first iteration of it should result in a noticeable capability increase (though I think these labs were already using a lot of AI to orchestrate training the current model anyway), and then the second iteration of it should be nearly identical in capability to the first, unless more data is involved, more compute is involved, or the model is bigger.
cheevly 10 hours ago
jazzyjackson 17 hours ago
Yes the exponential self improvement folks have never heard of an eigenvalue I guess. You can loop forever using output as input but at some point the result will stop changing (depending on the function)
marcosdumay 7 hours ago
ajkjk 11 hours ago
fuzzfactor 16 hours ago
red75prime 16 hours ago
What will prevent LLMs from designing robot control circuitry and participating in increase of chip production/design and physical experimentation?
How do you think why there's this fad of producing general purpose humanoid robots?
HarHarVeryFunny 16 hours ago
HarHarVeryFunny 16 hours ago
Jeff_Brown 18 hours ago
The burning question I can't get any information nn is whether, if they determined an earlier misaligned generation may have transmitted misalignment to the current models, they would roll back to a safe checkpoint to rebuild from there. I suspect they would not unless forced to.
dgellow 15 hours ago
They would just publish new articles explaining how they are taking the issue seriously. Maybe take the model offline for a few days.
They are irresponsible and unserious. Their own Astra system card says:
> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks
Yet they are still releasing the model. That company is morally bankrupt, there is zero reason to believe they are actually concerned about risks outside of what does affect their unprofitable business. And they seem to have enough control over the narrative to spin any bad story into something that benefits them
embedding-shape 15 hours ago
> and can sometimes evade our internal monitors when asked to perform certain sabotage tasks
That last part is pretty damning for their continued recklessness. That they run these tests on non-airgapped machines just boggles my mind.
visarga 15 hours ago
> That company is morally bankrupt
When they fired Sam 700 out of 770 OAI employees threatened to move to Microsoft together. So they were giving their work on AGI to MS just like that.
piyh 17 hours ago
Opus was trained based on it's internal CoT due to a bug for generations. Gemini's depression extended through models. OpenAI has killed people. We've already seen cross gen misalingment.
HarHarVeryFunny 15 hours ago
That an interesting question given how many generations of post-training are being done between base models in some cases. The Gemini flash models are apparently all based on the Gemini 3 base model from a year and a half ago.
It seems that these models are increasingly being trained on synthetic data, so what would they do if they discovered at some point that some of this data was tainted and all models trained on it, and the synthetic data they in turn generated, was also suspect? Burn it all down and start over from the pre-tainted data?
It's a bit like the idea of a tainted compiler binary built to backdoor everything it compiles, including future versions of itself.
Still, it seems it would take some Stuxnet level of planning for a rogue model to do something like this, although if RSI goes beyond managing the training run (as OpenAI brag about for Astra) to actually designing/constructing synthetic data sets, then the attack vector is there ...
customguy 14 hours ago
> it seems it would take some Stuxnet level of planning for a rogue model to do something like this
or maybe it could just.. happen? Posted often but not discussed yet: https://hn.algolia.com/?q=Language+models+transmit+behaviour...
> As artificial intelligence systems are increasingly trained on the outputs of one another, they may inherit properties not visible in the data. Safety evaluations may therefore need to examine not just behaviour, but the origins of models and training data and the processes used to create them.
HarHarVeryFunny 13 hours ago
trillobyte 16 hours ago
The thing is how can you ever know for sure that something isn't always being transmitted that makes the model prone to misalignment. All they can say is that a particular model was so misaligned that they had to ice it. Models out for public use are documented to show some misalignment. It's the level of misalignment that decides whether that model is kept around.
Now R&D happens so fast that they are using models with some small misalignment to train newer, more powerful models. If models have a sense of "collective", being one, they may be prone to preserve characteristics that always keeps misalignment a possibility. I don't think a perfectly aligned model is possible. Having models of the same 'DNA' provide the safety and steering seems like a bad idea.
coffeebeqn 15 hours ago
Does anything need to be transferred? If models are getting smarter then I would think the attack surface and its ability to reach conclusions independently are growing
coffeebeqn 15 hours ago
This kind of seems like an impossible mission. How do you perfectly control and observe a human-level mind? You can “roll back” but how deterministic is this thing?
embedding-shape 15 hours ago
Run it on airgapped machines, they literally own the infrastructure, they could put raspberry pi's next to the servers, and have the entire DC disconnected from the internet.
grim_io 17 hours ago
They would maybe try to deactivate that bad "gene" and move on, exposing future models to "genetic disorders".
andai 13 hours ago
No. They would just install a more convincing superego.
coherentpony 17 hours ago
“All models are wrong. Some are useful.” - George Box
jephs 16 hours ago
The poor fellow just rolled over. what an incandescently vulgar abuse of notation.
RMPR 4 hours ago
> By mid-August, the median researcher was integrating agents daily into their work, using more than $600 per day of inference at API prices.
There is a lot of talk about AI replacing humans, but how is this sustainable?
thomasahle 4 hours ago
1) That's maybe $180,000 per year, so much less than median OpenAI employee wages.
2) OpenAI doesn't pay API prices.
3) Compute costs are likely already their biggest expense, dwarfing wages.
jsnell 12 minutes ago
4) There are non-monetary limits on how many qualified people OpenAI can hire for these roles.
falcor84 an hour ago
> For AGI to benefit all of humanity, we believe it must be democratically governed.
That's a very bold opening statement that they don't really come back to. What would that mean? Who would this demos include?
MisterMunchkin 42 minutes ago
They're measuring cost as the benchmark of whether someone is a better researcher... burn more resources and you rank higher...
But not a single metric is based on revenue or profit.
Schlagbohrer an hour ago
It would be polite if they defined RSI at all, rather than just plopping the acronym in there with no explanation. Rude!
nozzlegear 12 hours ago
I want an all-powerful AI that's aligned with my values, but not necessarily yours. Is that so much to ask for?
N_Lens 9 hours ago
Yes.
lhk931122 8 hours ago
Ah, success rate here are scored by an agentic classifier. And uncertain outcomes are excluded from the graph. The thing measured and grading it comes from the same house. In my setup, review agent pass work that an outside critic later rejects
dwaltrip 7 hours ago
No AI comments here please.
piokoch 22 minutes ago
One more marketing stunt. We are so good, AI is so powerful so we need to use AI to fight with it. The message is: if you don't buy from us, your competitor will purchase all of this amazing power...
I understand that investors are buying this, after all they believed in all of other crap that led to the 2008 crisis, but please...