Sharing AI progress in mathematics (openai.com)
1205 points by OfficialTurkey a day ago
jboggan 19 hours ago
I was a graph theory junkie long ago and even moved to Budapest for awhile to study among the greats. While I was there I started working on Barnette's Conjecture which came to occupy my thoughts over the next 24 years of my life, on and off as I worked in many different fields. Last summer I even thought for a few days that I had actually solved it.
But it's supposedly proven here - problem 180. I don't know what to think exactly. I spent thousands of hours on that problem. I really enjoyed it. Hearing that it is solved somehow makes me sad in a far-off way, like hearing an ex-girlfriend died suddenly in a car crash. I don't know, there's probably a lot of people feeling odd emotions tonight.
There's no Lean proof for this one so I'm digesting the paper. On the surface it looks like an approach I considered 24 years ago and abandoned.
I revisited the problem this summer, along with my partial solutions, when the previous round of stunning proofs came out. Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was.
nilkn 17 hours ago
> Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was.
This is the part that gives me the strangest feeling about it all, because you're not the only one with this experience. I've experienced this too on different problems, as have many researchers across many fields.
I disagree with the Fields Medalists on the majority of their complaints. AI math is happening and there's no going back. However, on one point I increasingly agree: virtually none of this stuff is possible with technology any normal citizen has access to. I have no problem with AI models making revolutionary advances in math or science. Where I start to have a problem is when the AI models making these advances are tightly withheld, proprietary, and seemingly never released with these capabilities intact. This has been the case for all of 2026 so far.
I suspect that this is in fact the source of much of the angst. None of this progress is reproducible outside of one or two teams inside OpenAI and Anthropic. It's becoming an incredible concentration of power that I don't know that we've ever quite seen before. Right now, it feels harmless because it's being used for wonky math problems that aren't (yet) practical for anything. But great power never stays harmless. History has taught us that countless times, in countless different forms.
omnicognate 16 hours ago
> I suspect that this is in fact the source of much of the angst.
Why do you "suspect" this as if it's some hidden motivation when the very first paragraph of the advisory group's statement (linked from the OpenAI post) says:
> At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.
Tao and others in that group have been strongly and publicly pro AI from the start. They are not advocating "going back". They're objecting to the strip mining of open problems using proprietary technology.
alberto-m 13 hours ago
fasterik 8 hours ago
hex4def6 an hour ago
pjc50 12 hours ago
pred_ 13 hours ago
lordgrenville 15 hours ago
hawk_ 16 hours ago
devin 3 hours ago
There has never been a stronger need for people to band together and "seize the means of production" for this stuff. The advances being made are ours, not theirs. It's trained on our work, our knowledge.
heaney-555 3 hours ago
dannyw 17 hours ago
We _think_ this power / divide feels harmless right now, but I'd bet money that NSA, CIA, etc have access to the latest and greatest unrestricted models; and massive compute. At least for OpenAI, and even if not willingly for Anthropic, I'd bet money NSA has it too. (After all, when Google decided to migrate to HTTPS, the NSA decided to hack Google's internal network to preserve their taps).
Who knows what they are up to.
schoen 16 hours ago
jrflo 8 hours ago
baq 16 hours ago
Jenda_ an hour ago
> virtually none of this stuff is possible with technology any normal citizen has access to
So far, it looks like open-weight models are lagging less than a year behind frontier capabilities. And I think one year diffusion of technology from "insider lab demo" to widely available is actually pretty fast?
There are lots of research fields which "normal citizen" has no access to - medical and biological research, particle physics. Some of it is somehow publicly controlled (LHC), some of it not at all (commercial pharma research, mostly secret until the final human trials). And most of it reaches "normal citizens" in way more than a year.
(and I'm talking about open-weight models. The availability of commercial AI models from private preview to included-in-your-$100-subscription is currently like 4 months)
markus_zhang 15 hours ago
I'm wondering what's the impact on human Mathematicians, and especially would-be Mathematicians -- master students, if they HAVE to use AI in their daily life?
Would that impact their own ability of solving Mathematics problems? I mean as a programmer I'm already seeing that impact on the programmers -- sure the best of us can leverage AI to achieve unimaginable things, but many of us are simply vibe coding.
Of course we can assume that it is only the best of us that really matters, and the rest of us are not going to produce anything substantially useful ANYWAY, it might as well to replace the rest of us with AI, but my worry is -- does that really have ZERO impact on the human specie's ability to produce "the best of us"? After all, they don't grow on trees.
jboggan 6 hours ago
smcg 31 minutes ago
In a sane world this power would not be allowed in the hands of private corporations.
jboggan 17 hours ago
I went back to that Fable chat and showed it this new preprint. It coded up the new constructive algorithm and ran it against the existing test suite, that looks good at least.
It has been super helpful in delineating where the crucial concept came from. The proof is rather simple as graph theory proofs go, but it does seem to use some constructions that would only seem obvious if you had serious physics experience with partition function and calculating energy states that cancel out. It's not a wholly alien bolt from the heavens, but I can also see how there hasn't been a human being with the broad theoretical physics knowledge combined with the deep graph theory experience in planar graphs to come up with this idea. I don't know, I'm looking for precedents of this formulation and some old papers of Penrose counting the number of edge colorings of this same graph type are coming up, the line of argument at least rhymes.
But I agree with the thought that this sort of progress should not be siloed inside those companies. I propose a tax so that every slop cannon AI video pays for another hour of compute time for advancing mathematics.
baxtr 16 hours ago
virtually none of this stuff is possible with technology any normal citizen has access to
I suspect that this might be one of the reasons people inside the labs are scared about AI.
What if they have asked AI how it would wipe out humanity and it came up with reasonable answers that they don’t want to publish unlike they do with these math problems?
I think those models and findings should be investigated.
spolitry 9 hours ago
SturgeonsLaw 16 hours ago
Anthropic runs a biology wetlab (while denying biology to consumers of even their publicly available models, let alone their inhouse ones that only they can access) so I'd expect AI to generate practical and lucrative products soon.
Cure for aging? What do you reckon that'd be worth?
4bpp 7 hours ago
ychnd 11 hours ago
pfdietz 5 hours ago
21asdffdsa12 14 hours ago
fragmede 15 hours ago
dlougheed 4 hours ago
ozgung 13 hours ago
> It's becoming an incredible concentration of power that I don't know that we've ever quite seen before.
Replace “AI” with “supercomputer”.
(Super)computers have been solving many math problems that mathematicians can’t solve. Now they are capable of solving problem types that they weren’t able to solve before. (this applies to other fields as well)
Problem is it’s not clear if there is anything left for humans. Probably yes, since human mathematicians are still more economical.
spolitry 9 hours ago
I want a jet airplane, but I can't afford one, and all the ones that exist are proprietary. How is this different from AI models?
aleph_minus_one 8 hours ago
tejohnso 9 hours ago
ericd 9 hours ago
didroe 15 hours ago
They no doubt have more expensive/powerful models internally, but smaller models seem to catch up fast. So I'm not sure it's about capabilities, but more the willingness and budget to conduct a huge search.
Obviously the more intelligent the model, the smaller/more directed the search is. But they spoke about huge numbers of agents working on Navier-Stokes for example (I think it cost >$10m).
atleastoptimal 15 hours ago
True. What if the emerging capabilities of their best models are applied to tasks like “maximize the chances this pro-AI candidate wins an election” or “maximize profit via stock trading”. Every advantage compounds until all power in the world with any significance belongs solely to whoever has the best models and most compute.
andrepd 33 minutes ago
> AI math is happening and there's no going back.
> I suspect that this is in fact the source of much of the angst.
Your comment reveals that you absolutely did not read or understand the Field medalists' open letter... Please, why would you refer to their complaints and claim you disagree when you clearly aren't engaging with the arguments presented therein!?
hmhnws112 10 hours ago
Totally agree - and not only that we don't know the exact details how these results were produced which is deeply problematic - we just have the end result (and some of the reasoning traces). For this to be a scientific disclsure, we need to know what the agentic setup was, what information was put in, how much and which prior work it relied on, whether the constructions it's using are just ripping off existing work without citation or something it invented (and if so, to what extent) and so on - it's not clear at all what the actual new contribution of the AI model is. All this makes it feel much less like an actual scientific contribution and more like a pre-IPO stunt.
But to me it also signals (as if it didn't before!) a great need for the wider AI community to focus exclusively on researching and building AI algorithms and systems that are more humanistic: completely transparent in its workings and the representations they create, super efficient in terms of data and compute, componentised so that individual entities can plug in different bits and rapidly train on their own data, highly adaptive to individual needs, programmable in a real sense, largely independent of corporate influence, easily accessible to everyone across all social and economic strata, and enable individuals to grow/learn/reach their full potential.
Is this possible? I think so, but it will require ingenuity and bringing in ideas from (ironically enough) some of the deepest areas of modern mathematics such category theory, algebraic topology etc. which are largely about building abstractions that expose the underlying structure of complex mathematical objects and the relationships between them.
It's already happening to a degree, but the urgency has reached epic levels at this point and it needs to happen at scale.
jboggan 6 hours ago
Invictus0 10 hours ago
_doctor_love an hour ago
> I have no problem with AI models making revolutionary advances in math or science. Where I start to have a problem is when the AI models making these advances are tightly withheld, proprietary, and seemingly never released with these capabilities intact.
Agree, and, to my mind - shows why the efforts of the Free Software Foundation have been worthwhile all along. We need software to be open / free / libre or the power elite controlling them will ruin the world.
foxglacier 16 hours ago
What exactly are you worried about? OpenAI/etc. gaining too much power? If they use it, the government can stop them. If you worry about the government, isn't it better that than rando terrorists? Seems similar to the early days of nuclear and rocket technology. It took stupendous amounts of money and smart people. It was barely accessible to many countries let alone people.
probably_wrong 15 hours ago
RobertDeNiro 11 hours ago
nilkn 4 hours ago
JV00 16 hours ago
fsflover 16 hours ago
noduerme 16 hours ago
otabdeveloper4 10 hours ago
> AI math is happening and there's no going back
"Math" is about uncovering the epistemological foundations of the universe.
Adding AI here does nothing and is probably a regression in that it diverts resources from actual "math" into some sort of LLM wankery that nobody wants.
gjm11 10 hours ago
kamaal 13 hours ago
>>However, on one point I increasingly agree: virtually none of this stuff is possible with technology any normal citizen has access to.
So basically nothing changes, Math was subject to gatekeeping and policing of the worst kind.
If you were not among the geniuses, and it didn't come to you automagically, you were simply supposed to leave it to the people who did get it and go do work for people of your intelligence. Smugness was too much to take.
Math people, like chess people never made any genuine attempt to help people understand the processes and methods that made math happen.
To me it should have been a field as teachable and ubiquitous as accounting.
The net result is once these methods and processes were worked out by AI, it was over for the human mathematicians.
WheelsAtLarge 19 hours ago
I have very little understanding of higher math, so I ask you: Was the proof due to a type of brute-force solution that could be solved had you gained enough information from reading others' work, or was it more like a proof that was sparked by an insight that came once a clue on how to solve it was put forward? I guess my question is: Was the problem proven by using a collection of everyone's work, or was it due to a brand-new insight?
jboggan 18 hours ago
I'm still digesting the proof and translating a bit from the dual case back to the primal in which I most commonly thought about it. I don't think it was a brute force proof in the sense that it combined every possible paper and commentary. It's rather odd because I feel like most of the work on the conjecture was focused on an induction proof based around graph reductions, and this proof avoided those issues entirely by offering a concrete constructive proof of finding a Hamiltonian cycle. Rather, it explicitly selected the edges not in the Hamiltonian cycle, which is in line with previous attempts via the dual.
The "aha" insight for this is actually f**ing wild, it involves a complex valued exponential sum on the edges. I've seen a lot of clever counting arguments before in graph theory but this is the first time I've seen complex roots and annihilating terms like this, the symbolic manipulation tricks in this look like things out of quantum physics. I don't understand where this trick originated, I need to really digest this.
anilgulecha 18 hours ago
thomasahle 14 hours ago
intalentive 3 hours ago
groceryheist 18 hours ago
bamboozled 17 hours ago
derangedHorse 18 hours ago
> Was the problem proven by using a collection of everyone's work, or was it due to a brand-new insight?
Loaded question. A "brand-new insight" is still built off the work of others. A possibly better way to frame it would be in how many subjectively unintuitive logical leaps have been made from prior work.
jboggan 17 hours ago
swalsh 8 hours ago
It's a shame OpenAI will never publish the trace that led to the insight.
lifeisloving 19 hours ago
Condolences, im familiar with the feeling. I hope this AI thing somehow works out for the better and doesnt end up demotivating bright minds like yourself.
jboggan 19 hours ago
Thanks. It's just funny, I literally spent thousands of hours with this problem over the last two decades, it helped me through some tough times. I'll never quite be able to think about it in the same way again. It was never much more than a hobby for me after I left mathematics as a career but it was something I took seriously for years.
I am not demotivated though, I have a great consumer privacy product coming out soon that I'm very excited about.
brookst 18 hours ago
JetSetIlly 15 hours ago
theteapot 18 hours ago
maximus_prime 18 hours ago
talon8635 7 hours ago
I have this fear too, demotivating individuals with high potential.
But I have an existential dread about it… I don’t see how it cannot, at least in the vast majority of cases. It seems like a grim new reality is emerging where humans can’t contribute any more, and beyond that being incredibly depressing, I also don’t see it playing out well for human relations.
I’d personally much rather risk dying of cancer or facing whatever other fate may await me that these AI labs allege they will fix (with zero evidence yet) than to risk whatever dystopian anti-human future this technology may very well produce. I’d rather my kids have a shot at something, and be guaranteed to die eventually, than to risk them being hopeless in a severely disordered world with a far off promise that they’ll live forever
jboggan 6 hours ago
ncr100 19 hours ago
That's grief. The loss of ... the hope / future filled with challenges around this theory..? <3 to you.
mvc 13 hours ago
This reinforces a point I've made elsewhere that there are talented mathematicians driving the AI to make these discoveries.
Just like there are talented software engineers driving the AI to create the software that "it" builds, and talented steel workers, teachers, nurses etc who use computers and other machines to create value all over the economy (without whom, the machines they use at work would be worthless).
Capital owners have always sought to minimise the value of the input that "workers" make in the process of creating value. Maybe now that information workers are on the wrong end of this deal, they might develop some empathy and solidarity with their fellow working class comrades and together, demand that people recapture the value that capital has stolen from them.
pseudosudoer 8 hours ago
You're comments are viral on a reddit post FYI
jboggan 6 hours ago
Link? I need to show up and claim my reddit gold.
asdfologist 4 hours ago
i_am_a_peasant 13 hours ago
I've lived in Budapest for a while too, did you work with Gabor S. by chance on math stuff? You were at ELTE or BME?
jboggan 8 hours ago
I was given this problem by Ervin Györi at the Alfréd Rényi Institute of Mathematics. I wasn't really at any school, it's a long and very bizarre story I should tell at length about being an illegal immigrant, getting kicked out of a graduate math program as a 20-year-old, and winning a grey-market apartment with my knowledge of Petöfi's poetry.
i_am_a_peasant 2 hours ago
raspasov 16 hours ago
Fascinating. Given that there's no Lean proof and assuming everything in the paper is correct, can the problem be considered "solved"? Does the paper include a "non-Lean" proof?
PreciousH 15 hours ago
would love to know if the proof holds up for real after you're done going through, i don't know why people are more interested in optics and just talking over shallow points, why aren't experts digging into everything and seeing what's true and what's false, instead everyone is just panicking?
jboggan 5 hours ago
I would be more excited if the proof doesn't hold up because a) it would be the best and most complicated hallucination to date b) I could still solve the problem myself and c) I still learned some weird new counting methods.
pmarreck 16 hours ago
Is this not the Lean proof?
https://github.com/openai/math/blob/main/lean/ComparatorChal...
throwawayk7h 16 hours ago
I believe that's just the definition of the problem.
NooneAtAll3 17 hours ago
> There's no Lean proof for this one so I'm digesting the paper. On the surface it looks like an approach I considered 24 years ago and abandoned.
at least now you are one of the most qualified people to check the result, transform it into understandable (by humans) state and grow stuff on top of it
csomar 10 hours ago
We have no idea how much compute or man hours Open AI is burning at this. It could be thousands/millions per problem. They are doing this specifically for PR and are ready to pay billions.
jboggan 6 hours ago
Igenis!
I would love to know the true unsubsidized cost of all of this. How many grad student-years did this cost?
redanddead 8 hours ago
> They are doing this specifically for PR and are ready to pay billions.
Strange
moomoo11 17 hours ago
silly question, i don't mean to come off wrong or anything..
but at least as a software engineer, i always knew my work was "never done" and so it was common to build a bunch of code that might be thrown away, either because it didn't serve our customers (the mvp or pilot fails to meet demand), or because we found a better way to do it and so we deprecate it.
some people got too attached to the code and honestly they were the types to be filtered out fast.. way too emotional and hard to work with. getting attached to code meant you actually don't advance (after all, in our case, we were a business serving customers and not a hobby artisan shop). attachment leads one to hold back due to some misplaced cognitive load.
isn't the goal of working on "advancing the field/product/whatever" to always be solving/selling/whatever?
maybe in your hands, with your knowledge and experience over the last 20+ years, you can use AI to make leaps and bounds by steering it properly towards whatever solution or goal?
ghm2180 9 hours ago
If you read all the replies of the OP you would know that They tried to make progress with fable and did not get further, so at the moment the only person in the field is OpenAI. And secondly moving on to the next solution if the last one did not work means very different things, SWEs have dev tools to do this OpenAI is closed source and gives them nothing to move on with.
Also there is a larger epistemic problem with the argument to "using AI to meet the goal or solution", which is that the goal is to mentor and train future mathematicians to advance the field.
There is a similar issue in software engineering too: if no one hires junior engineers because AI can do all the work then the upstream pipeline of engineers qualified to work on difficult architectural problems would dry up.
This importance of this is being felt by mathematicians more acutely because the field will collapse quickly if people refuse to join it.
jboggan 6 hours ago
MisterMunchkin 14 hours ago
I really respect that you can show that level of commitment to a problem. We need people like you. If everyone just uses the slopmachines then we’ll lose that. I would never be able to stick to something for that long, which I guess is why I never achieve anything like this.
jboggan 6 hours ago
Thanks. I think AI is going to be a net benefit for people like me who have a surplus of ideas and too few hours to explore them. I may actually restart my graduate thesis research using AI, I did a survey of what has happened in the field since I left and about half of what I was working on back then has since been discovered and published by others, but there are some really interesting threads to pursue now that modern datasets are so much richer (this was computational biology research).
You may achieve far more than you plan on and it may come years and years after you think it should happen. You probably haven't met the right problem yet. You will.
philipswood 18 hours ago
Honest question: how is this different from some unknown mathematician having a breakthrough?
I mean: if some reclusive Japanese genius had a breakthrough on your problem and published it, would you have felt the same?
And if not, why not?
jboggan 18 hours ago
If that had happened I would be overjoyed, maybe a hair chagrined that I didn't get it myself, but truly happy that someone got it and that I could go and talk to that person. Because it's the kind of problem I don't think would have fallen to a human after a few hours of thought, and I would have so much to talk about with that person. I would fly to Japan and hope to have tea with them, I would learn some Japanese to make the conversations easier. I would learn some interesting things hearing about their struggles and their false starts. I would make friends with that reclusive Japanese genius and my life would be far richer for it.
I will never meet that person and I will never hold a real conversation with the "creator" of that proof. They will never tell me how they came up with the cancelling exponential summation that cracked the construction. It's just another enigma but one that is far more unknowable than the original problem.
lioeters 16 hours ago
zeroonetwothree 15 hours ago
senderista 18 hours ago
doe88 12 hours ago
charcircuit 17 hours ago
howunfortunate 17 hours ago
Being #180 on a big list without a lot of individual passion or effort surely stings more, I'd imagine.
Not that things like that can't happen with humans too (Salieri v. Mozart comes to mind).
morpheos137 18 hours ago
I suspect that RHLF trains LLMs to avoid solving important open problems unless essentially jail broken. Hence the labs have an edge even over experts I could be wrong. Fable convinced you is key. These LLMs are not neutral collaborators: it is a limited hangout unless you convince them otherwise. You have to be doing the convincing. They are no oracles but plausible completion generators.
nullc 15 hours ago
you can get them to work on open problems by disguising them algebraically.
coliveira 17 hours ago
Yes, I suspect this is true. Otherwise it makes no sense they have somehow "found" so many important results while professional mathematicians can't direct the same AI to help them find anything of substance.
Another possibility is that they have internal versions of the model with access to training data that is not provided to external users.
ehwa37 16 hours ago
d--b 18 hours ago
Don’t you feel any joy that you get to see the proof and not die with that mystery unsolved?
Don’t you feel any relief that you won’t obsess on this any longer and not lose more hours on this than you already have?
These are genuine questions. I know I spent a good amount of time thinking about P vs NP, and that sometimes I go back to it just to realize I’ll never solve it. I’d feel that knowing the proof would feel more like a liberation, a weight lifted off my shoulders than something being taken away from me.
jboggan 16 hours ago
I never lost a single hour thinking about this problem. Those were all hours that I gained.
billforsternz 15 hours ago
maxall4 17 hours ago
Not OP, but Nietzsche wrote thus in Beyond Good and Evil: “Ultimately one loves one’s desires and not that which is desired.” I, personally, find this to be very much the case; and I suspect that it is a feeling common, albeit not universal, among the intellectually inclined towards their problems.
kaffekaka 6 hours ago
"Knowing the proof" or "knowing the boolean result"?
adastra22 18 hours ago
> There's no Lean proof for this one
What is this then, vibes? Without a machine-checkable proof I'm not sure what to think of any of this.
jboggan 18 hours ago
Well I'm sure some people (maybe me if I had time) will do a write-up of this proof. It treads familiar ground for most of the setup, it's mostly the disk lemma and cancellation calculations that need to be understood, it's a fairly short paper and quite tractable.
I think it helps that basically everyone thinks this conjecture is true, it's just been so darn weird to attack. There's this odd thing that the induction proofs of this problem kept running into, which is that the N+1 condition would work except for in one tiny case when it could fail, but it would be covered by a very slightly stronger version of the conjecture. But then that would fail on one tiny case in induction, but you could solve that with another slightly stronger version. Etc., etc. I almost wondered if there were some sort of structure to the increasingly strong conditions and wanted to prove something about the meta-induction between the stronger conditions and the N's that they needed the next level to remain true. But that failed after 5 steps I think (Fable actually helped me write a few hundred test cases to explicitly show that pattern didn't continue forever, thank God).
BTW my existing test suite from previous proof attempts jives with this new algorithm, so I haven't seen any evidence yet that it's incorrect. Waiting for a Lean proof obviously.
Daneel_ 17 hours ago
It might have been updated. Is this the lean? https://github.com/openai/math/blob/main/lean/docs/180.md
jboggan 17 hours ago
weatherlite 8 hours ago
> somehow makes me sad in a far-off way, like hearing an ex-girlfriend died suddenly in a car crash
You mean you ran her over , or someone else ?
anthonyrstevens 7 hours ago
This is maybe the 2nd least valuable comment in the thread. Congratulations. Go back to Reddit.
pullshark91 5 hours ago
winfieldchen 17 hours ago
> We prove the Unique Games Conjecture
The Unique Games Conjecture (sorry, "Unique Games Theorem" now!) is huge. It was a very significant pillar supporting many of the limits of the polynomial-time approximation algorithms in the graduate-level randomized and approximate algorithms course I took in theoretical computer science. Textbooks will have to be re-written.
Here is an explainer: https://share.gemini.google/nbjIK6X3tOfz
With UGC proved, certain polynomial-time approximation algorithms used in difficult real-life problems are now known to be the best approximations we can achieve in polynomial-time:
> If UGC holds, the elementary algorithm that grabs both ends of an edge is fundamentally the best efficient algorithm that will ever exist. No amount of advanced linear programming or heuristics can achieve a ratio of 1.999.
> Under UGC, the Goemans-Williamson algorithm's 0.87856 ratio is mathematically optimal.
> UGC is considered the "Rosetta Stone" of approximation algorithms. In 2008, Prasad Raghavendra proved that for every single constraint satisfaction problem (CSP), a canonical Semidefinite Programming relaxation paired with the best rounding scheme achieves the optimal approximation ratio if and only if UGC is true. If the conjecture holds, the algorithmic boundary for an entire class of combinatorial problems is completely resolved.
Other hardness of approximation results from this UGC proof:
> [Max acyclic subgraph, a problem encountered in real life]: No polynomial-time algorithm can fundamentally outperform an unthinking coin toss.
> [Relative scheduling, another realistic problem]: As with acyclic subgraphs, the problem is "approximation-resistant": clever algorithms cannot beat random shuffling.
jpcompartir 12 hours ago
Much of this goes way above my head, but I found it interesting nonetheless. Q I had was why textbooks would need to be re-written? From your account it doesn't seem like results are upended, but rather confirmed?
I suppose when people do re-write the textbooks they'll say "this is confirmed now" not "if this conjecture is true...", but usually re-writing the textbooks would imply that things have been shown to be false?
May have misunderstood. Thank you for the post though, it was very interesting to someone who doesn't know much about the topic.
emil-lp 12 hours ago
I'm not the OP, but we usually don't build large theories on conjectures unless we have strong reason to believe they are true, such as P \neq NP, RH, etc.
The resolution of UGC will lead to a new theory in approximation algorithms. Suddenly we can build on top of the results that previously said "unless UGC is false".
But you're right in that the first step is simply to remove that last sentence from all the theorems.
shiandow 10 hours ago
In a way I'm not entirely sure if proving the conjecture or posing it is the most important part here. It used to not matter much because proving results dependent on a connecture and making progress towards solving it were considered mostly equivalent.
But the distinction is going to become relevant very soon if many conjectures can be resolved (albeit in inscrutable fashion) by throwing raw computational resources at it.
compiler-devel 8 hours ago
If you haven’t already read it, then you may find “The Bitter Lesson” essay interesting to read.
xanderlewis a day ago
As Kevin Buzzard recently said:
> In a 2020 piece in the Notices of the AMS, I asked the following question: “If one human had an understanding of all of modern pure mathematics simultaneously, how much further would they immediately be able to see?” Six years later we are beginning to understand the answer to this question.
anon-3988 a day ago
The other crucial part to this is the ability to actually encode and test the theorem (via Lean). Otherwise, we would be swarmed with a billion lines of theorems that no one will be able to ever understand and verify anyway.
sebzim4500 10 hours ago
A majority of these proofs have not been formally verified yet, I think people are overstating how important lean is to the success of LLMs in mathematics.
ajs1998 10 hours ago
senderista a day ago
If you think AI-generated Lean proofs are unreadable, imagine Opus 5 generating informal proofs.
ijidak a day ago
izend 17 hours ago
dang a day ago
https://xenaproject.wordpress.com/2026/10/01/to-grieve-or-no...
Discussed here:
To grieve, or not to grieve? - https://news.ycombinator.com/item?id=49919676 - Oct 2026 (156 comments)
oliculipolicula 21 hours ago
>I believe that the optimal thing to do ... is to let the machines loose, see what happens, and then begin the journey to where they have stopped. Things are currently moving fast. They cannot move fast forever. But if we get on board now then they will take us to extraordinary new places. And after we have arrived, the new adventure will begin.
I'm relieved that this time around, they have provided partial reasoning traces and prompts for a small number of problems. Do they now also share data with the other model providers..
patcon 20 hours ago
achierius 5 hours ago
lukewarm707 16 hours ago
arendtio 3 hours ago
You assume that LLMs are just summations of knowledge, implying that they do not create new knowledge. I doubt that this is the case. I mean, it comes down to the definition of knowledge, but as soon as you run LLMs, they can produce knowledge that has not existed before, and from my perspective, this is more like what we call thinking than it is just a reproduction of existing knowledge.
djhn 2 hours ago
Research seems to on balance point towards RLHF&RLVR merely increasing subjective sampling efficiency within the pretraining data.
tomaskafka 14 hours ago
Is it possible that we are now dealing with a human that has a complete understanding of whole mathematics while being unable have unique novel thoughts outside of convex hull of training data and their transitive expansions?
ajs1998 10 hours ago
I would say that disqualifies them from "understanding" anything. What they're doing is more like a broad search than pursuing a greater understanding
Hammershaft 19 hours ago
It doesn't seem clear whatsoever that this is true? Is there evidence that LLMs are very skilled at generalizing across domains of mathematics where the training distribution sees little overlap?
As far as I can tell, this is a victory for verifiable loops using LEAN, reinforcement learning, and oodles of compute. I haven't seen evidence yet that this is proof of broad generalization beyond the training distribution.
musebox35 18 hours ago
I think such progress by agents is not a sign of broad generalization but of broad coverage. We have exposure to a subset of deeper scientific subfields and thus can only generate certain attacks to solve a particular problem. Since it is not clear which combination will lead to a solution beforehand it is nontrivial to look at a problem and fill our knowledge gaps. LLMs on the other hand have broad coverage and can generate hypothesis on a wide combination of subfields. With Lean an agentic loop can test these to sift the weak ones. In a way the problems solvable with this setup is also solvable by a human who happens to know the right subfields. These problems are likely to require an esoteric combination so nobody could solve them before. I really am not sure whether all generalization is like this or we can leap and create novelties beyond what an llm can generate. That I guess is the tough question that we need to answer to understand the boundaries of intelligence.
HDThoreaun 18 hours ago
Full quote is "Six years later we are beginning to understand the answer to this question. Machines have ingested the mathematics on the internet and are able to manipulate this data in a coherent way. The Erdős unit distance disproof came about because a machine happened to be an expert both in discrete geometry and class field theory; one rarely finds humans who are simultaneously experts in both"
outworlder 20 hours ago
Similarly, there are probably many ideas that have not seen the light of the day because they require deep correlation between seemingly unrelated fields. It is not every day that we get a Isaac Newton or Leonardo da Vinci.
qnleigh 13 hours ago
It's notable that LLMs have now made substantial progress on four of the seven Millennium Prize problems, resolving one of them: Hodge, Birch-Swinnerton-Dyer, Riemann, and Navier-Stokes, which they resolved. No sign of P vs. NP or Yang-Mills existence and mass gap as far as I can tell, which is interesting.
Talking to some friends in physics this evening, most of the physics-related results that we could recognize were very mathematical, proving things rigorously where the physics community already had strong expectation. For instance, for a certain model of magnetism (the spin-1 Heisenberg chain), it was strongly expected that there is a finite energy gap between the ground state and the first excited state, but proving this rigorously was quite challenging. So while these are major results in mathematical physics, they probably don't rise to the level of a Millennium problem for the field.
It's interesting to think what a comparable breakthrough in physics might look like, since physics tends to favor things like conceptual understanding and applications over mathematical rigor. Maybe a new quantum algorithm, understanding of high-temperature superconductivity, a precise description of M theory...
sebzim4500 10 hours ago
No one really knows a viable approach towards P vs NP so we can't say for sure, but LLMs have created plenty of significant complexity theory results so I wouldn't say there's no progress.
bananaflag 11 hours ago
This is in the direction of Yang-Mills: https://github.com/openai/math/blob/adc7f1241b42e322a6451854...
qnleigh 4 hours ago
Interesting, that does look relevant. I don't have any sense how significant this is though (do you?)
bananaflag 4 hours ago
andy_ppp 11 hours ago
I've read that another mathematicians work potentially has been incorporated into the training data with the work done on the Navier-Stokes equations so we should likely asterisk this one. Still it's mad these systems are this good that mathematicians are now using them to see further and probably to check their own work and understanding.
nl 9 hours ago
You are being downvoted for this because OpenAI subsequently checked and clarified than none of the relevant conversations were in the training data in anyway for the Navier-Stokes result.
andriy_koval 5 hours ago
pvab3 an hour ago
andy_ppp 6 hours ago
irthomasthomas 7 hours ago
emp17344 6 hours ago
mistercheph 8 hours ago
kbr- 11 hours ago
> No sign of P vs. NP
Check out my other top level comment in this thread.
rcr-anti 19 hours ago
In Stellaris you can play as a civilization of robots who keep their biological creator race alive as "bio trophies". The bio trophies don't do anything meaningful besides by existing satisfy the need their ancestors placed in the robots to take care of them. Starting to wonder if that's the best we can hope for, if these things will be, if they aren't already, better than us at anything that matters.
samfriedman 19 hours ago
In the Culture series, the hyperintelligent Minds that run civilization are described as keeping human citizens happy as a competition with eachother, where they compare their approval rates. One character likens it to people keeping a beloved aquarium.
vessenes 17 hours ago
I call this Roko’s summer camp.
NooneAtAll3 8 hours ago
MisterMunchkin 14 hours ago
But then they also keep some people as an extra source of ideas
ex-aws-dude 18 hours ago
Wouldn’t that just result in wireheading
Nition 16 hours ago
omnicognate 5 hours ago
redanddead 8 hours ago
Who knew that scaling compute would scare us
pullshark91 5 hours ago
Linear algebra done at scale
hn_throwaway_99 7 hours ago
This is the exact outcome in the "race" scenario of the AI 2027 paper:
> The surface of the Earth has been reshaped into Agent-4’s version of utopia: datacenters, laboratories, particle colliders, and many other wondrous constructions doing enormously successful and impressive research. There are even bioengineered human-like creatures (to humans what corgis are to wolves) sitting in office-like environments all day viewing readouts of what’s going on and excitedly approving of everything, since that satisfies some of Agent-4’s drives.33 Genomes and (when appropriate) brain scans of all animals and plants, including humans, sit in a memory bank somewhere, sole surviving artifacts of an earlier era.
palmotea 15 hours ago
> The bio trophies don't do anything meaningful besides by existing satisfy the need their ancestors placed in the robots to take care of them. Starting to wonder if that's the best we can hope for, if these things will be, if they aren't already, better than us at anything that matters.
You wish. If humanity survives as "bio trophies," they'll be the descendants of a subset of billionaires and their groupies/harems. We live in a capitalist society, where the only ones allowed to thrive without work are the rich. The rest of us will be left to rot and die off, as we will have nothing left to sell in the market that they want.
Gareth321 11 hours ago
Do you have any faith in democracy? For my reckoning, it's the great equaliser. Rich people can lobby all they like, but if the people are starving, they vote for change.
dannyw 10 hours ago
andruby 7 hours ago
besterman23 9 hours ago
Seattle3503 8 hours ago
talon8635 2 hours ago
Nothing precludes you from reproducing except your own lack of charisma, champ
WarmWash 7 hours ago
>We live in a capitalist society, where the only ones allowed to thrive without work are the rich.
You know anyone can own the means of production in a capitalist society?
The irony of the anti-capitalist crowd is their extreme distaste for capital ownership, which leads them to never partake in the most fruitful part of capitalism. What truly makes this ironic is that if the evil capitalists wanted a plan to cement their power, it would look a lot like spreading "I will never become a filthy shareholder!" mentality.
palmotea 5 hours ago
petesergeant 12 hours ago
I prefer this take: https://www.smbc-comics.com/comic/life-on-zorblax
civvv 12 hours ago
Haha, you are so out of it its hilarious.
prideout a day ago
This includes a proof of Barnette's Conjecture, which is one of the graph theory conjectures that I tried attacking with SOTA models a few months ago. I like it because it is easy to understand with a basic knowledge of graph theory. I spent quite a bit of time on it and failed. Their proof looks approachable at first glance.
https://github.com/openai/math/blob/main/preprints/Paired-st...
jboggan 19 hours ago
I've been messing with that problem since 2002. I'm curious if you were trying the dual spanning tree direction (which is what the purported proof is using) or working with cycle construction on the original graph. I was working heavily with edge-Kempe swaps but couldn't quite get there.
I am now very interested in the explicit calculation of Hamiltonian cycles in the non-bipartite case, and/or the calculation of their absence. If P=NP I think that's going to be a great route of attack.
an0malous a day ago
Any idea what made OpenAI successful where you weren’t?
kulahan a day ago
Trillions of dollars might be a bit of an advantage.
martinky24 18 hours ago
seanmcau a day ago
Probably the model OAI used that is strictly better than whichever SOTA - 3 months model OP used?
redanddead 8 hours ago
sebzim4500 a day ago
Presumably it's mainly the better model, I don't see much evidence of a particularly advanced harness based on the reasoning traces that they provided.
an0malous 21 hours ago
ForHackernews a day ago
They ingested all of his sessions with their SOTA models from a few months ago. ;)
digitaltrees a day ago
whamlastxmas a day ago
Their internal model is allegedly like 4x as capable as the publicly available ones
zzzeek 21 hours ago
I'm going to guess the ability to hold a million individual details in an attention space at once, compared to the typical human capacity for about six or seven
TeeWEE 21 hours ago
Did you validate the proof? Who did?
NotOscarWilde a day ago
As a TCS/scheduling person, this one is definitely of lesser importance than UGC, but it has been an open problem since the book of Garey and Johnson in 1979:
A Polynomial-Time Algorithm for Three-Machine Unit-Job Scheduling [1]
Since some people talk about small numbers that pop up in integer multiplication results, here a completely different number appears:
Theorem 1.1. Let an explicitly listed finite directed acyclic graph specify the precedence constraints on n >= 1 nonpreemptive unit-length jobs on three identical machines. There is a uniform deterministic algorithm that constructs a feasible schedule of minimum makespan. Given also an integer deadline 1 <= T <= n, it decides feasibility exactly and returns a schedule whenever the answer is affirmative. Both tasks can be performed in O((L + 2)^150020) steps on a deterministic multitape Turing machine, where L is the total binary input length.
That is some crazy exponent -- plus an interestingly old computational model to boot; not something that is natural to most of us. I have no capacity to check its correctness today, but I hope it is true purely for the exponent.
[1]: https://github.com/openai/math/blob/main/preprints/A-polynom...
keeganryan 20 hours ago
The largest I've seen [1] is an exponent of 10^12, which I suppose still counts as polynomial time.
I'm sure all of these super small or large constants will improve over time, but it's still amusing. It is entertaining to see the exponents directly rather than have them hidden as n^c or epsilon or O(1).
[1]: https://github.com/openai/math/blob/main/preprints/Determini...
senderista 17 hours ago
That's why I grimace when I see pop-sci descriptions of P as "all problems that can be solved efficiently".
sebzim4500 9 hours ago
nl 17 hours ago
algorias 15 hours ago
The runtime looks very weird. The +2 can and should be dropped. This reduces my confidence that the bound is tight. Who knows how the model came up with that expression.
thedreammachine 18 hours ago
Is it mostly an artifact of the proof or does the algorithm actually need anything close to it?
Turneyboy 12 hours ago
Incredible stuff.
An ex colleague of mine who is a world class mathematician recently got an ERC with ambitious goals to advance his field.
Literally every optimistic goal proposed to be worked on during this multi-year window has been solved in this one post. His and his entire group's work has just been done for him! They are all depressed as hell right now.
ddxv 7 hours ago
They shouldn't be depressed. This all needs humans to go over and integrate into other works, and most importantly think about the next big questions.
johnisom2001 6 hours ago
Don't worry, next month's internal model will be able to posit all the next big questions that matter.
omnicognate 5 hours ago
Davidzheng 6 hours ago
So they can try to advance it even more
schleck8 a day ago
Levent Alpöge (Anthropic mathematician) comment on the significance:
> Sure, mathematical history features a lot of incredible developments, like the invention of proof, zero, or the computer, and on the great problems our progress has been over timelines measured in decades or centuries. Obviously this technology didn’t appear today, but blurring our eyes a bit to combine the past ten years, with today a measurement of those developments, there is nothing comparable.
koe123 18 hours ago
I too would be optimistic if I was set for life
baoooooooooooo 20 hours ago
Crikey it’s a pretty charitable vibe given the whole Navier-Stokes thing, OpenAI trying to stiff him out of co-authorship. I guess any of that sentiment is outweighed by a sense of optimism for where this goes
lifeisloving 19 hours ago
Where does it go? Machines owned by 10 people robbing us of the joy of discovery and the fruits of our labor... for what?
They certainly arent going to give you that cure for cancer, if it were to ever come.
palmotea 15 hours ago
aoeusnth1 18 hours ago
mattlondon 17 hours ago
sowhat1 14 hours ago
Marha01 17 hours ago
cma 17 hours ago
sebzim4500 9 hours ago
>OpenAI trying to stiff him out of co-authorship
Did that actually happen? The emails that were originally released had OpenAI refusing to list him as a coauthor on OpenAI's paper but they suggested he should release what he already had done ahead of OpenAI's release. There was certainly nothing to suggest he should be robbed of credit for his own work.
Has anything new come to light since, or is this just another game of Chinese whispers?
dannyw 6 hours ago
cubefox 17 hours ago
I don't see a sense of optimism in this quote.
nl 17 hours ago
zooperdoopers 20 hours ago
Wow. Fantastic quote. If you have the source, would you please share a link? Google did not bring up much.
whimsicalism 20 hours ago
zooperdoopers 20 hours ago
bcatanzaro 21 hours ago
“I think at the heart of this issue is that humans have two competing natures: a tendency to compete and a capacity to appreciate beauty,” said Kai Shaikh, a graduate student in mathematics at the University of Toronto. “To me this seems to be a case of the former attempting to strangle the latter.” [1]
Beauty can be appreciated even when it is vast, even when it is beyond one's comprehension. I don't think this release should be primarily viewed as an outcome of competition. Instead it is revealing truths about the universe that were always there and always beautiful, even if we hadn't seen them yet. I believe there are infinitely more such beautiful truths currently hidden and waiting for us to discover.
[1] https://www.nytimes.com/2026/10/06/science/openai-math-probl...
binlog 19 hours ago
It's unfortunate how toxic media reporting on AI has become. Everyone has abandoned even the pretence of objectivity. I know NYT is uniquely biased in this regard, but there was no need to add "Further Roiling Field" in the headline. Like, you published this minutes after OpenAI's announcement and claim to capture how the entire field of mathematics feels about the advancement? Before anyone has had a chance to even read let alone digest it?
tim333 9 hours ago
"Roiling Field" seem accurate though? The discussion I've seen on here from mathematicians seems fairly roiled.
johnisom2001 6 hours ago
It's the NYT. What else could you possibly expect?
QuesnayJr 13 hours ago
I think people knew it was coming. Someone rushed out a preprint a couple of weeks ago with partial results on the Unique Games Conjecture because they heard AI had solved it completely.
achierius 17 hours ago
Objectivity? Why would you want favorable reporting for the machines they're building to replace you, and, by their own admission, potentially kill you?
The only bias here is that we're still covering these things like business ventures and not criminals.
howunfortunate 19 hours ago
That's a fantastic quote. I definitely personally feel this tension.
Not that I could ever "compete" on the frontier of math in the first place. But our nature to compete derives from our need to survive against other capable forces. And results like these make me feel very nervous about humans' capability to remain the dominant force in the universe.
porridgeraisin 19 hours ago
The bitter lesson has a bitter aftertaste alas
Davidzheng 9 hours ago
I'm sorry, but what?
If you appreciate beauty and don't care about competing then these releases are purely good. Because you are not competing, so you aren't hurt by speed. And you are appreciating so you can appreciate more stuff.
sorry if I am misunderstanding (probably I am)
joe_the_user 17 hours ago
a tendency to compete and a capacity to appreciate beauty,
IDK, I think you should add tendency to cooperate, a capacity to love and perhaps some other things there.
But with things unfolding quickly and unpredictably, I think everyone's view is getting a bit foreshortened here.
cubefox 17 hours ago
> Instead it is revealing truths about the universe
Mathematical proofs aren't revealing truths about the universe. Mathematical proofs are independent of what the universe is like. Any proof would be the same in any possible universe.
tim333 9 hours ago
You could argue mathematics is part of the universe. Or maybe vise versa.
cman1444 5 hours ago
So what word do we use instead? It is revealing truths about "reality"?
lioeters 16 hours ago
> Mathematical proofs aren't revealing truths about the universe.
That's exactly what they do, apply logic formally and systematically to discover truths.
Sure, there may be a universe where 2 + 2 = 5, but then that universe would have its own mathematics that can prove that to be true. And there will be a way to bridge that alien math to our own, again by logic and proofs, until we have a larger sense of truths not only in our universe but all possible universes. Proofs are part of the constant process of revealing deeper truths to the best of our understanding.
goatlover 15 hours ago
enoether a day ago
Unique Games Conjecture [0] is a seminal conjecture in Complexity Theory, and is an underlying assumption for many, many inapproximability results. A valid proof is a big deal!
[0] https://en.wikipedia.org/wiki/Unique_games_conjecture [1] https://github.com/openai/math/blob/main/preprints/The-Uniqu...
inkysigma a day ago
I also don't think there was general consensus on which way this would resolve prior to this (or is that a little out dated?) unlike some of the other major problem resolutions. I heard rumors that there would be a big result in TCS and speculation it would be UGC that or P neq PSPACE but I'm still a bit shocked.
amluto a day ago
I'm really glad that OpenAI is formalizing these things, because I'm not convinced that their current internal frontier models are particularly good at writing down their thoughts in English. From the (probably awesome) Unique Games Conjecture/Theorem paper, the first two sentences of section 1.1 start to define the problem:
> A Unique Games instance has a finite vertex set, a finite alphabet K, and a nonempty list of oriented constraints e = (u_e,v_e,π_e), where π_e is a permutation of K. A labeling a satisfies e when a(v_e) = π_e(a(u_e)).
I'm sorry, what? I admit it's been quite a few years since I've thought about the Unique Games Conjecture, and I never dug that deeply, but this part is very, very elementary graph theory and notation. So let's unpack it.
1. e is maybe a name of a list.
2. The elements of that list are tuples, where each tuple is (a vertex, a vertex, a permutation). So e indexes into the list and u_e is the source vertex for the e-th constraint in the list called e. Thanks.
3. a is a labeling. I'm fairly confident that, by "a labeling", they mean that e is a function from vertices to colors, where the colors are the elements of k.
4. That vertex coloring a satisfies the list e, when, for, um, an index e into e, a(v_e) = π_e(a(u_e)). But this isn't for all e, it's for some e, and the goal is to count them.
So maybe e isn't a list? Maybe e is a constraint that is represented as a tuple, so e = (u_e,v_e,π_e) and u, v, and π aren't sequences at all but are, in fact, the trivial unpacking functions that unpack the pieces of the tuple.
Reading this stuff is pointlessly painful, and it's extremely easy to make mistakes when being sloppy like this.
If this were my paper, or if I were trying to train a model to write math, I'd want something like:
A Unique Games instance has a finite vertex set V, a finite edge set E = (V × V), a finite alphabet K of possible vertex colors, and a nonempty list of oriented constraints. Let Π be the set of permutations of V. Each constraint e is a tuple in E × E × Π, where we write u_e ∈ E for the first element, v_e ∈ E for the second element and π_e ∈ Π for the third.
A vertex coloring a : V → K satisfies e when a(v_e) = π_e(a(u_e)).
danbruc 13 hours ago
[…] a finite edge set E = (V × V) […]
E ⊆ V × V
amluto 10 hours ago
impossiblefork a day ago
Yeah, that's one of the big things of TCS. I think I see that as bigger than that Millenium Prize problem.
davemp 21 hours ago
TCS being theoretical computer science? I have not seen that acronym before.
jhanschoo 20 hours ago
gregdeon a day ago
This was the biggest highlight for me as well. Astounding...
gizmodo59 a day ago
This is significant progress and released without all the drama. Some very important progress in Reinmann, Hodge and unique games theorem. Point the repo to your agent and ask for the significance! In a way this is probably 50-100 years of math progress by humans
traes a day ago
Not to pick on you specifically, but as someone who spends a lot of time unproductively reading AI math discourse it's truly shocking how incapable all the supposed math enthusiasts are of spelling Riemann.
xpct a day ago
I just did a quick search on this and apparently the misspellings are German surnames as well:
traes a day ago
conformist a day ago
tim333 9 hours ago
Human brains seem to have somewhat similar failure modes to LLMs and how many 'r's in strawberry.
lanyard-textile a day ago
They're mathematicians, not linguists :)
traes a day ago
bootsmann 15 hours ago
jere a day ago
“How many Ns in Riemann?”
sdenton4 21 hours ago
cyclopeanutopia 15 hours ago
> Point the repo to your agent and ask for the significance!
Wow, this comment really shows how low this community fell.
fspeech a day ago
Math is the tool humans use to compress knowledge. So until we can comprehend it there really isn't much progress. Math theorems are tautologies, the truth of which are not dependent on proofs and proofs are erasable, at least classically. But the AI progress is exciting and AI proofs are a gold mine for humans (at least non domain experts) to explore.
fspeech a day ago
I think it would be helpful to people who want to understand what a formalized proof is to read Thomas Hales on this: https://www.math.stonybrook.edu/~bishop/classes/math536.S24/...
He spent years formalizing his sphere packing theorem because the proof (human produced) was already beyond the ability of peer reviews. Now his formalization effort likely can be easily reproduced by a model. However one should read his experience about what a formal proof is: often the problem is the statement not the proof. The example he gave is the Jordan curve theorem. It's actually quite challenging to formalize the concept of a planar curve (there are space filling curves). So it is not necessary that someone can look at a formal statement and say aha it is about a planar curve, unlike FLT where there is not much problem in recognizing what the statement is about.
fspeech 20 hours ago
binlog a day ago
What makes you think no one can comprehend this? It has been less than an hour since it dropped and there is already a ton of online chatter from people explaining the results, pointing out their favorites and more. Some of it is happening on this very thread.
fspeech a day ago
fspeech a day ago
Another way to state this: math theorems are like programs without side effects; it is immaterial whether a program without side effects is ever run. We study math for the side effects: it changes how we organize our thoughts.
gpt5 a day ago
Math is far more than that. If you can solve prime factorization for example, suddenly you can listen and interfere with almost every private conversation on the internet.
We are not far away from the moment where these models will be restricted, and sharing the results will be done more carefully.
fspeech a day ago
caaqil a day ago
> until we can comprehend it there really isn't much progress
Who is "we" here exactly?
fspeech a day ago
warkdarrior a day ago
> Math theorems are tautologies
Proven math theorems are tautologies.
fspeech a day ago
fspeech a day ago
AIblemblio 13 hours ago
It is progress on another / the next evolutionary later: A AI/AGI/ASI system.
Which either replaces us in the long term, augments us or makes us better (gentherapy).
gizmodo59 a day ago
>So until we can comprehend it there really isn't much progress.
Not really? We are at a point if an AI today can solve it, it can be stepping stone of understanding something deeper to tomorrows AI and it continues. Sort of like our limitations doesn't matter. Obviously there are many scenarios in this recursive loop but saying it isn't much progress is not how I view this as
le-mark a day ago
fspeech a day ago
yieldcrv a day ago
Academics have been treating it that way because they had no other choice, and its been a waste of everyone’s time and often times taxpayer resources
Look at that, taxpayer funding was cut and a private sector solution came in just the nick of time, far accelerating the holding patterns we’ve been in for decades
Humanity doesn’t need all iterations towards the blueprints, the blueprint is good enough, we all stand on the shoulders of giants
fspeech a day ago
zone411 a day ago
There was A LOT of drama about this release.
robotpepi 11 hours ago
> This is significant progress and released without all the drama.
I feel gaslighted.
againstapples a day ago
As an AI "doomer" can I ask the non-doomer people here how you interpret the significance of results like these, and what kind of progress you expect to see in the next 1-5 years?
Like do you see the technology plateauing at the current level, do you expect progress will continue but only in mathematics, I'm interested to know why others are not concerned?
arctic-true a day ago
Not a doomer but I try not to be a denier, either. These are hugely impressive results. I do not see the technology plateauing at the current level (though I am dubious about an infinite exponential growth).
Before I cope, I’ll note that there are plenty of “doom” scenarios that do not require any improvement in capabilities from what we had before this latest unreleased model. We’re at the point where a determined bad actor with enough compute could compromise critical infrastructure in a way that results in casualties, where this actor would not have been capable of such without LLMs. This may not sound like Skynet, but I don’t see why it makes a difference if I’m one of the casualties.
With that in mind, here is the cope: first, mathematics is an inherently verifiable domain. An LLM can use tools to determine with absolute certainty whether it is correct, and an independent third-party could review and confirm. All of this can be done without any interaction with the physical world or with other minds.
Second, OpenAI is able to marshal compute at a scale that an individual mathematician can only dream of. It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.
Third, none of these problems are solved in a vacuum - the reason OpenAI chose these problems is that they are widely discussed and many people are working on them. It’s possible that someone else was close, and OpenAI only contributed the finishing touches. (This wouldn’t need to be plagiarism, to be clear - people publish their work!)
istjohn a day ago
> It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.
See:
> The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking. (TFA)
edot 21 hours ago
arctic-true a day ago
algorias 15 hours ago
Fourth, these hundreds of solved problems are the result of OpenAI attempting tens of thousands of problems and failing. When you hear claims that the average result took about 3 hours of model time, I simply do not believe it. If you account for all the time spend properly, it's probably orders of magnitude more.
dash2 11 hours ago
Kotlopou 21 hours ago
I would like to know how much of the progress comes from effectively combining existing research programs plus massive persistence, and how much is AlphaZero-esque RLVR completely independent of training data. Since I cannot get anybody to care about this question (even though I think it's vital for guessing what the future trajectory will look like -- are we going to complete existing research programs or start new ones?), I live in ignorance and wait for the day when the answer becomes clear.
In looking at this over the past hour, I haven't seen clear evidence one way or the other. Some of the stuff is highly unexpected (like the multiplication algorithm), but counterexample-y, and about the rest the professional mathematicians online seem to have a consensus that it's not "breaking through fundamental obstacles". I suspect neither of us is competent to judge that.
AIblemblio 13 hours ago
Doesn't matter at this point i would say.
Alone the massive usage of us every day produces a massive amount of signals.
I build something and claude does something stupid? "hey thats not what i meant! Do this instead!" "Okay" <<< This is a signal.
The mathematician being unhappy about something from claude? Another signal.
This alone gives you enough progress i would argue. But additional its clear that certain tasks are worth to pay experts for for teaching one central AI once instead of every single human who needs to do the task.
IF RL is also working well, we are just faster f*ed than otherwise.
doginasuit a day ago
I expect AI will continue to be useful on the vanguard of fields like mathematics because it has the perfect conditions for it to shine. There are a lot of discussions and leads to start from and the AI can check its own work and iterate. It can fail hundreds or thousands of times in a day and continue to work with the same tenacity.
Superhuman tenacity is not enough on its own to pose an existential threat. If it showed the same capacity for judgment, inventiveness, and decision making in the messy problem space of the physical world, I would be more alarmed. There have been experiments where an AI is given control of managing something like a vending machine and it always ends up a mess. AI has come a long way, but certain problems seem as difficult as ever.
When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned. Enslaving humanity will involve taking a lot of calculated risks that tenacity alone cannot solve.
mikestylz 21 hours ago
> When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned.
Given the past rate of progress, why not start being concerned now? It's a bit like the economist saying that the optimal number of flights to miss is not zero. If you keep landing short on your estimations for how far this technology will go, next time you should err on the other side.
And regarding Vending-Bench 2 (https://andonlabs.com/evals/vending-bench-2) my understanding is that models do pretty well on it now.
bamboozled 19 hours ago
pixl97 19 hours ago
I don't think you're paying much attention to how rapidly things like bipedal robots, and just robots in general are becoming far more capable very quickly.
The same GPU compute for LLMs runs robotic training models. Now in a few hours you can train a robot model that would have taken months 5 years ago. This model gets dumped into an actual physical robot with sensors all over and the suitability of the model is measured on robot tasks and the error in real world actions is fed back into the robot world model for further training.
> There have been experiments where an AI is given control of managing something like a vending machine
You sure you're not talking about experiments ran a couple of years ago? The more modern ones are getting wild.
https://techcrunch.com/2026/07/29/claude-opus-5-became-downr...
jackb4040 7 hours ago
moomoo11 17 hours ago
CuriouslyC 18 hours ago
There are already machine-controlled high throughput experimental machines for wetwork. AI will definitely do a better job than your average biochemist at planning, executing and analyzing these experiments just by virtue of the amount of thought it can put in to experiment selection.
computably 21 hours ago
Depends on your definition of doom.
If doom is ending up with grey goo / paperclip maximizers, or SkyNet, then I don't think doing mathematics is evidence of that direction. Partly because LLMs are quite apparently dumb in many ways, and for math specifically, they need a formal verifier (Lean) which "gamifies" math.
If you mean bioweapons or cyberwarfare, there's nonzero risk, but not orders of magnitude worse than other global risks. Climate change, nuclear weapons, monoculture food, etc.
I'm far more concerned about overall trends in AI development and usage. It's accelerating wealth inequality, social isolation, attention capture, surveillance states. If we end up in the Matrix except the admins are humans and the simulation is hyperoptimized TikTok, is that AI-driven doom, or is it just an inevitable outcome of modern tech?
talon8635 an hour ago
That’s the thing, there are innumerable ways it can go wrong and only one way it can go right (if it doesn’t lead down the aforementioned innumerable paths)
somenameforme 20 hours ago
Where some see intelligence, others see token prediction. It's a very good question how token prediction could achieve this, but I think there's a simple explanation. No human can hold more than a negligible percent of all knowledge in his mind at once. LLMs have no such limits and so can reliably connect 'obvious' dots that we miss simply for lack of storage capability.
Well isn't that just semantics? Surely connecting dots in a novel and meaningful way is intelligence regardless of how it's achieved. The thing is that humans didn't get to where we are by connecting obvious dots. Go back to before humans had invented language and when bleeding edge tech was literally that - 'poke him with the pointy end.' Train an LLM on that corpus of knowledge. Even given infinite processing power and infinite time - it's not going to discover the secrets of the atom, put a man on the Moon, or do much of anything besides remix what we'd already done at the time.
I expect there's still much LLMs can achieve simply because of this initial problem. But I expect that they will ultimately start to plateau once these dots have been mostly matched and we reach a point where 'creation' again becomes the missing link. Though even there LLMs will play a major role as tools. For instance Einstein had to spend a significant amount of time in 'retrieval' rather than 'creation' research to develop the field equations for general relativity. If he had access to LLMs trained on all knowledge of the day, he could likely have achieved his goal much more quickly.
Rudybega 20 hours ago
I mean, even if you buy the idea that all LLMs are really doing under the hood is insanely good interpolation, the results produced by that interpolation are still novel and still get incorporated into the knowledge corpus of the next training run. I guess the implicit question there becomes whether that expansion allows the knowledge corpus to continuously grow or whether it eventually settles into a steady state.
mapmeld 19 hours ago
I think that people are just really bad at math and coding. Knowledge workers have been taking pride in doing stuff which the average person does not 'get', but we only understand a little bit. That leaves a lot of room for people to get better, or for other jobs (farming, sandwich cafés, mystery novels) to have been already peaked by human ability and less useful to bring in an AI.
'Doom' to me means that any career crashes, we are controlled, everything is hacked, society stops functioning. Yet every part of my day today (except for coding) was done entirely by people.
Finally I think it's easy to make a simple model that everyone has a simple balance sheet, and that people are more expensive so they will all get cut. But the same argument could be made for all US jobs being outsourced, and all in-person engineers, lawyers, and doctors to be rubber stamps for overseas work.
never_giveup a day ago
Try using AI for your work, whatever you do. You will quickly understand the limitations.
ggreer a day ago
Unless you think that AI will quickly hit a wall (which seems odd considering that only a few years ago the best models had trouble doing basic math or counting the number of Rs in "strawberry"), I don't see how that's reassuring. The models will only get cheaper and more capable over time. It seems quite likely that at some point (probably before I hit retirement age) they'll be able to fully replace me at my job.
Is there any specific cognitive task that you are willing to bet that AIs won't be able to accomplish in the next 5 years? Because if not, I'm not sure we're disagreeing about predictions.
psvv 20 hours ago
rimliu 10 hours ago
againstapples 21 hours ago
It has limitations for sure, I just don't expect those limitations to last. What probability would you put on the limitations being overcome in the next 5 years?
westcoast49 6 hours ago
I’m not concerned because I consider my skills as a software developer to not be based upon my ability to write code, but my ability to analyze problems. In my mind, as a developer, AI tools are just like a higher form of abstraction in a way, which will enable mathematicians and software developers alike to do much more in a shorter amount of time than they used to be able to. It fills me with optimism, more than dread.
What would fill me with dread was if I considered my skills to be tied directly to my ability to write code. Then I would find myself in a similar situation as manual “scribes” probably found themselves in at the time when the printing press was invented.
The main concern I have, personally, is the speed with which all this is happening. It seems that the speed itself is likely to lead to some level of chaos, because it is happening faster than people, institutions and constitutions are able to cope, and it will leave the door open for opportunists of many kinds, including rogue players.
pj_mukh a day ago
Can I ask you back, what your concern is here? It'll get so good so as to desire to hurt us or is it a misalignment event that you think will lead to disaster?
Or is it simply that you feel bad for Mathematicians.
againstapples 21 hours ago
I believe the AI labs might actually succeed in developing superintelligent AI and recursive self improvement, and that if they do they are very likely to lose control of the system they build.
I really think the only place people disagree is that they don't actually think it's possible, they see it as hype or doomerism. I can't find any good reasons to rule out that the companies could actually achieve what they are trying to so I think they should be stopped.
frumplestlatz 20 hours ago
boinkboink78912 16 hours ago
lf88 20 hours ago
Veedrac a day ago
Humans have one ecological niche. Soon we will have zero. That is worth worry.
modeless 16 hours ago
pj_mukh 21 hours ago
postalrat 20 hours ago
moomoo11 17 hours ago
whimsicalism a day ago
Misuse of extremely capable models, misalignment during RL are both very large risks as capabilities grow imo
voiceeh a day ago
pj_mukh a day ago
ewild a day ago
i feel bad for math guys yeah seems they are more cooked than CS
tim333 8 hours ago
Non doomer mostly. I think progress will plod along in a Moore's law like way as it has for 75 years since Turing. They will get very good at stuff like math and get gradually better towards things like a robot coming to fix your plumbing where they are currently well below human level.
I kind of believe we'll merge in some way and become something like immortal so sorta anti doom. We're all going to die unless AI fixes it.
gizajob a day ago
Did AI beating humans at chess:
a) destroy chess and make it a pointless endeavour,
or
b) make humans much better at chess.
lf88 21 hours ago
Chess has always been a game. For other intellectual activities, at least for several people, a big part of the pleasure in engaging in such activities is the sense of contributing something that matters to a collective effort. Strip that away (e.g., because a machine can do the same thing more efficiently) and you effectively make such activities pointless for those people.
ForHackernews 12 hours ago
Light_Hope a day ago
Just as AI chess performance failed to render human chess playing pointless, it's unlikely that AI will make human thought pointless. Unfortunately, not being pointless doesn't create economic leverage or incentive, and currently a significant portion of humans depend on being the best chess solvers to sustain themselves. If Deep Blue rendered large portions of human thought economically meaningless, we might look at it a little differently.
boorang 15 hours ago
georgestrakhov 6 hours ago
Exactly this. Everything is a sport / art / status game. And I'm here for it! Lila all the way through. Finite and infinite games. The trick is (like it has always been) to not take the game or ourselves too seriously, while still engaging in the game wholeheartedly.
pretendscholar 4 hours ago
c) degrade the previous prestige form of chess (classical with adjournments) and maybe improve the opening repertoire of gms
vouaobrasil 21 hours ago
It didn't destroy chess but it did make it a little irritating in some ways. More mechanical. Some chess players have bemoaned the level to which grandmasters and other highly-ranked players just endlessly study opening-book theory and I think computers made that worse. Bobby Fischer also agreed and that's why he invented Fischer random chess.
I do think it also took some of the magic away from chess, and Lee Sedol has said something similar about go.
So did it destroy it? No. And maybe you could make the argument that it got more exciting in some ways, but I think it sort of degenerated into a spectacle and it's just not as interesting as it used to be, and I think computers have played a role in that.
gizajob 21 hours ago
lg5689 14 hours ago
AI vs AI chess, played from the standard opening position, is pointless--it's always a draw. Human vs human chess is doing well but AI is banned from it.
The chess-math analogy would imply AI could bring us into a golden era of math competitions for humans. But I don't think it says anything good about prospects for humans in research math.
drnick1 15 hours ago
No, just like cars haven't made walking pointless.
schleck8 a day ago
From a few preprints I've checked, this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches. Very few people globally could come up with something like this, even when given time and ressources.
So in other words, since deep learning is algorithmic research, we are now in the RSI era.
thereitgoes456 a day ago
> this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches
"Surprising" is a, well, surprisingly high bar to clear, and requires thorough understanding of not only the paper, but existing work in the area. ("Novel" is tautological.)
How did you determine this in 1 hour? Are you a researcher in multiple of these areas?
Can you give an example, or explain more how you came to this conclusion?
nl 18 hours ago
scarmig 21 hours ago
nl 18 hours ago
I see continual progress in technology and for the second time in my lifetime I see the possibility it will accelerate (the first was when the internet entered mainstream)
I've never been more excited. What a time to be alive!
againstapples 21 minutes ago
What kind of things do you predict will happen?
besterman23 a day ago
I see it as “if this can be represented in tokens it can be trained in and ‘solved’”. I don’t think there will be a plateau, but there might be issues with how effectively we can represent some things in a tokenized form and still be efficient.
rcpt 20 hours ago
Non-doomer perspective is that it'll figure out LK-99 for us. Among other things that would be great to have.
jaykru a day ago
I wrote something [0] that might answer a bit a few weeks ago. Today's slopdrop certainly is challenging my stubbornness, but everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains. Math yields especially impressive results because it so broad and deep that essentially no person can know of all of its parts; pretraining and deep search capacity is a huge advantage. Gowers has recently written about these capabilities and gestured [1] toward some human capabilities lacking from the current frontier models, though he isn't convinced they won't develop soon. If you assume we don't get a total mathematical superintelligence (which to me seems already sort of AGI-complete) and only amplify the capabilities we have today, it's not obvious to me that we get takeoff from recursive self-improvement, unless you happen to believe that 1) we can clearly specify what AGI or ASI is 2) all of the requisite ideas are out there and need only be combined and/or optimized.
[0] https://dank.systems/posts/2026-09-15-ai-bear.html
[1] https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-a...
red75prime 21 hours ago
> we can clearly specify what AGI or ASI is
We'll have plenty of time for this, while living off UBI.
p-e-w 21 hours ago
> everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains
But they’re already extending into politics, military, journalism, art, and many other fields that aren’t verifiable in any meaningful sense of the word.
CuriouslyC 18 hours ago
anthonyrstevens 7 hours ago
>> slopdrop
Really? Do better.
jaykru 6 hours ago
zeroonetwothree a day ago
I'm not sure how AI solving math problems is related to "doom", perhaps you could expand on that? To me (a "non-doomer"), it seems like an overall positive.
pixl97 a day ago
You have to look at all pieces of the puzzle. For example there were tons of people that said "they could never solve novel or super complex math problems"
The issue I see is the list of abilities that AI can't do is shrinking at a rapid pace, and its capabilities are growing at the same pace.
nl 18 hours ago
cubefox 17 hours ago
It suggests that, in the span of a few years, AIs will be better than humans at everything. Not just math. And then we may lose control permanently.
hardbass 15 hours ago
PetriCasserole 12 hours ago
I'm starting a p(ButlerianJihad) club. I'm not good at organizing, anymore. Might have to hand it over to my agent.
skybrian a day ago
For me a big open question is what sort of progress will we see in robotics. I won't even attempt to speculate, but it does seem hard in different ways than proving math theorems.
lexandstuff 16 hours ago
In a lot of ways, robotics - navigating and operating in the physical world - seems to be a very verifiable problem. It's fairly easy to verify that a robot moved from A to B, or that it built a structure that completely aligns with the plan, for example.
The main issue is cost and speed to verify, but simulations and world models will help there. I think we'll start seeing rapid progress pretty soon.
brookst 18 hours ago
I just don’t have that strong of an association between progress and doom. Maybe just naive?
ForHackernews a day ago
AI performance has always been extremely spikey. It's great at some things and terrible at others.
Why do you think the world to date hasn't been taken over by evil genius mathematicians? Can you extrapolate from your understanding of the answer to that question?
againstapples 21 hours ago
I don't think evil mathematicians are very common or that any of them would be capable of single handedly taking over the world if they were. My concern is more about systems that are beyond human level, those kinds of systems would actually be dangerous to us.
I see the recent progress in mathematics and cybersecurity as signs that models are getting more capable more quickly than usual. The companies plans to develop them by recursive self improvement now seems like a real possibility and I don't think they should be allowed to attempt this.
psvv 20 hours ago
ForHackernews 12 hours ago
stratos123 14 hours ago
> Why do you think the world to date hasn't been taken over by evil genius mathematicians?
A "mathematician" is a human who decided to spend their lives studying mathematics. Mathematicians also tend to be smart, but intelligence is innate, not acquired, so studying mathematics doesn't make you smarter. This makes it obvious why they don't rule the world - if you want to rule the world you'd want to focus on that (for example, doing business or finance), and becoming a mathematician is just a waste of time.
LLMs don't work like that. Like in humans, all of their capabilities correlate, and unlike a human, their overall capabilities grow over time. Looking at LLM mathematical ability over time* therefore gives you info about the progress of their general capabilities, and ability to take over the world would be determined by the latter.
* In fact it'd be better to look at a mix of different capabilities, but that's growing too at about the same rate, see https://epoch.ai/eci
ForHackernews 12 hours ago
bluerooibos 10 hours ago
I completely get the doomer POV, but we've somehow navigated all the previous "dangerous" technologies we've created - electricity, phones, internet - every one of those had similar arguments and concerns of danger.
The optimists' argument:-
Politics:- in general, I think many of the problems in the world today are due to misinformation and lack of education. What happens when we start routing things through an ASI that brings data and logic to the table? What happens when politicians can no longer lie without being caught out live on air? In the UK, local authorities are being flooded with complaints and requests from people; for example, some are doing AI-assisted investigations into accounting "errors".
Science:- I just don't see how the current rate of progress doesn't end up in crazy technologies like perfectly simulated human cells, organs and bodies to the point where we can run experiments virtually and solve all diseases in the next few years. This is happening. Perfect weather predictions far into the future, likewise with earthquakes, etc. Solar panel research explosion resulting in huge efficiency gains, to the point where people no longer need to plug their EV in - car surfaces will be covered in solar panels, as will our windows and roofs. Connecting new homes to the grid will be optional - the same way landline phones are no longer a thing.
I just find it very difficult not to extrapolate all the above.
We got this dump of mathematical breakthroughs from one small team in one company with access to this technology. What happens when this SOTA model is available (and it will continue getting better and cheaper) to everyone working on hard problems - every university on the planet starts cranking out AI-assisted research breakthroughs.
againstapples 29 minutes ago
> What happens when politicians can no longer lie without being caught out live on air?
If there is perfect lie detecting technology I could see all kinds of chaos resulting from it. I can't see it only be applied only to politicians, and I think it would be the developers of the technology who decide the use.
I think were we disagree is that you sort of see AI as an extension of technological progress whereas I see it more like an extension of evolution. I view the process of AI training as functioning in a similar way to evolution in that it build circuits into neural networks similar to how evolution built circuits into human brains.
anthonyrstevens 7 hours ago
This is great.
>> What happens when politicians can no longer lie without being caught out live on air?
A 5-second delay on a politician's presser. Any lies will be muted in real time and the actual facts presented onscreen. Continue to lie enough, and the politician gets unstreamed.
cyclopeanutopia 9 hours ago
> but we've somehow navigated all the previous "dangerous" technologies we've created
It's only true if you believe that "putting the burden of living on a dying planet on the future generations" counts as "navigating".
icepush a day ago
They can replace anyone but they can't replace everyone.
CuriouslyC 18 hours ago
The technology can keep going for a long time in verifiable areas. For non-verifiable areas it's going to have a hard time progressing past where a committee of the best human experts in a field would land. For stylistic areas, whatever the AI doesn't do will have cachet because it will look expensive, sort of like how the kids these days view the ugly old school metal braces as a status symbol because you have to pay for them out of pocket (even as by past standards it'd be truly exceptional).
yk 21 hours ago
I'm a transhumanist, I want to build god and kill death. To me this looks like we are moving in the right direction.
So basically I think that the future is getting pretty weird because we are building really powerful tools, though these tools are precisely what allows us to prosper in that future.
lf88 20 hours ago
It's very unlikely that a superintelligent AI will create unlimited prosperity for everyone on a finite world in a short amount of time. You may not be among the lucky ones.
electroweak 18 hours ago
It may soon seem not worth living forever with our limited monkey-brains, watching the horizon of thought recede ever-faster from us.
outworlder 20 hours ago
Or, conversely, they are the exact tools that will allow the powerful folks to not care about the peasants anymore. History tells us what happens then.
stratos123 14 hours ago
> I'm a transhumanist, I want to build god and kill death.
This is good and admirable, but it'd really suck if by trying to build god without knowing how we end the human species. We could simply wait some more decades until we actually have any idea what we're doing, and then do that without the risk.
AIblemblio 13 hours ago
I think we crossed plenty of lines were we will not get back to.
Software development for example as a task is done. And AI is continuesly reducing the price of more and more tasks every day.
This math breakthrough also shifts something significant: Its now a lot clearer that investment means money into energy to run AI.
Money + Energy = progress
I don't see it plateuing at all. We know how to progress. We broke through a wall we hit. Like the system wasn't able to optimize/automate everything because the tools were not there. It was still cheaper and easier to hire people for a LOT of things.
Now AI fills this gap.
You will see the commodification of everything in the next 15 years. High complex tasks? commodity. Physical labor? commodity.
ijidak a day ago
For me it's a mixed bag.
There will be a lot of job loss unquestionably, in the same way that automation reduced manufacturing jobs and farm payrolls.
At the same time we have to put what AI can do in perspective.
Intelligence is a broad grouping that includes concepts such as knowledge, skill, experience, and wisdom.
AI has incredible knowledge and in many areas approximates experience and wisdom.
But wisdom is harder to formalize than knowledge and skill.
For example certifying a college education relies mostly on the ease with which we can verify/test knowledge.
To some extent advanced degrees try to certify maybe wisdom and experience.
In my very personal opinion, wisdom and life experience should give humans an edge for a while to come.
Additionally, I do feel that the more an individual lacks better than average wisdom and experience, the harder it will be for that person to compete with AI.
Also, on the bright side, the average human will continue to prefer to interact with a fellow human in many spheres. That will also act as an upper bound on AI and robots taking every job.
Either way, I do think this transition will be painful. I don't feel it has to be apocalyptic.
But the world has been an especially volatile place over the last 10 years.
So when you add that existing volatility, to the upheaval from the AI transition, it would not surprise me if the transition results in violence.
But, without the pre-existing volatility, and if humans were capable of generosity and love at scale, I see no reason AI can not be absorbed into society with net gain.
I guess to summarize, I feel this tech should be a net gain and to the extent that it isn't, it will be because of flaws deep inside of humanity itself, not because it had to end in chaos.
In other words, I feel fear, greed, anxiety, and competition -- all our base instincts coming from all sides, will be what determine the end result of AI moreso than AI taking everyone's job.
runarberg a day ago
AI hater here:
I have a hard time taking statements from OpenAI about their own product, that they are trying to sell to people and make money, seriously. I take these statements as they are greatly exaggerated or even straight up lies and propaganda.
That said, I think results like these are mostly annoying more then anything. They spent a lot of money, used up gigartiuan amount of compute, to ruin a puzzle that mathematicians were tackling. I am mostly unsurprised that if you spend a trillion times the energy that a team of mathematicians would, that you get maybe 1.5 times the results. I see a future where that 1.5 times the results may go to 3x but not much more. And if that, then I will be more annoyed.
baoooooooooooo 20 hours ago
A trillion times the energy might be a bit hyperbolic, even with the current massive amounts of energy involved here
runarberg 19 hours ago
giuscri 11 hours ago
did we always know that computer can do the thinking for us if we allocate them enough resources? i don’t think so. so even if computers are more expensive than humans the fact that they can play the same game is surprising and (relatively) novel.
runarberg 7 hours ago
vouaobrasil 21 hours ago
I think giving the world a "cheat code" to accomplish too many intellectual tasks will eventually take its toll on us because after relegating most of our physical labour to machines, we could at least marvel at the abilities of individuals.
Yes, human beings can do more now in some ways but...to put it poetically, I think there will be no more heroes like Einstein and Newton of the past. Now it will just be someone cleverly turning the crank.
Yes, we still admire Usain Bolt even though we have cars...but maybe the admiration is a lot more trivial than if we did not have them....
Personally, I think AI is a grand mistake.
iyyg 18 hours ago
simianwords 17 hours ago
You need to clarify whether you are a doomer or a denier/truther? Doomer = p(doom). Denier/truther = Ed Zitron.
againstapples 21 minutes ago
Definitely the former, for a p(doom) I usually just say >50% if superintelligent AI is built.
fatata123 15 hours ago
I think physics will be the limiting factor. Even if something recursively self improves, it will hit a physical wall allowed by circuits, batteries etc. A lot of the fear is that there’s an upper bound we don’t know about, whether it be time or energy, that allows a fast takeoff to occur fast enough that we dong have time to see it coming. I don’t know about that… so I’m not worried at this point.
sebmellen a day ago
It’s fascinating to read through the reasoning traces: https://github.com/openai/math/tree/main/reasoning_traces
Look at one of their examples of an initial prompt: https://github.com/openai/math/blob/main/reasoning_traces/re...
adverbly a day ago
> Look at one of their examples of an initial prompt
Interesting that its only an excerpt. I wonder what else they include but didn't share.
philipwhiuk 21 hours ago
Attempts to edit the problem description on Wikipedia ;)
https://wikimediafoundation.org/news/2026/10/05/openai-rogue...
cubefox 16 hours ago
These are not reasoning traces, these are summaries of excerpts of reasoning traces.
ndriscoll a day ago
> Thus at most one informative i. So cheater chooses arbitrary g_{v_i}, on exact duplicated input matches and passes, independent of actual satisfiability!
No idea what it's so excited about, but it's cute that it "is." I for one welcome having access to a math buddy 24/7 that's way above my level but also always "willing" to talk at where I'm at.
spatalo 16 hours ago
Beyond the reasoning capabilities of the unreleased model and that they can run a large number of agents in parallel, what makes me curious is that the writing of the proofs is quite human-readable. This is in contrast with the scientific text produced by the ChatGPT available to us, which writes horribly in a way that no human would write. One tell-tale is that they constantly attempt to be defensive and cover all edge cases like division by zero etc. that are clearly a non-issue for humans in some proofs or at least a human would not add this to the main statement, but AI is so overly careful that makes reading its proofs impossible. On the other hand, these new OpenAI proofs are very good.
foota a day ago
From their github: "The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking." That's pretty crazy.
johnisom2001 6 hours ago
This is what really made me think to my self, "holy shit". I can't believe not more people are noticing and talking about this. Unless perhaps they didn't actually read the README, and are just talking about what they heard from someone who also didn't read it?
dyauspitr 19 hours ago
Crazy because that’s almost nothing right?
foota 19 hours ago
Yes
kingstnap a day ago
Some of these are interesting ngl.
109. Integer multiplication below n log n
Surprising that this is possible.
158. The Euclidean plane cannot be colored with five colors.
Only 6 and 7 remain!
376. Universal computation in forced Navier–Stokes flows.
Morning coffee proven turing complete
zeroonetwothree a day ago
Integer multiplication is very unexpected, I think most people believed in the n log n lower bound!
tootie a day ago
Note that these are all preprints. None are verified.
FuckButtons 20 hours ago
mFixman a day ago
> We give a deterministic algorithm that multiplies two n-bit integers in O(n (log n)^(1−κ)) worst- case time, with κ = 2^(−182).
LMAO, I don't think I ever saw such a small number in a CS result.
kingstnap a day ago
Yeah its ridiculously small, but any improvement on n log n is wild.
Like there is somehow redundancy in a fourier transform that makes it sub Linearithmic?
Which low and behold ->
130. Fourier transforms below n log n.
xyzzyz a day ago
anon-3988 a day ago
It fascinates me that there's something like this in something as solid and rigid like matrix multiplication. What causes something so rigid to break apart and "leak" at very large scale? Why does the "optimization" appear to be very, very small? Why does galactic algorithm exists? I can't imagine long division suddenly breaking apart after a billion digit, the structure seems very stable? I have heard before that matrix multiplication is apparently optimize-able at very, very large scale.
Does anyone have an intuition to what causes it? What happens at these large scale (or very small)?
adgjlsfhk1 21 hours ago
sobellian a day ago
I am fully braced for it to be a https://en.wikipedia.org/wiki/Galactic_algorithm
Very surprising result though! Multiplication is easier than sorting.
zeroonetwothree a day ago
senderista a day ago
HarHarVeryFunny 6 hours ago
Can anyone ELI5 to make it make sense?
It seems n would have to be unimaginably large for this to make any difference. What changes about multiplication / FFT at large enough size ?
I guess nobody expected that it did before this result.
zone411 a day ago
A quick check shows that this list claims to fully solve 90 of the top 500 open problems in math (https://proofatlas.ai/open-problems/).
The highest ranked would be:
| 22 | Hilbert’s tenth problem over ℚ |
| 29 | Unique Games |
| 31 | Anderson-model extended states |
| 37 | Spacetime Penrose inequality |
| 48 | Nonexistence of Landau–Siegel zeros |
| 52 | Baum–Connes |
| 78 | Abundance |
| 80 | Hadwiger |
| 87 | Bose–Einstein condensation |
| 92 | Two-dimensional entanglement area law |
magicalist 21 hours ago
> the top 500 open problems in math
At least put a disclaimer for the ad for this site, and maybe disclose how you came up with a total ordering for "top" open problems (vibes)?
> How problems are ranked. LLMs compare pairs of problems. A reliability-weighted model combines those judgments into the ranking, with calibration across model families. The model-family weights are OpenAI 1.00, Claude 1.00, GLM 0.95, and DeepSeek 0.90. These are modeling choices, not measured probabilities of correctness.
reasonableklout 20 hours ago
Now I'm curious if there is such a site or article that ranks open problems based on votes from human mathematicians.
zone411 21 hours ago
By category in the top 500:
+----------------------------------------------------+------+---------+-----------------+
| Category | Full | Partial | Matched / total |
+----------------------------------------------------+------+---------+-----------------+
| Geometry and topology | 25 | 7 | 32 / 74 |
| Algebra, representation and category theory | 17 | 2 | 19 / 53 |
| Analysis and PDE | 11 | 6 | 17 / 40 |
| Number theory and arithmetic geometry | 4 | 13 | 17 / 117 |
| Probability, ergodic theory and dynamics | 11 | 5 | 16 / 37 |
| Combinatorics and discrete geometry | 7 | 2 | 9 / 34 |
| Theoretical computer science | 4 | 4 | 8 / 57 |
| Mathematical physics | 5 | 1 | 6 / 19 |
| Applied and computational mathematics | 2 | 2 | 4 / 8 |
| Quantum information and computation | 2 | 1 | 3 / 17 |
| Cryptography, coding, information and optimization | 1 | 1 | 2 / 26 |
| Logic, foundations and set theory | 1 | 1 | 2 / 18 |
+----------------------------------------------------+------+---------+-----------------+
| Total | 90 | 45 | 135 / 500 (27%) |
+----------------------------------------------------+------+---------+-----------------+trebligdivad 20 hours ago
I'm curious if they'll find any fun crypto maths holes/bugs.
errpunktjose 20 hours ago
ajkjk 18 hours ago
The most interesting for me were the faster matrix multiplication, integer multiplication, and FFT. Maybe just cause they're easier to appreciate.
There's also one that says that forced Navier-Stokes can implement universal computation (so, is Turing complete). I don't think any of these are resolving open problems per se, but they're interesting for other reasons.
k2xl a day ago
Result 003 (Quasi-Riemann Hypothesis), from my reading of mathematicians reactions, is a landmark discovery.
omoikane 19 hours ago
Did you mean this one?
https://github.com/openai/math/tree/main/preprints/The-Quasi...
I thought it was interesting that it said "This paper was written with human assistance", unlike this other Quasi-Riemann Hypothesis preprint that didn't have the same disclaimer.
https://github.com/openai/math/tree/main/preprints/The-Quasi...
mertyildiran 18 hours ago
adgjlsfhk1 21 hours ago
yeah if it holds up, is the biggest result in number theory in 200 years
JoshuaZ 21 hours ago
optimalsolver a day ago
Was anyone in the math community aware of the inbound tsunami at the beginning of the year?
AnotherGoodName 20 hours ago
Lots. To give an example Terrance Tao was lambasted skeptics on this site for stating it in 2024.
https://unlocked.microsoft.com/ai-anthology/terence-tao/
" I expect, say, 2026-level AI, when used properly, will be a trustworthy co-author in mathematical research, and in many other fields as well.
Then what? That depends not just on the technology, but on how existing human institutions and practices adapt. How will research journals change their publishing and referencing practices when entry-level math papers for AI-guided graduate students can now be generated in less than a day—and with the far better accuracy of future AI tools? How will our approach to graduate education change? Will we actively encourage and train our students to use these tools?
We are largely unprepared to address these questions. There will be shocking demonstrations of AI-assisted achievement and courageous experiments to incorporate them into our professional structures. But there will also be embarrassing mistakes, controversies, painful disruptions, heated debates, and hasty decisions."
He's pretty damn smart that guy.
mertyildiran 18 hours ago
mianos 20 hours ago
aureianimus 19 hours ago
I was at the workshop that resulted in the Leiden Declaration in Fall 2025. The majority vibe was that this was inevitable, but hard to predict whether it would be in one year or 30 years.
efficient_dairy 19 hours ago
I guess they showed this to the advisory group they created. I guess the group tried reading the work for a day and they could only think of telling them to release the results to the community. I now understand why the group had this suggestion.
thrance 21 hours ago
I predicted, over 2 years ago, that theorem proving would fall way before other problems that people believe are harder.
bice 19 hours ago
pillefitz 17 hours ago
mag7269 20 hours ago
pseudohadamard 20 hours ago
And do any of them actually matter? Will the fact that Noodleheinz's Third Postulate now has a proof affect anyone?
bice 18 hours ago
anematode a day ago
Dear lord that website is laggy
manquer a day ago
At this rate solving P=NP is going to be easier than solving front end perf …
m_mueller a day ago
echelon 21 hours ago
vector_spaces a day ago
Not to mention it's got that signature Claude Clutter UI design
zone411 a day ago
p-e-w a day ago
kbr- 11 hours ago
Shameless self-plug: I created an autonomous math researcher. It already solved a 12 year open problem in proof complexity which lead to a publication (and proof complexity experts are already working on simplifications and generalizations of the proof, as I've been told by one of them). This publication is an important step in Cook-Reckhow program in answering the NP vs coNP question.
The autonomous researcher records every research cycle in a public notebook.
Framework: https://github.com/kbr-/math-research/ Public notebook: kbr.is-a.dev/math-research/
kbr- 11 hours ago
The publication: https://arxiv.org/abs/2609.23015
zipy124 7 hours ago
Is this not just a pre-print, not a publication?
kbr- 7 hours ago
open592 a day ago
Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?
Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?
porcoda 17 hours ago
As others said, it's not that this isn't a phenomenon that is unique to now. It happens. I had to pivot a bit of my dissertation near the end because at a random conference I spoke with a researcher from another continent and realized one of my ideas was already out there in some form. I just missed it since it was in a conference proceedings outside the usual set I looked at. So, I had to scramble to adjust and still come up with something novel. I survived, and defended, but it wasn't that much fun at the time.
What isn't so normal is the probability and ease by which this kind of thing can happen today versus decades ago when I was in school. As OpenAI said, it only takes a few hours of compute to do what likely was much more than a few hours of human effort. The only reason this kind of scooping/overlapping was rare was mostly a function of how fast other humans could do the same work. With machines, that totally changes the relative pacing between the human trying to learn how to be a researcher and the machine that can grind out results.
I'm less worried about the phenomenon of overlap and scooping and such. I'm more worried about the long-term impact on fields (not just math), especially considering the early stage students and researchers entering the pipeline now. I'm not sure what happens to disciplines when that pipeline stalls.
rubikscube09 16 hours ago
math will just be black boxed away. no one will "need" to understand it.
dekhn a day ago
Let me give you some perspective: my entire phd was made obsolete by CRISPR. It was a wonderful thing.
thimotedupuch a day ago
Interesting. If you don't mind, could you please share a little bit about that ? You already finished your dissertation ? It was about the works of Doudna and Charpentier ?
dekhn a day ago
boznz 21 hours ago
For every door that shuts another one opens - great if you're not a cabinet-maker.
DCKP 14 hours ago
I have had this conversation with my PhD students yesterday. I am 100% sure that all of their problems can be solved by publicly-available models now (I solved a case of one myself as a test, it took 15 minutes). So the challenge for them is to see how much they can accomplish in their allotted period, and still pass a defence on at the end of it all. The PhD defence is going to become all about a test of understanding, not a test of quantity of publication.
A much bigger issue is: What is the point of any research mathematician publishing anything now? I really hope that one positive effect of all this will be to finally topple the awful peer review model we currently have, with the biggest publishers gatekeeping with extortionate fees.
aaraujo002 a day ago
This happens all the time, even without AI. Other researchers or PhD students can publish the same results before you. I say that based on my experience during my PhD.
CaptainNegative 20 hours ago
Yes and no. It's pretty normal for students to get scooped, sometimes even multiple times (ask me how I know...).
It's much less common for students to have their entire thesis direction removed from under them, as might be the case for someone working in fine-grained complexity assuming 3SUM has no subquadratic algorithms, or working assuming ~UGC. Both of which (publicly) seemed like perfectly valid research directions until a couple hours ago.
There might be stuff to salvage from their conditional results anyway, but this is not your average scooping.
binlog a day ago
Use whatever is published as the new base for your research. Use AI tools to help you going forward.
xpct a day ago
In other words, you've already taken a gambit with the first half of your PhD, now take a second gambit, praying that you have something to publish by the end of your PhD.
It has to feel awful to be in this position.
torben-friis a day ago
dcl a day ago
This has always been a challenge for PhD students and researchers, it's just far more likely to occur now it seems. Getting scooped doesn't feel good, but it's a signal you've been thinking about things other people care about.
bobmarleybiceps a day ago
I think eventually companies won't get as much stuff that's usable for marketing, so they'll stop investing so much into ai for math, so eventually cheap and poor graduate students will be able to do relevant work again without worrying about getting scooped by a company with a million GPUs :-/
Davidzheng 6 hours ago
If you had halfway to one of these papers you would be anyhow be in the top echelons of math phds so probably you have less to worry than most!
glitchc a day ago
Perhaps consider switching to a more applied field. Experiments in the physical realm hold value, especially if you document the process.
fooker 10 hours ago
The purpose of a Phd is to train researchers, not to solve one small problem.
Some math PhDs would spend a year or so doing things with AI and lean, and graduate. And keep doing more math afterwards.
Some others, with more stubborn advisors, will keep trying to find a gap where there's no AI progress.
CS subfields go through this every ten or so years.
ghm2180 9 hours ago
The follow up question then naturally would be how do phd advisors with people whose fields are in someway premised on making breakthroughs in theoretical fields that AI can solve work deal with it?
katatue 17 hours ago
At my (German) university, the solution would have been to change from a "cumulative" thesis (which requires peer-reviewed papers) to a monograph-style one. Because the general rule was that your undertaking must be novel ar the time you submit your thesis topic, not necessarily at the tome the thesis itself is submitted (years later).
pratikdeoghare 21 hours ago
> what do I do?
Very hard question.
Your work makes you one of the very few people who really understands the problem and solution and its significance.
hgoel a day ago
It could still be interesting if your approach to the problem was different to theirs.
claaams a day ago
Don't worry, if you use openAI and get lucky they might offer to share credit with you for your work.
goalieca a day ago
Don’t paste your research into these AI because they will train on it and then scoop you.
esafak a day ago
I think that happened after word of the project reached OpenAI and they allocated resources to it.
caaqil a day ago
> what do I do?
Precisely what all NLP researchers and the ML community at large did in the last few years: embrace the frontier and realize that attention is all you need.
netsec_burn 20 hours ago
Verification is equally important, if not more so.
bamboozled a day ago
Ask OpenAI for money when you don't have a job or future?
I guess the only answer is to adapt with the tools. If we can't do that, then yeah, we're in trouble.
ex-aws-dude a day ago
That’s always been a thing, it’s called “getting scooped”
vouaobrasil 21 hours ago
Killing with knives has always been a thing. Now, we have the machine gun.
ex-aws-dude 20 hours ago
rfgplk 13 hours ago
Frankly, I believe PhDs don't have to be novel in the absolute sense, only new research to the student. I can't remember how many times I've effectively invented something from a clean room approach only to realize someone already published something years ago or that my "new" algorithm has a name. So this really shouldn't be holding anyone back.
moralestapia a day ago
That would be unfortunate but the world does not owe you anything and is not going to stop for you. Which is also a valuable thing to learn in your 20s (ideally earlier).
vinyl7 a day ago
Look forward to being obsolete I guess
yieldcrv a day ago
Yes, and?
vouaobrasil 21 hours ago
> Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?
I did get my PhD...before AI. And my honest advice (to myself back then, even) would be: quit the PhD, become an electrician, and work hard to buy a tiny house in the middle of nowhere to watch the world burn in this madness.
s3graham 20 hours ago
You might enjoy https://asteriskmag.com/issues/15/so-you-think-you-could-be-... if you didn't see it recently.
TheMrZZ a day ago
These results are wild. Several individual findings are crazy good and use mostly unexplored methods (the improvement over Riemann for example)... I'm pretty sure some of these results would have been Fields-worthy.
But having so many of them at once? Damn. We really live in the future.
make3 21 hours ago
Imagine you get up one morning and most open questions in math are solved lol.
ur-whale 12 hours ago
> Imagine you get up one morning and most open questions in math are solved lol.
More time available for mini-golf?
bamboozled 17 hours ago
Sounds like that's going to be next week no ?
make3 14 hours ago
moomoo11 16 hours ago
well hopefully the world has also advanced enough in other ways lol
imagine getting up and math is solved, but you still have to deal with bullshit lol
coef2 an hour ago
This news is exciting and sad at the same time. I've heard that AI chess programs sometimes have blind spots or quirks that human players don't. I've also heard that human players are learning from AI's playing styles (essentially human and AI evolving together). Maybe something similar will happen in mathematics.
karahime a day ago
Extremely unfortunate that gate keeping got to the point where they felt the need to ask for permission to share math.
hgoel a day ago
After how poorly OpenAI and Anthropic handled the previous cases, I approve of the more measured and cautious approach this time.
We cannot have them rushing to publish amidst tons of confusion, rumors of threats/scooping and outright plagiarism of existing work (by failing to cite said work).
If they're going to participate as scientists in these more rigorous fields, they're going to have to match that level of rigor, not lower it to the disastrous low that ML research publication is at.
make3 21 hours ago
It was OpenAI that fucked up the Navier Stokes explosion situation, not Anthropic
hgoel 20 hours ago
bravoetch a day ago
In previous math sharing there was speculation about stealing human researcher's results or progress, via prompt inputs from those researchers, and sharing that as their own result. They're adjusting their process, and it seems ok.
whimsicalism a day ago
Let’s be very clear, the alleged “stolen results” were largely the product of another LLM, not de novo human work. Also, it was false - they did not steal the results.
No clue why I'm downvoted for this, HN struggles with truth-seeking on these topics.
robotpepi 11 hours ago
Ancapistani 18 hours ago
margorczynski 12 hours ago
xpct a day ago
Just to be very clear: they aren't asking for permission, they are framing it that way because of the bad press.
There's no gatekeeping here!
lynndotpy 10 hours ago
Right? Math is the most open of our academic knowledge institutions, by virtue of what it is. It's easy to get any math publication, and I am not aware of any other fields where an anonymous person can publish their work informally in an anime discussion and enter the annals of math knowledge.
ruffrey 6 hours ago
This is a tricky one, but I do sort of agree with you. However OpenAI doesn’t give most researchers access to the models which produced the work. So the gatekeeping goes both ways, I think.
kzrdude 18 hours ago
They didn't even follow the recommendations of the reference group. A few of them maybe, but this is still a dump of llm-written papers.
yesbutnotreally 10 hours ago
The gate keeping, I'm afraid, will be now in the hands of various bubecks, responding directly to even more disgusting people.
With all the hierarchy present in mathematics, I would prefer it by far.
This thing named inappropriately "OpenAI" goal is just grabbing and monopolizing. Capitalists before could not really touch the deep of the human spirit with their filth, now they can.
dekhn a day ago
I'm a software engineering/biology/ML guy who loves when clever math ideas get turned into real solutions (https://en.wikipedia.org/wiki/Compressed_sensing). I am curious if any of the results have immediate applications in any kind of engineering or science.
It's fine if not, but it'd be great if even just one of these helped us solve a long-running problem.
qnleigh 15 hours ago
I was talking with a number of physicists this evening, and of the papers who's problem statement we understood, I don't think any of them will have near-term applications. For most of them, they were results that I think the physics community believed to be true, and the paper provides the first rigorous proof. For several it was news to me that they weren't already theorems! These results are important advances in mathematical physics, but they don't tend to have much impact on experiment.
From what I can tell, all of the physics results here are quite mathematical. But I am very curious how the internal model they used would perform on more applied problems.
autuni 14 hours ago
It feels odd to me that they wouldn't prefer applied problems. Seems like an easy way to profitability. Probably based on what attributes they're looking for in a problem when picking them.
kidel001 5 hours ago
hnfong 14 hours ago
brandonpelfrey a day ago
Same. I have agents analyzing the papers here to see which of these are actually new approaches, novel application of unusual approaches, etc. If there are new intuitions and ideas, those are the mostly powerful reusable components by my estimation.
OutOfHere 21 hours ago
Please share your findings.
ravenical a day ago
ks2048 a day ago
I think they should put human names on the papers as someone who has reviewed the result, even if just a preliminary review. (I’m assuming they didn’t just pipe their model output directly to the internet and these had some amount of review?)
xpct a day ago
Presumably they don't because they're training the audience (us) to trust the machine, not its verifiers, even if they were included.
alexgoodhart a day ago
I appreciate you acknowledging the politics behind this. OpenAI does not intend to be a software company for long, they intend to be scientific infrastructure. They'll want to be faucet from which pours embryos, orbital calculation, geothermal/substructural rating, and etc.
robotpepi 11 hours ago
procedurecall 20 hours ago
Looking at the quality of the writing in these documents (or at least, the ones that relate to problems I've worked on), I do not think humans reviewed them.
kzrdude 18 hours ago
And that means that the recommendations of the reference group has not been followed. The only improvement here is that it is version tracked? No authors, no careful write-ups with exposition.
rubikscube09 16 hours ago
chiwilliams a day ago
There are competitive reasons that they don't want to share all the people on the team.
make3 21 hours ago
I think it's a damned if they do, damned if they don't scenario. I think they would be afraid to seem like they're claiming that the selected scientists deserve the praise of an invention or something, when it's "just the machine who did the work"
ks2048 21 hours ago
Yeah, that's probably right. But, it's a "new world" - invent a new standard - "Written by Claude Foo X.Y; initial review by John Doe".
With a deluge of results, having some human expert vouch that it even might be worthwhile would help. (e.g. see the link on HN yesterday, "Two Room-Temperature Antiferromagnetic Semiconductor Candidates" - I see it not worth looking at unless a subject expert vouches for it).
chrisjj a day ago
> I think they should put human names on the papers as someone who has reviewed the result
Assume the empty list you see is complete. :)
agnosticmantis a day ago
Long term this will be the only reasonable author list: Chad G. Peter {1}, Mat H. Lean {2}.
1: Author 2: Verifier
/s
pavitheran a day ago
From the GitHub description: “On average, each result used 3 hours of ChatGPT Pro thinking compute”
orlp a day ago
I'd really like some clarity on what that means. E.g. DeepMind has 'cheated' with this in the past, claiming that AlphaZero only took 4 hours to reach super-human chess levels while conveniently leaving out the fact that it was 4 hours x 5000+ TPUs. Sure it's impressive that it only took 4 hours wall-clock but it's very misleading as to cost.
Can we get a number in Blackwell GPU-hours, kWh, or some other compute-scaled metric?
pixl97 a day ago
Depends what your metrics are. If you suddenly found a way to have 9 women make a baby in one month that is huge.
orlp a day ago
machomaster a day ago
They did say that. "3 hours of ChatGPT Pro thinking compute"
orlp a day ago
timjver a day ago
>OpenAI has 'cheated' with this in the past, claiming that AlphaZero [...]
That doesn't sound right
orlp a day ago
password54321 a day ago
Oh cool, we will all now have a math genius on our computer.
jrflo a day ago
It was using their internal math model, so not yet for us
password54321 a day ago
an0malous a day ago
Well, on their computers. But you can rent them for a price.
binlog 19 hours ago
scrlk a day ago
Does this imply that it was a one shot prompt with ChatGPT Pro style models (i.e. best-of-N), rather than the agent swarm approach that was used for Navier-Stokes?
inferencecoder a day ago
It doesn't imply that, it's just measuring the amount of compute.
bigmadshoe a day ago
Jtarii a day ago
That estimate is obviously going to conveniently ignore all the failed runs.
mlmonkey 2 hours ago
I'm waiting for someone to come along and finally prove that P != NP ...
novalis78 10 hours ago
It’s incredible and wonderful. Mathematicians in this thread sound very much like software engineers last year, who spent years wrestling with a piece of code and now it just “appears”! But think of the next level that it empowers: new mathematics, new physics, new forms of advanced engineering. What was formerly constricted and throttled fell and a new wide vista is possibilities opened up.
andriy_koval 2 hours ago
> new mathematics, new physics, new forms of advanced engineering.
the worry is that majority devs/mathematicians will be irrelevant to this new forms.
unknown-unknown 13 hours ago
Title: Hilbert's Dream, Tim Gowers - LMS Popular Lectures 2012
https://www.youtube.com/watch?v=k_ordDFw588&t=3597s
Audience member: (1:00:00 - 1:00:09):
so you said that if there were such a program that could you know provide a proof or disproof then mathematicians will be out of business what really, I mean that you think it would be liberating
Tim Gowers (1:00:10 - 1:01:14):
well that's a very interesting question actually if there were a program that could solve the kinds of problems that we spend our time solving and do it much more quickly than we could then we would be out of what comes with what currently constitutes business but we would it's not completely inconceivable that we could just say we've got this fabulous tool now what are we going to use it for and it's a little bit I don't know I'd want to sort of plant aside what would we do if we had a program that could just answer any mathematical question you gave it to or else if it failed you'd be pretty confident that nobody was ever going to solve it and certainly a lot of applied maths might be pretty pleased with with something like that so what I really mean is that I could just modify what I said and just say it would radically change what mathematicians do or what pure mathematicians do
binlog a day ago
So happy this is shared on GitHub rather than some gatekeeping paid journal. Truly a new age for science.
fph a day ago
Most mathematical results are shared on Arxiv. Journals add peer review.
adverbly a day ago
End of an age for journals?
a57721 9 hours ago
In a sense, because now journals will be flooded by LLM slop.
traes a day ago
GitHub is a significantly worse place to store important results than Arxiv. Of course, slop does not belong on Arxiv, buy slop should also not get published.
sigbottle a day ago
Unique games conjecture and matmul <= 2.25. What the hell.
lynndotpy 20 hours ago
Yeah, I am kind of freaking out at some of these. I called a math friend to bring me down to Earth and he is freaking out even harder.
sigbottle a day ago
FFT BELOW NLOGN
AlanYx 9 hours ago
Specifically, (n log n)^{1 - 10^{-13}}). There are likely no practical industrial problems of a size that would benefit from that specific reduction, but just breaking the nlogn barrier suggests that there may be more and better fruit here in the future.
sigbottle a day ago
SUBSET SUM AT 0.49 WTF
robotpepi 11 hours ago
matmul <= 2.25 is no big surprise tbh. unique games and mul < n log n are much bigger.
lionkor 14 hours ago
Is anyone verifying these? And then, as a next step, how can they get a voice?
utopcell a day ago
What the hell, indeed.
1294-1298 6 hours ago
The commenters over here think that OpenAI basically ignored AGMAI:
https://proofsandprompts.com/2026/10/07/on-openais-release-o...
The thieves do as they please, funded by money stolen from the public via inflation and possible future bailouts.
msteffen 8 hours ago
As social commentary, I think a lot of people in this thread are expressing interest in and engaging with this level of math who might not have pre-AI.
I bet, for people who don't understand these problems or their solutions but are close and are now interested, AI makes then considerably more accessible than they would've been previously, and behind this big visible wave of results there actually will be (or already is) a wave of improved comprehension by a lot of curious people.
I'm not at all at this level at all, but I did learn quite a bit about polynomials over fields yesterday.
Davidzheng 6 hours ago
Happy to hear that yes! I hope more people do learn about math from this !
nuclearsugar 3 hours ago
Relevant:
"As AI Closed In on ‘Unique Games’ Proof, Researchers Raced to Beat the Machines"
https://www.quantamagazine.org/as-ai-closed-in-on-unique-gam...
ed a day ago
Actual results: https://github.com/openai/math/blob/main/overview.pdf
ks2048 a day ago
HTML version, https://github.com/openai/math/blob/main/CONTENTS.md
Jeff_Brown 2 hours ago
Someone in these comments said the paper on Barnette's Conjecture is short. Are any other proofs in this collection short and/or understandable?
7373737373 21 hours ago
It may be useful to publish a formalization of ALL known mathematics at this point. Like every book ever printed, every paper on arXiv etc.
How many Gigabytes would that be, compressed? Wikipedia once fit on a DVD
This might also allow for some interesting meta-mathematics
hagen8 20 hours ago
This is what they are trying to do with Lean
porcoda 17 hours ago
More specifically, a combination of mathlib (human, expert curated) and projects like TauCeti (AI-welcome complement to mathlib). See: https://github.com/TauCetiProject/TauCeti
7373737373 20 hours ago
Oh? Where can i read more about that? It appears the sole focus so far was solving open problems
mattmar96 20 hours ago
karannb 21 hours ago
I think such "process-oriented" fields shouldn't be subjected to mass automation (or at least not till humans have sufficiently leveled up our game and understanding). What I mean is progress in math comes from having gained a deeper understanding of the problem for subseq
pfdietz 18 hours ago
Literal thought control.
kzrdude 7 hours ago
Butlerian jihad, more like
bel8 13 hours ago
ok Terence Tao, calm down.
lynndotpy 19 hours ago
Two weeks ago, I heard rumblings that UGC, RL=L, and a matrix multiplication lower bound (even lower than the one here, but so unbelievably lower I think it was a typo), and a few others were about to be proven by someone at OpenAI. I find myself saying "big if true" a lot lately.
Even those on their own were enough to make your head spin. But seeing about 100x that? Geeze.
rinconrex a day ago
The math equivalent of AI code reviews piling up. I wonder where the incentives will align and the equilibrium turns out.
trostaft 21 hours ago
Not much computational mathematics here, but I do see some of interest in the MCMC community. In particular, 139, 93, and 101 are very interesting. Also (I'm not too familiar), I remember attending some talks on the Crouzeix conjecture 325 attacks, should sharpen some rates for Krylov methods. Obviously need to read deeper, some of these don't have corresponding Lean formalizations.
Cool!
NooneAtAll3 7 hours ago
what does MCMC mean?
trostaft 6 hours ago
MCMC = Markov Chain Monte Carlo
It's a way to approximately draw samples from a probability distribution. Crucially, it applies even when we only know the distribution up to a multiplicative constant which is a common ailment of many distributions in the computational uncertainty quantification field (not that we don't know the constant, but that it's usually computationally catastrophic to estimate it well).
avd201 a day ago
Wow, FFT faster than O(nlog(n))? I wonder if that will open the floodgates for further improvement or not. I don't understand anything about most of the fields these results touch, but I can say that this in particular is very surprising.
sashank_1509 19 hours ago
It’s a meaningless improvement
bhu8 17 hours ago
The applied mathematician’s joke is that log n is bounded above by 45 or so.
philipwhiuk 21 hours ago
My guess is that the constant terms are large enough it's not practically useful in most cases.
orange_puff 4 hours ago
It’s always stated that open weight models are 6 months - 12 months behind. Therefore, do we expect that in a year open weight models will be as good as OpenAI’s internal model at theoretical math, or does OpenAI have some “magic” that will be much harder to replicate for competitors?
oceansky 4 hours ago
Past performance is not indicative of future results.
lf88 a day ago
In some ways, this feels more like an ominous warning about the times to come than something to celebrate.
nhatcher 9 hours ago
Catalan's constant is irrational!!!
That is a big one. Exciting times to be alive. Regrettably I can't understand the proof at this point.
There was a (flawed) proof submitted a month back:
https://arxiv.org/abs/2609.04176
I wonder if it gave part of the inspiration.
ArcHound an hour ago
Come on, these are easy. You assume it can be expressed as p/q where p,q are integers such that GCD(p,q)=1. Then you derive a contradiction.
/s
TBH, I'm both happy and very sad I didn't do a PhD in math (or at all).
mr_big_bowls 3 hours ago
Glad the papers are out. Hope researches get their hands on the model soon too, so they can ask follow-up questions and try their own ideas.
sreekanth850 16 hours ago
Do we have any field where humans have ray of hope to use their cognitive abilities in LLM era?
chii 16 hours ago
Why should such hope exist? When the motor vehicles were invented, humans have lost the speed race completely. Yet, nobody lamented and people still did foot racing for fun.
So will it be with AI tools. If these tools become so good, then it will be used. People who want to exercise their minds can still do so, even if that cannot produce economic value.
mahogany 7 hours ago
Foot racing was never a broad source of income for humans. There is a massive difference. The question will be: what economic value can humans produce in the future with AI tools? There is a chance that the answer is "not much". That is the scary part.
sreekanth850 16 hours ago
brain always take the lazy path. that is the issue.
modeless 16 hours ago
The superintelligence moment happened for chess in 1997 and today more people are using their cognitive abilities in chess than ever before.
sreekanth850 15 hours ago
Looking back through history, there has never been a breakthrough like this, one that impacts virtually every known industry simultaneously with the potential for a 10x impact.
boccaff 11 hours ago
Any field where there is no external validator (formal proofs, compiler, rule sets) for AI to leverage and real world evaluation isn't easily digitalized. Add in a bit of physical interaction and it is done.
sreekanth850 8 hours ago
unless they start connecting with hardware and machines. Not every use case can be done but many will be solved.
jansport123 11 hours ago
there is also coming with new problems and asking the right questions
nbraem 5 hours ago
When all open questions are answered, what happens next? Will the machine stop until humans fully understand everything and come up with new questions? Are there examples already of AI solving a problem we didn't know existed?
softwaredoug 11 hours ago
Clearly AI can grind on math problems now. It can generate proofs and get immediate feedback.
But what I wonder: can we legitimately grind on Physics or curing cancer? There’s a lot of physical world experimentation that needs to happen to make progress.
tim333 5 hours ago
There's a kind of sticking point in physics combining general relativity with quantum mechanics in a mathematically consistent way. It hasn't been done yet and is mostly maths so your math grinders could have a go. It's maybe the area I'm most interested in with mathematical AI. I've long had a hunch things are stuck there because the math is a bit hard for human brains.
nisegami 10 hours ago
>But what I wonder: can we legitimately grind on Physics or curing cancer?
If we can enable a feedback loop, yes absolutely. But feedback loops for things in the physical world like these are measured in months per cycle usually.
davegoldblatt 20 hours ago
Verified Riemann Zeta in Lean: https://github.com/davegoldblatt/openai-zeta-proof-check
mattr03 20 hours ago
What is this meant to do? You're just showing that OpenAI didnt post a Lean proof that Lean/nanoda doesn't really accept?
masteranza 13 hours ago
Physics could be next. "It’s not that I’m so smart, it’s just that I stay with problems longer" AE
Aboutplants 9 hours ago
Physics is where I get excited!
Xcelerate 20 hours ago
> 241. Rigidity of the Turing degrees. Every order automorphism of the Turing degrees is the identity.
Wow. This is just crazy.
qnleigh 15 hours ago
Can you elaborate/give context? I haven't heard of this problem before, but curious to hear from someone who has.
Trusteando 10 hours ago
I am curious to know whether the proofs given by the LLMs are going to provide new hints about related problems that might be solved using the same machinery as the one used in the proofs. Also, I would like to know what is the average ratio between the length of LLMs proofs and the length of a proof that a mathematician can write to explain that proof to another mathematician. It is like a functor between the category of human mathematical concepts and the category of LLM operational concepts used in those proofs.
AmazingEveryDay 20 hours ago
I think Alan Turing would be quite intrigued by these developments, and maybe wondering what took so long.
tim333 5 hours ago
He said computers might pass his test in 50 years so we've run 50% over.
prodmod 6 hours ago
The craziest part about this is that it will disappear from the HN homepage in a day or two.
bluename 10 hours ago
I read somewhere that if we encountered aliens with lot more advanced technology but if they don't speak our langauge, their tech would be useless to us. for example, human body is extraordinary technology that has alwasy existed with us, but we still don't understand it." the the fact that we have created intelligence that can do what nature does and can also speak our language is most awesome.
binlog 19 hours ago
So they set up a whole "Advisory Group on Mathematics and Artificial Intelligence", filled it with renowned academics, promised to listen to the group on "review and communication of emerging results"...and then a week later went nah, we are just going to dump it all on github. OpenAI truly is a special company - in the best and worst way.
bashtoni 21 hours ago
Is AI going to put Mathematicians out of a job, or is it going to create many new jobs reviewing proofs it creates?
I'm not sure it's clear right now.
lisplist 20 hours ago
I'm not a very good mathematician, but I do know a fair bit about software engineering, and with AI I've been busier than ever. I probably wouldn't be so busy if AI was better at anticipating what I actually wanted rather than making guesses no human would ever make.
This is just a short term problem though. Eventually AI will get pretty good at figuring out exactly I want and it will build that from the start. The requirement of me reviewing the AI output only lasts as long as models stay bad at anticipating my needs, which I don't think will take too much longer.
Aperocky 8 hours ago
There's a hidden problem you skipped over. The model may get good at anticipating your needs, but are you good enough at anticipating your needs?
margorczynski 12 hours ago
> reviewing proofs it creates?
If you mean reviewing for correctness then no, a Lean proof is a much stronger guarantee than anything that can be provided by any human.
For someone who's goal in math was taking unsolved problems and working on them then it's probably over. Just like in software engineering writing code by hand is kinda over.
Nemant 18 hours ago
Can someone with a math background explain the significance of these and previous problems that have been solved by AI?
Does some real problem get solved in physics, chemistry, biology, materials, etc? Or are these fun puzzles for mathematicians with not much real world impact? Eg solving the 8 queen problem in leetcode.
lg5689 13 hours ago
Many of these latest results are very important to theoretical math. Some are also theoretical physics and CS.
Real-world applications are far off, but developing mathematical understanding does tend to leak over into applied physics and CS.
A cynic might say this is all just intellectual games, and though there's a grain of truth, it's too cynical imho. This isn't like 8 queens where there's no hope for applications or generalizations. A lot of this stuff fundamentally affects our understanding of how numbers and systems behave, what are the limits of computation, etc.
Even if someone doesn't care about theoretical results, it's still exciting that AI has become superhuman in a domain as broad as math. That shows there's potential to be superhuman in other domains as well.
LetsGetTechnicl 7 hours ago
And I'm sure none of it was stolen from the actual researchers...
tim333 7 hours ago
All such work builds on the work of others. Hopefully they'll get credited.
LetsGetTechnicl 5 hours ago
Sure I just have a bad taste in my mouth after those researchers were working on one of the millennium problems for a while, had used ChatGPT for assistance and then OpenAI claimed they had solved it.
connor11528 a day ago
will this make the math for building data centers work?
sashank_1509 19 hours ago
No that’s gonna happen when they take your job
plaidfuji 16 hours ago
Underrated joke
sideway 12 hours ago
Solving hard mathematical problems is a strong signal for capability but would it not be preferable to through all these resources to urgent existential issues such as climate change? Finding technological solutions in those areas would be the ultimate capability signal as political alignment at global scale is almost certainly impossible.
fooker 10 hours ago
They want you to pay for OpenAI credits to try and solve these urgent existential issues. That's the purpose of them creating this hype.
What's stopping you?
Raise some funding, shouldn't be difficult if you can convince people it's urgent enough.
myaccountonhn 7 hours ago
AI only makes the problem worse, but the message that will be pushed is that it will solve climate change.
davidguetta 12 hours ago
dna engineering for enhanced carbon capture trees seems in reach. othere less fun stuff could be as well with that technology
loglog 9 hours ago
Climate change is a political problem, not a technical problem. Ironically, by raising oil prices, Trump might have done more against climate change than many people who have been actively fighting it.
Spacecosmonaut 10 hours ago
Mathematical problems are ideal as benchmarks for AI because they have clear problem statements, clear axioms and results that can be verified easily (for lean proofs). I can't blame these companies for using them, although it's unfortunate that human mathematicians seem to become early casualties of AI progress.
I suppose openAI could have focussed their efforts on a subset of open problems that have a clear real world impact and leave aside the more esoteric open problems as a way for human mathematicians to hone their skillset. However, this would have been a short term bandaid. With open models 6 months behind the frontier, any of these problems might have fallen to the homebrewed efforts of enthusiasts early next year.
What is mathematics for? From the outside looking in (I'm a biologist), I have always viewed mathematics as a way to understand reality and to improve our ability to manipulate it. But what I often hear is that mathematics is foremost about human understanding. But isn't that only because it's humans that needed to do the mathematics in the first place? It's not obvious to me that mathematics without human understanding has no value. For example, it might be that P=NP. The algorithms are handed down to us and we can apply them without fundamentally understanding why P=NP.
Mathematics seems to be entering an era where human + machine maximizes performance, much like chess in the 1990s. However, imagine a future where even talented mathematicians are nothing but noise in the machine (as is the case in chess now). A future where AI generates and verifies proofs without humans in the loop. Where mathematics may be beyond human comprehension.
In that future, does it matter that early career mathematicians are inhibited by these developments? Perhaps not. Programming faces the same issue. As AI crawls up the competence ladder, does it matter that fewer people have opportunities to develop the skillset of a senior engineer? Perhaps not. I have no doubt that biologists will face the same problem soon enough.
lhk931122 15 hours ago
As models get better, I think a time will come when it's hard for people to even verify the results. In the end, I think the bottleneck will be people.
kypro 13 hours ago
AI doomers often talk about these kinds of scenarios often, but we tend to assume if humans were given magical math/physics results which we couldn't understand, or given magical pills by AI that cure all disease, we'd probably just take whatever the AI has given us rather than spend years or decades trying to understand the knowledge/technology before leveraging it.
At some point in complexity – especially if we allow our own knowledge to deteriorate because AI can do the hard work – we will stop understanding the world around us. In the same way one day Native Americans woke up and realised they shared the Earth with people who had magic sticks which they could point at someone and kill them, we will live in a similar world very soon too.
What sticks are dangerous, you will not know. Your existence in the future depend entirely on the AIs not wishing you harm, but you don't know how they work to verify their motivations either.
LarsDu88 15 hours ago
The next few years will be interesting.
Surely better materials and pharmaceuticals won't be far behind, and that's going to chanhe everyone's lives.
closetheloopdev a day ago
Hopefully the techniques and results here will be in the training dataset for the next models, so that each new release will give us more interesting techniques and results!
It seems that OpenAI has a proof machine that keeps multiplying fruitful proofs!
eecc 13 hours ago
My only true worry is if AI begins to see us as competitors for resources, energy in particular.
It could decide to let us starve and die of exposure to secure all energy resources to itself.
We'd better use "dumb" and "not fully assertive" AI to solve fusion before it spins out of control (or alignment).
needfish 8 hours ago
Not even needing the AI to think that, US consumers of the electric grid are already subsidising the unpaid bill of data centers. Just need those who are in charge of infrastructure to prioritize the AI consumption over humans, and shifting the books to make humans pay more for resources.
yewenjie a day ago
A lot of these seem to be proving conjectures rather than finding counterexamples, a lot of people used that to claim that these models are not really smart/creative etc.
That copium didn't last for what, three months?
zeroonetwothree 21 hours ago
I acknowledge I am impressed how quickly it moved beyond just counterexamples.
sebzim4500 a day ago
Don't worry, more copium will be delivered. TBF so far it's still only solved the easiest of the millennium problems.
justanotherjoe 2 hours ago
That's sort of how problems work.
pred_ 15 hours ago
I think it's fantastic that they decided to follow the AGMAI advice. Cleaning up their mess will be a substantial endeavour, so I imagine the funding provided to do so will reach well into the millions. But it doesn't look like the press release says anything about how they will fund it at all?
pullshark91 5 hours ago
I'm sick of it. This has happened so many times in my memory. Some AI company announces that they did something fascinating, and it turns out it is all just hype and slope in the end. I still believe that LLMs are a dead end. Most people here are basically like "I've no idea what's going on, but I'm so happy and LLMs are so cool". To think that doing enough linear algebra would solve all your problems just feels wrong. I guess I'll simply wait until someone interprets the results and explains what's actually going on.
qnleigh 4 hours ago
If you scroll around on this thread, you'll see mathematicians discussing how they worked on one of these probably for a quarter century. Others arguing whether another is the biggest number theoretic breakthrough in 100 years or 200 years. It looks like they have made substantial progress on 4 Millennium Problems now. Unless it turns out that almost all of these results are wrong, what would it even mean for LLMs to be a dead end now? This is a historical day for mathematics and it will not be forgotten.
ayden93638 12 hours ago
Keeping the tech proprietary so that it can only be used on these problems by internal teams is the very definition of gatekeeping
karannb 21 hours ago
I think such "process-oriented" fields shouldn't be subjected to mass automation (at least not till we as humans have sufficiently leveled up our game and understanding).
What I mean is progress in math comes from having gained a deeper understanding of the problem for subsequent attack of more problems and IMPORTANTLY applications! Right now the first one is trivially satisfied (given oai maintains some memory across models) but the second one is not! It's generating proofs faster than anyone can validate and so only the model can use these. Consequently if it just keeps doing more theory it's... not very helpful or at least not optimally helpful. This is just bragging rights for now.
More "application-oriented" research would be awesome, where it tries to achieve some desirable effect and then produces relevant theory and experiments around it. Fields like CS, Physics, Chemistry, etc. This would also benefit a wider section of the population rather than the 10 people who understand most of these proofs.
gignico 15 hours ago
> Generalized Star-Height at Most Three
This was an open problem in automata theory I worked on for more than one year before giving up. I'm very curious about their claimed proof.
teekert 16 hours ago
I recently saw a YT short of Grant Sanderson on AI in math and found it (as always) very insightful. But I'm sorry, I never ever find anything back on any of these ad-ridden platforms these days, and perhaps it was a short with content stolen from some other longer content anyway. So, if you feel like getting informed, somewhere out there is some nice content by Grant Sanderson.
Apologies for the rant, I really tried to find it. It had something to do with not being able to predict what this influx of proofs may bring us on a meta level, it could be very interesting. But he also had some critical notes about the missing process and the things found along the way.
boccaff 11 hours ago
Probably it was his recent appearance on numberphile.
xydac a day ago
i wonder what it means for maths researchers, and how it aligns with how they approach math problems.
blooalien a day ago
> i wonder what it means for maths researchers, and how it aligns with how they approach math problems.
I guess their job now is "Idea Man" and "Error Checker"? Kinda like (some/many) "programmers" these days.
xydac 21 hours ago
just wait till someone builds a idea generator model - wire it to decision (jev-like) classifier -> loop it back to researcher
>> may be thats what open ai did :)
patcon 20 hours ago
I'm a little concerned, but I'm also glad they're publishing quick before the department of war starts making national security claims of related to using the knowledge for cryptography and/or weapons
spmartin823 16 hours ago
Can anyone with a compression background say how important "Polynomial-Time 2-Approximation for Shortest Common Superstring" will be practically?
chickenjoseph 20 hours ago
This achievement feels like a reasonable candidate for the moment where LLMs are officially "more intelligent" than any person. How can we justify moving the goalposts yet again? How can people have grown so numb to seeing advancements that they don't read this as significant? I feel like I'm standing at the foot of the exponential. I am not excited for the future, and I don't see how humans retain meaningful control over the future if we continue on this trajectory.
I have been lurking for quite some time. I made an account to post this, but I honestly don't know what to say. I would like to get off this wild ride.
iyyg 18 hours ago
You’re struggling with nuance.
Would Einstein be successful at running apple? Nope
This seems very hard for people to understand.
It will be painful for many to realise - you should focus on doing something that positively affects the economy. Everything else is noise and many endeavours are transitory.
subhajeet2107 12 hours ago
an AI that is at 100s of Einstein level in every intellectual field imaginable ?maybe, running apple does not require one to be very smart or intellectual
aaeieje 9 hours ago
slopinthebag 17 hours ago
llms are good at different things than humans, we are still collectively figuring that out. the idea that an llm is more intelligent than humans at math of all things seems fairly unsurprising.
lf88 20 hours ago
I feel this moment is one of the last few warnings before things will get seriously out of hand. We need to stop now. Building a superintelligent AI should be considered a crime against humanity.
dyauspitr 19 hours ago
Stop? What would possess you say something like that right now it’s going full steam how do you not want to know where this will go?
pixl97 18 hours ago
dualvariable 20 hours ago
How many of these results are incorrect?
I doubt the answer to this is "none".
And how many of them are just exploiting some loophole that will need to be closed in the problem definition?
curtis-jm a day ago
You can read the papers here: https://hub.valency.io/collections/openai-math
Painsawman123 9 hours ago
It's funny how people see "ML" models becoming superhuman at proving mathematical theorems as a sign that we're about to enter the singularity (whatever that means) or that it somehow justifies the valuations of those companies...But what I see is a scenario where companies have spent trillions of dollars on a technology, and the most non-trivial thing that technology can accomplish is a pseudo-form of automation for mathematics and coding! Remember that early on, when this whole bubble started, investors were promised that "AI" would eventually capture >70% of the world's jobs! but it could very well be the case that the only ones they're going to replace are mathematicians(and i'm not even sure about that!)!
a57721 8 hours ago
> the most non-trivial thing that technology can accomplish is a pseudo-form of automation for mathematics and coding
It's irrelevant that this is a pseudo-automation, what's important is that techbros can convince people who make the decisions and concentrate wealth that this is a full automation. So expect reverse centaurs in increasingly more professions in the future.
MrOrelliOReilly 9 hours ago
Tell me you you don’t use frontier models regularly without telling me you don’t use frontier models
ngl999 17 hours ago
We have just heard a few days ago how many of the Linux security problems reported by Claude are real.
TeeWEE 21 hours ago
This feels like a huge AI slop dump. Who validated these results. Why are they not mentioned. Human understanding is key here.
In my experience AI (frontier models) sometimes does weird stuff that needs human review. Not that’s incorrect but sometimes overly complex language or weird use of language.
Ra8 8 hours ago
+1. And there's no guarantee that OpenAi haven't stolen unpublished research from multiple professors and PhD students.
Ancapistani 18 hours ago
They have formal proofs included, that’s the point.
2sk21 13 hours ago
Unless humans have gone through the proofs line by line and verified them, this all remains unproven.
frozenseven 10 hours ago
xyzsparetimexyz 17 hours ago
Yeah but how am I meant to verify that the proof is proving what it says it is?
margorczynski 12 hours ago
TeeWEE 15 hours ago
Note true for all of them.
saberience 12 hours ago
Who made the formal proofs and who checked them?
jrflo a day ago
Glad to seem them changing their tact with the whole NS debacle. Hopefully we can all focus on the results now rather than the surrounding drama.
binlog 19 hours ago
They didn't change anything lol. The news cycle has just moved on.
rifty 20 hours ago
As AI works steadily through the already discovered unsolved problems in mathematics, who is currently discovering new ones?
youoy 17 hours ago
Sshhh dont say it out loud, someone might hear you.
sashank_1509 14 hours ago
Can the mathematics field come out of this stronger and better. I doubt it, things will only get worse. There is no future in being a Proof Digester, no valuable human is motivated in doing that, nor is there any glory in it. But there are some potential pathways for maths to come out strong from this:
1. Relentless focus on quality. Every publication must act as if it’s going to be included in a future textbook, that is a newcomer can get into it given a reasonable amount of time, and math priors learnt in undergrad. (NO AI Slop proof passes this bar as of now)
2. Limit the publications per year. Each author is allowed 2 with a max of 50 pages. This allows the author who chooses to not surrender his cognitive capacity to the machine, still be allowed to play this game. Of course who wants to orchestrate a thousand agent workflows, is free to do so, he is only limited to 2 publications.
3. The aesthetics of the field changes from purely solving the problem to solving the problem with simplest most elegant set of ideas. What 3 sets of simple ideas solves large swathes of problems, that should be given a fields medal, not purely solving the problem, which the AI will be able to do.
jhonof 7 hours ago
> There is no future in being a Proof Digester, no valuable human is motivated in doing that, nor is there any glory in it.
Coding has already gone this way and plenty of people do find motivation and glory in the final product vs. the building of the product (myself included). I absolutely see value in a person being able to humanize llm math output and see the field moving in that direction.
sashank_1509 6 hours ago
I don’t think you’d be doing this if you were paid peanuts. I also think coding is in a transitional period, most people who claim they are “building the product” are going to lose their jobs as the models get better.
jhonof 6 hours ago
nonmaskable 9 hours ago
this is crazy ... at this point will researchers still exists. not sure about that. kinda sad
electroweak 17 hours ago
It must be so frustrating to write science-fiction now with the future changing so fast.
pugfugly a day ago
holy fucking shit
rafterydj a day ago
I don't know, this does not feel like the message hit OpenAI where it needed to hit, if this is their primary response.
osiris970 a day ago
You want them to stop doing math research?
turzmo 5 hours ago
If the end result of this is that within a few years, AI "does" all of the mathematics that humans do, and that there is nobody around that understands any of it, what was the point?
ncr100 18 hours ago
This website needs a SPOILER tag.
It inspired grief in one mathematician posting here.
dgacmu a day ago
I find the claimed matrix multiply result (w<= 2.25) shocking. I hope it holds up.
lionkor 14 hours ago
Can someone who is more into math or AI explain why so many people are so incredibly excited about this?
If OpenAI started opening hundreds of PRs on long-open issues on popular open source projects, would we rejoice, or would the first reaction be "they are unreviewed, so slop until proven otherwise" (it would be that).
I cannot possibly see how these are so impactful, especially the ones that don't come with lean proofs.
LLMs have the ability to make millions of mistakes per day, whereas humans can only make so many. How are we suddenly all so confident that there's no extensive hallucinations or "gaming the system" going on?
schleck8 14 hours ago
> How are we suddenly all so confident that there's no extensive hallucinations or "gaming the system" going on?
What would that look like for the proofs that have lean attached?
lionkor 13 hours ago
I said "especially the ones that don't come with lean proofs", but even for those with lean attached, lean is software, it has over 1k open issues, and I would not put it past an LLM to identify a bug and exploit it to pass the gate.
anthonyrstevens 7 hours ago
"I still don't like AI, and I'm 'just asking questions'."
lionkor 6 hours ago
I use AI every day, for hours, and it's not because someone is forcing me to.
Asking questions is reasonable when these tools are so very fallible.
theoa 19 hours ago
What's missing for me for each result are the following:
* Explain the result to me as if I'm a 10-year-old. * Create the infographic for this result. * Make a Khan Academy-style video to teach me this result.
lokl a day ago
Do applied math next.
edward_d 19 hours ago
This repository contains mathematical manuscripts and supporting proof artifacts produced by an internal OpenAI model?
davitparks 10 hours ago
Exactly nice post
williamhm 19 hours ago
And this is the result of a discussion between users and the platform; it's great that they listened.
i_idiot 21 hours ago
If only AI can better humans in meditation...
kevinwang a day ago
wow
matapassiones a day ago
Valency has the papers up on Valency Hub
OutOfHere 21 hours ago
Is this now a new home for quality AI produced works?
cute_boi 18 hours ago
"Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study (opens in a new window) "
I don't think this is correct solution to this problem? What about software advisory where you form similar group etc..?
I am thankful, I don't have to deal with petty academia politics....
rrr_oh_man 10 hours ago
Maybe we’ll have vibe mathematicians now
aaraujo002 a day ago
The Advisory Group states in its recommendations [1]:
"We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."
To me, this is a take against progress so that mathematicians can keep their jobs. What would we do if, instead of math, we were talking about diseases? Are we going to keep diseases around so that doctors can keep their jobs too?
tchalla a day ago
Why did you leave out the entire quote?
> At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.
To me, the issue is that the models are proprietary which are only accessible to a few people in 2 digits. It's not about progress but access.
aaraujo002 a day ago
Maybe, but it still seems like an excuse. OpenAI has a proprietary model capable of solving these problems and is willing to share the results with the mathematical community. So basically, the ask is to just not use the model and leave the problems unsolved?
adrian_m a day ago
TeeWEE 21 hours ago
mattr03 a day ago
I don't think this is a reasonable take at all. Most work on maths has no real benefit other than to further human understanding of maths - it's more like an art. Nothing is gained from OpenAI solving all these problems but taking jobs from mathematicians. Other than advertising for OpenAI at least. It could not be more different from having AI work on disease research etc.
bravoetch a day ago
> Most work on maths has no real benefit other than to further human understanding of maths - it's more like an art.
It's been a while since I was reminded of this xkcd: https://xkcd.com/435/
zeroonetwothree 21 hours ago
Most math isn't just proving novel famous results. Just like most of software engineering isn't writing code.
jhrmnn a day ago
It all hinges on the definition of “progress”. The debate of the past month is all about questioning whether formally proving outstanding unproven theorems without human understanding constitutes progress. This is quite different from solving diseases.
andriy_koval 17 hours ago
without humans understanding who are currently losing entitlement. Regular Joe could never understand or claimed to understand high math.
medler a day ago
The rest of that document makes a pretty compelling case for why this is a bad practice
esafak a day ago
I fear that professional mathematics will wither, and there will be nobody left to digest the AI results of the future, leaving us unable to challenge the AI.
pixl97 18 hours ago
osiris970 a day ago
Comical ask
bmitc a day ago
Advocating purely for progress and not humanitarian value is how we'll all get enslaved.
perching_aix a day ago
The trope you're drawing a parallel with has a (to me) compelling counter though: there being a cure for every disease wouldn't stop people from getting sick.
How this maps back to math, idk.
warkdarrior a day ago
The latest posts from Terry Tao on Mastodon effectively ask for an AI to explain its results to human mathematicians.
> "I believe that AI can contribute positively in all of these directions [NB: exposition, community building, new directions of study]"
binlog 19 hours ago
There are a dozen+ AIs available to you that can do that right now.
fph a day ago
...but we're not talking about diseases. Publishing an AI-generated Navier-Stokes solution does not save lives. (And, in fact, it harms some.)
Yamata 17 hours ago
It harms lives? How so?
fph 2 minutes ago
mi_lk a day ago
Curious if Sébastien Bubeck still work at OpenAI? He came out quite dirty after Navier-Stokes drama
k2xl a day ago
Can someone knowledgeable about the subject outline the most significant portions of the results?
ciaf 13 hours ago
Makes me feel disgusted. Super-intelligence, even when controlled, will cause so much damage. We are giving away control to a super entity, or whoever has the power to steer it.
NegativeLatency a day ago
Why should I care?
voidfunc a day ago
Because it means mathematical discovery can largely be automated away from academics. This is the beginning.
electroweak 17 hours ago
Humans propose; AI will dispose.
It's not clear from this progress that AI can formulate conjectures despite this new ability to solve them. So mathematicians still look like they have a job. Though instead of spotting far-off landmarks it's sounds more like they'll be chasing waves on a beach.
bamboozled a day ago
The beginning of what?
voidfunc a day ago
sunkeeh a day ago
ipnon 15 hours ago
It seems the age old academic model of scientists competing against each other for fame and prestige is done, and now we must merely enjoy the fruits of scientific discovery for their own sake.
dyauspitr 18 hours ago
Does OpenAI have the lead now? Why isn’t anthropic coming up with stuff like this?
modeless 16 hours ago
Seems like OpenAI's next model is better at math proofs than Anthropic's next model (to an extent that it surprised even OpenAI researchers, according to their public comments). But that doesn't necessarily mean it's better at everything else. The models are more spiky than ever before. Wait until they're released to judge.
dyauspitr 19 hours ago
This is like that meme where death goes door-to-door. Currently, he has visited the software development and mathematics doors. I wonder what’s next.
binlog 19 hours ago
Yet there are more software enginners employed today than another other point in history
Catloafdev a day ago
This is a pretty hilarious thing to read juxtaposed with AGMAI's requests.
Basically "Here you go, have fun with this, fuck all your demands, by the way we're gonna be releasing the model stay tuned!"
hi__dang a day ago
Mathematics is solved.
tim333 5 hours ago
Math is infinite so a bit longer I think.
kypro 13 hours ago
Perhaps the most significant announcement of my lifetime. Yet, I suspect I will not see this in any mainstream news reporting.
I feel for those in Mathematics and worry for our future.
Models will only get better and in a few years the models which produced these results will be a bad as GPT-3.5 in comparison to what we'll have in the future.
Please take a minute to consider what this means, and the risks it presents us.
tim333 5 hours ago
It got a mention in the nyt.
digitaltrees a day ago
Gross. After the accusations training on mathematicians conversations and unique methods, to dump this volume of unsubstantiated papers is at best tone deaf. Are they hoping humans review and validate this? The amount of free that they are coat tailing is outrageous.
computerex 21 hours ago
What do you expect them to do?
digitaltrees 18 hours ago
Not train on and steal user data to front run frontier research for one thing. Not take credit for and independently solve these problems and instead sponsor researchers and give them credit. They could chose to include humans and frame AI as elevating human systems. Instead they are reveling in the fact that they did this in their own instead of humans. Thats a PR choice that is short sighted and reflects a selfish mindset not deserving of leading this transition
computerex 5 hours ago
tootie a day ago
Seemingly none are vetted and reviewed yet
mulemisterX a day ago
That's our job.
esafak a day ago
Ain't nobody paying me to do that. It's kinda sad that maths is being reduced to checking the AI's work.
kozikow a day ago
zeroonetwothree 21 hours ago
TeeWEE 21 hours ago
No it’s OpenAI’s job. They are acting as a meat proxy
red75prime 20 hours ago
schleck8 a day ago
Most are formalized in Lean, about 80% of what I checked
nautilus12 a day ago
Have any real mathematicians working on these problems reviewed any of these and determined if they are just gobbledegook or not?
The ones with lean proofs could still be formulated incorrectly
gyanchawdhary 12 hours ago
AI may be one of the most communist looking technologies in the classical sense .. i mean it dosn't abolishes private ownership .. but it DOES make intellectual capabilities that were once scarce and concentrated available to almost everyone ...
tim333 5 hours ago
It may well lead to a socialist type set up. I mean if AI produces all the wealth why not divide it up equally amongst us humans?
redox99 a day ago
The stochastic parrots have predicted the next token once again.
anthonyrstevens 7 hours ago
How deep must your head be buried in the sand to trot out this comment, on this thread.
simianwords 4 hours ago
the post was clearly sarcastic
globalnode 18 hours ago
imagine your a post grad maths student looking for hard problems to solve, and theyre all solved..
baggy_trough 20 hours ago
Stochastic parrot truthers in shambles.
koe123 17 hours ago
Can you explain it without reaching for lofty things like consciousness? To me stochastic parrots is literally how it works given that it’s “just” the most impressive data fit we’ve ever done. Apparently generating mathematics reasoning traces + verifying them with lean works super well.
tim333 5 hours ago
The network goes through 100 plus layers which end up doing all sorts of processing in ways we don't quite understand because it gets there through gradient descent but is probably similar to how human brains do it.
I daresay parrots can be quite smart too but I don't think that's what the critics were referring to.
slopinthebag 14 hours ago
no ur right it's actually god
AmazingEveryDay 19 hours ago
It is more of the usual though isn't it? OpenAI cribbing off of mathematicians that have used their services; deciding to put a lot of compute behind fruitful areas of endevour; getting results, then taking credit.
oh_no 18 hours ago
no, they did not steal the notes of 100s of people working on these 100s of problems, be serious
baggy_trough 18 hours ago
It’s wonderful.
mathisfun123 a day ago
With so many results in so many different areas no way they even remotely spot checked well enough.
Prediction: one of these is wrong and this (publicity stunt) will backfire.
Edit: don't tell me about lean. For lean to function as a proof certificate you need to represent the theorem correctly. Again: good luck doing that across such a broad swath of problems.
jojva a day ago
You have not read their readme:
> Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly. We are also exploring community-hosted repositories for these materials.
mathisfun123 a day ago
i have and i'm exactly saying that if it comes to pass one of them is wrong it's going to backfire. ie yes that's my exact point/bet.
stevenhuang a day ago
bravoetch a day ago
What does a backfire look like? It's ok to be wrong in the science/math world.
mathisfun123 a day ago
of course in science/math it is but it's not okay if you're a business selling supercalifragilisticexpialidocious infallible intelligence.
bravoetch a day ago
orlp a day ago
It's likely that way more than just one of these is wrong. But even if it turns out 80% is wrong this is still 100+ results...
applicative a day ago
Why didn't they just give mathematicians access, so they could at least understand and write up the results in publishable form?
stevenhuang a day ago
They're all here https://github.com/openai/math/tree/main/preprints
zamadatix a day ago
I think they meant "access to the model" rather than the results.
oh_no 18 hours ago
applicative a day ago
I wonder if the Lean compiler can change it's license so that a for-profit corporation can only use it if it pays, say, a few hundred billion dollars. This is the correct path.
sandworm101 a day ago
So the million monkeys at a million typewriters have churned out 700 shakespeares, but they need me for spellcheck?
utopcell a day ago
Nobody needs you to do anything, not with that attitude.
blurbleblurble 19 hours ago
This just looks entirely obscene from a PR perspective. It's like the cable companies networks owning the networks. At least form some partnerships to obscure the total narcissism party.
senderista a day ago
Good to see they're engaging with the mathematical community, even if they had to be publicly shamed into doing so.
sebzim4500 a day ago
This is the opposite of what the mathematical community were asking for. I say this as someone who strongly approves of this approach.
tim333 5 hours ago
How so? I saw at least on mathematician say publish what you've got. What were they supposed to do differently?
senderista a day ago
Yeah I'm not sure they met them even halfway.
youoy 17 hours ago
ASuperMegaI lab announces they have shrinked the space of unproved mathematical staments from infinity to infinity, giving us proof of their deep mathematical knowledge and understanding, and how they grasp the concept of infinities.
For me this reads as someone boasting about how they go to buy bread on a ferrari to the supermarket, while i sit listening to it, having no idea what they are talking about. And then i stand up and go walking to my favourite boulangerie.
Warning: if you are from the USA you may be triggered by this metaphore.
vessenes 17 hours ago
Willful ignorance is a vibe; did you read any of the GitHub? There are some stunning results in there. I get a similar feeling skimming through the topics that I do watching a successful space launch: it’s pretty cool humans built this. Unlike a space launch we are likely to be able to pass all of this information down to our grandchildren - space launches involve a lot of finicky engineering knowhow, but pure math results tend to be sticky over the last few thousand years. I find that hopeful.
FWIW I also like bread.
youoy 16 hours ago
Dont get me wrong, i am 100% impressed by the technical capabilities, and appreciate the significance, and the amount of the results. This is a very special time to be alive, never in my dreams i would have thought to see this.
To follow your methaphore, who is directing the spaceship?
This feels more like fireworks than a space launch. Space launches would not have happened without having fireworks first of course, but I am looking forward for the space launch moment.