I think it can both true that 1) OpenAI is being inconsiderate/harmful/<pick-whatever-adjective> with their math releases, and 2) there is now a treasure trove of mathematical results ready for the taking.
Yes, the situation sucks overall and mathematics as a whole is in a turbulent time now.
But it also sucks when mathematicians, who are considered experts on a particular problem, refuse to engage with breakthrough results about that problem. #2 above is still true regardless of where it came from or how hard it can be to absorb.
While reading "The Mathocalypse" post [0] by Scott Aaronson, Scott described his wife Dana's reaction to one of the newly solved results in her primary domain of expertise, on which she'd been working for decades.
After her initial shock, and annoyance with the format/style, she decided to start using Astra - for the first time - to help her understand the new result. And he reported in the comments that she had made a lot of progress understanding it in one day, and may be even excited to give a talk about it!
That seems like a much healthier attitude towards these new results.
Yes, everything else sucks about this messy period. But there are still diamonds (in the rough) in this drop that perhaps should be looked into. If the author is too busy, perhaps one of their students can take a look? Someone will, eventually.
Why are you two-siding this. There is only one misaligned agent here and that is the company OpenAI. What OpenAI is doing sucks, and people are calling OpenAI out for sucking. However mathematicians (the main victims of OpenAI’s lousy behavior) behave is not the issue here.
If some mathematicians complain about OpenAI sucking, that is fine actually, and if others are more “mature” about it, that that is fine too. Neither of these reactions should be at put as an equivalence to the blame OpenAI deserves for this stunt.
I know I am, though maybe not as visibly so as your parent. Us anti-AI Luddites are arguing against a massive propaganda machine with trillions of dollars on the line. It is not easy, and it is mentally draining. Especially since AI-companies are behaving in an obviously anti-human, anti-science, and anti-industry ways for their own short term profits.
And worse yet, we have seen this before (albeit on a lesser scale). GMO was supposed to solve global hunger, Carbon capture was supposed to solve climate change, etc. etc. These were obvious lies and propaganda back then as much as the lies where AI are supposed to help humanity today. It gets tiring, and at some point we simply crash out.
unfortunately technology will never fix humanity's tendency to ruin itself (see: war, climate change, etc; all of which we know the solutions for ie 'stop killing people' 'stop using oil' etc) but that is not a fault of the technology itself (gmo, carbon capture, crispr, llms, etc)
This is simply not true. Plenty of technology exists which has helped humanity unambiguously. But there is a difference between good and helpful technology and a grift. GMO could have been helpful but it was mostly just a grift, GMO solving world hunger was always an obvious marketing stunt, and it worked, even though many of us (particularly on the political left) saw it as such and called it out as such. Ditto Carbon capture.
AI is currently the grift of choice.
Also I fundamentally disagree that there is any tendency for humanity to ruin it self. There are perverse intensives, which rewards undesirable behavior, there are unfair, undemocratic, or otherwise hostile powerstructures which allows one class of peoples to exploit another class of peoples. And some technology sometimes plays a pivotal role in maintaining or furthering these hostile power structures/perverse intensives.
Are they fully “doing math” in an aligned way though? They are finding mathematical proofs, which is important, but contextualizing results in the prior literature and clearly communicating the approach and implications is just as much part of mathematical research.
You might argue that these aspects of math are less important in the new AI accelerated math world, because agents will inevitably be smarter than humans, but I think clear framing and communication is even more important than before because with this technology we can choose to augment our intelligence instead of defer it
> Are they fully “doing math” in an aligned way though?
No matter what you think "aligned" means here, publicly releasing advancements in a scientific field shouldn't be gatekept or be perceived as misaligned in any possible way.
If anything, the true misalignment comes from people trying to prevent these advancements from happening or being disclosed, or putting research behind BS paywalls.
I don‘t think this is gatekeeping. OpenAI is filthy rich, they can afford to pay mathematicians to go over these papers and produce results consistent with the scientific method. The fact they don‘t is evidence of their misalignment. Their motives are obviously ulterior and have nothing to do with advancing knowledge. If they were they would adhere to the scientific method, hire experts, and produce reproducible results.
> Putting research behind BS paywalls.
Yes, that is gatekeeping. But there are more then one way to be misaligned. And the practice of publishers is not being discussed here. No need for whataboutism.
Do other researchers or research institutions “pay mathematicians (other than those in their employ) to go over these papers” or do they simply publish their work for review as part of producing “reproducible results”?
AI may be taking over Mathematicians’ jobs, along with everyone else’s, and it’s OK to hate that, but the scientific method says nothing about that, or hiring experts, or ulterior motives, or “being filthy rich.”
Dredging every inch of the lake to catch every single fish and dumping them all in one big pile the town square is just “solving fishing and releasing the fish to the public”, this anti-trawler hysteria is getting really funny.
Paraphrasing Hardy, 'exposition is for second-rate minds'.
I don't believe this myself. But I do believe that if you've formed your very ideas about what is good and desirable on the basis of a culture that has held certain values dear for hundreds of years, and have fought against every doubt and difficulty in life for decades to mold yourself into that image, that it does not 'suck' that you are unable to adapt to a new reality overnight.
Very few people that would love to be craftsmen would love to be factory foremen. It is far too insensitive to the human experience to expect people to just deal.
> Paraphrasing Hardy, 'exposition is for second-rate minds'.
Forgive me if I have no sympathy for current mathematicians who think this way. It's a pretty ugly kind of arrogance.
Some people told themselves they were the pinnacle, the first-rate mind, as opposed to all the second-rater. Well guess what, now your first-rate mind is a commodity and exposition is more valuable. They'd better learn to live with it.
It turns out your an expendable commodity, and my computer system is predicting that society will profit from your annihilation, really sorry about that and I hope there's no hard feelings.
> Forgive me if I have no sympathy for current mathematicians who think this way. It's a pretty ugly kind of arrogance.
Thankfully, very few mathematicians share Hardy's opinion, just as very few share his opinion that "mathematics is a young man's game" (and indeed we now have prizes like the Abel Prize with no age limit).
In fact, many of the greatest mathematicians throughout history have taken exposition very seriously, e.g. Euclid, Euler, Lagrange, Cauchy, Dirichlet, Kolmogorov etc. all wrote textbooks. Many mathematicians today carry on that tradition of taking exposition seriously and write books and freely share their lecture notes.
So we should not take Hardy's opinion as representing the opinion of all mathematicians or even most mathematicians. In fact, Hardy's statement is somewhat self-contradictory since he himself wrote several expository books (e.g. "A Course of Pure Mathematics").
His statement was self-deprecating (in a not so endearing way), as he was referencing his young self as one of those first-rate minds but his current old self doing exposition as second-rate.
I know that it was self-deprecating, but that doesn't redeem it much to me. I also still find it pretty contradictory/ironic because he wrote A Course of Pure Mathematics when he was in his early thirties.
A first rate mind communes with mathematics not matter terrestrial. An Uber-Erdos entering to lecture is indifferent to the auditorium as well as it's contents. He's at the board silently communicating new mathematics. Not for himself, not for you, but for mathematics.
> But it also sucks when mathematicians, who are considered experts on a particular problem, refuse to engage with breakthrough results about that problem.
It might be a shock for you but they are very few in numbers. Most of researchers I know are always busy with something. They cannot just drop other responsibilities for something like this. They will take their own time getting through the proofs (if they want to).
> That seems like a much healthier attitude towards these new results.
Another thing to consider is not all mathematicians are from US or with good funding. The PI or graduate students cannot afford to pay 200/month.
I think the role of specific _human_ mathematicians at OpenAI should not be understated.
TFA was about a niche topic that OpenAI doesn't have in-house expertise in.
Otoh Aaronson is the co-author on Lijie Chen's (reasoning lead at OAI) top cited paper. OAI have deployed their resources more effectively against UGC that some of their staff are already familiar with
That's true. OpenAI has some of the best talents. In my experience, domain experts get the most benefits from the models. They can work much faster, catch false positives, and stir the model in right direction.
I wish they take a bit of more time to communicate the findings effectively.
There are good reasons not to delay publishing at all:
> They should release all their results immediately. (Imagine working on one of the problems they already solved.)
This is the most popular answer to a question regarding AI advisory group and immediate access on a popular website for professional mathematicians: https://mathoverflow.net/a/515442/473286
The whole debate regarding the behaviour of OpenAI is a red herring. Mathematics need to redefine their profession and how they work (like us software developers too). There are very good reasons to believe mathematics has an important role to play. If they could just stop talking about OpenAI and get back to work - they are very much needed, in particular now!
> If they could just stop talking about OpenAI and get back to work - they are very much needed, in particular now
This kind of phrasing sounds particularly empty. We are not in WWII researching the nuclear bomb. What are they so urgently needed for to drop everything and work on understanding openai's proof on partition principle and axiom of choice?
Where did I say they should "drop everything"? I hope to read more about how they envision their future, instead of all this regretting and whaling how one company (that I very much dislike too) published a large amount of proofs. They will get more of them, very soon - if they like it or not -, and I wish that would be the primary subject of the discussion. And yes, I'd also hope they engage with the published proofs. The more raw these proofs are, the better. If AI companies start selecting mathematicians to write nice expositions of their proofs, this is doomed to become a very elitist science.
Says who? The professional mathematician writing the article disagreed. Why should people start dancing the tune that openai wants to play for their own reasons and interests? And I do not see how taking the time and effort to write a proper exposition makes it "a very elitist science" when this exact effort and time is needed to actually get other experts understand and build on a result. Unless you equate spending time and effort learning math as "elitism", which is the ai-shilling moto some time now with everything time and effort related. I cannot see how spending time and effort to understand a field and then spend time and effort to make a proper exposition so that other people can also understand it as "elitist" vs throw everything out there "in raw form".
This was my opinion. But anyone arguing to publish "results immediately" is likely to imply something like it. I guess in chemistry we have the situation you envision - for different reasons: Laboratories holding back their data, until their scientists have published their papers or developed their products. There is a real danger AI companies will do something similar too.
Who do you expect they will select for the exposition?
On which basis do you want a (likely US based) AI company to decide who is to untangle a proof that their latest internal model has just spit out?
Do you expect this to fall to an aspiring, but still unknown mathematician at - say - the mathematics department of Nairobi university?! This is what I meant with my rather unclear "elite" reference: The first publication will always show the name of a mathematician already known to the field, more likely than not to come from the same country as the company ("Our message ... is, you’re a great American company, but you’ve got to hire great American workers"). Are you not worried at all? Don't you think it would be good, if anyone in mathematics had a chance to write that first paper on a new proof?
There are parts of mathematics where the result is the important part. That's not what we're seeing here. Knowing whether the partition principle implies the axiom of choice doesn't meaningfully shape downstream knowledge and decisions. For these more foundational problems, clever proof techniques and the exposition around them are literally the point. Without that, neither humans nor AI can take this slop and derive anything useful.
And if AI can do that then great. I don't care about being elitist or not, and I'm fine with AI taking over math. That's not what it's done though, at least not yet.
> Mathematics need to redefine their profession and how they work (like us software developers too). There are very good reasons to believe mathematics has an important role to play. If they could just stop talking about OpenAI and get back to work - they are very much needed, in particular now!
The root issue is OpenAI et al.'s thoughtlessness in their engagement with a field.
OpenAI has resources.
That they fail to allocate enough of those to cleaning up pre-print papers (that seem to be a corporate PR priority for them to release) so they can be consumed and engaged with by the field they're targeting is... acting like a jackass?
It's the same "Meta / Alphabet can't vs won't hire more human reviewers" problem.
OpenAI could, at an immaterial salary level to them, pay a ton of PhD students and mathematicians just to clean up their proofs and papers.
Not doing so is a leadership and financial choice.
I'd prefer AI companies don't decide who's "cleaning up pre-print papers". This should remain the job of mathematicians at universities, which I am happy to pay with my taxes. Ideally there was something like a Bermuda Principles declaration for mathematics (https://en.wikipedia.org/wiki/Bermuda_Principles). This gave mathematicians even at poor universities and beyond the chance to participate in mathematical progress.
What do you think is gained, if AI companies manage "cleaning up"? Tax money?
No one is "responsible". If the paper is read, depends on the interest of the individual scientist. Many mathematicians interested in a theorem proven/disproven in a new AI paper will be curious, even if it is utterly cumbersome to extract the relevant line of thought. But don't you agree that the job can only be done by a mathematician anyway - be she/he paid by the company or by a university?!
What you are arguing for is somewhat like the "proprietary period" in astronomy (e.g. see https://www.scientificamerican.com/article/nasas-plan-to-mak... ). But there is no mathematician who asked for the run of AI, i.e. there is no mathematician, who can be regarded as the owner of the result, even if ownership was temporary. Complex proofs might take years to explain in detail (think of the ternary Goldbach problem, https://en.wikipedia.org/wiki/Goldbach%27s_weak_conjecture ). It's just unfair, if AI companies held back with their data until a proof has been put into a nicely readable article, and at the same time mathematicians elsewhere are spending all their time trying to solve it.
I'm not. To the extent my taxes are used for maths research (which is minute as a proportion of them) I want them to be used for the development of human mathematical understanding, culture and education, not trawling through a mountain of slop mechanically generated by a Silicon Valley startup that's about to IPO for trillions in the hope there might be some nuggets of insight hiding in there. If the latter is an activity with some value the startup in question can pay for it.
> But it also sucks when mathematicians, who are considered experts on a particular problem, refuse to engage with breakthrough results about that problem.
For me the problem is that rigth now the structure of incentives that has been built (e.g. you publish more = you get a grant; good exposition < solving a conjecture) is now broken. So, for instance, you would be very irresponsible if you throw your student into one of those AI papers, it's too much the risk. This part is mathematician's responsability, they need to change this incentives structure.
In any case, OpenAI is being a dickhead here. They throw millions of dollars at these problems, but they can't afford basic literature reviews (the drafts barely cite previous work)? Or checking that Lean's formalizations really correspond to what they claim to prove (even for Navier-Stokes they made this mistake)? It's obvious that for them this is just a PR stunt.
Let us not forget that there's much more to this than just OpenAI being lazy and incompetent and willingly ignoring the high standards that researchers usually holds themselves to.
There's also the case of ethical violations, straight up scientific misconduct, as when OpenAI steals results of others (their customers) and present them as their own.
> The result is also contained in a paper [8] released by OpenAI on October 6, 2026, in which the proof strategy and specific choices of notation are identical to a preliminary version of the present paper that was uploaded to ChatGPT on September 8, 2026.
Of course it's hard to say what to make of that without knowing what exactly went into the machine, but it certainly looks bad. And there's obviously a non-zero probability that it is indeed another instance of plagiarism, given that that's how they operate.
In this case, the author is a grad student, so what we're looking at is a company willing to steal from a student, ignoring whatever impact that could have on their career prospects, for a tiny piece of marketing material.
At this point I would be surprised if internal sandboxes are not trivially by-passed and that openai's agents do not (at the very least) have complete read access to all user accounts, chat histories and uploaded documents. Orthonogally, openai could still be wholesale lying about not training on this user data, of course.
So... If you use ChatGPT for anything of value, including abstract stuff like obscure maths problems, you should assume that at some point in the future OpenAI will include that in their training set and sell it onto other people.
But this is even worse, because there is no way that OpenAI "trained" on this data between September 8, 2026, the date Chenglong Ma uploaded the paper to ChatGPT; and October 6, 2026, the date that OpenAI released a paper with "identical proof strategy and specific choices of notation" (Ma). That's one month, that's not the timescale for model training.
So this implies _not_ that OpenAI is training on user input, in the conventional sense of adjusting weights; but rather that they are *straight-up channeling ideas from user input*, and with a very short lag. You would think there would be about a million controls to prevent this.
This is next-level alarming. I would be very interested in knowing whether Chenlong activated the privacy (do not train, etc) options in ChatGPT, and any other details of their setup (which plan, etc). Also, note that "do not train" might be, in a lawyerly sense, considered by OpenAI to be strictly about weights, and not covering "we hoover up your results and regurgitate them".
Impossible to know without either more information about the actual occurrence or a deep understanding of the person making the claims. Just as OpenAI could easily be acting improperly, researchers who (understandably) feel deeply attacked by their problem getting solved right from under them might not be the most unbiased, either.
> Part of me would like a little longer in the world where the problem is still open and I am still looking for its solution. But reaching a summit, even by someone else’s route, comes with a view. From here I can see new mountains, and I look forward to climbing them with my students, collaborators and the machines.
I think eventually there will be sort of a centralized more or less automated repository for ingesting and sorting ai-generated lean proofs and making them searchable and re-usable. On some level it kind of doesn't matter if mathematicians can ingest the results, if coding agents can just search for them online and use them in their own proofs.
I actually think it would be very smart for the big AI labs to get together to fund an independent organization to manage such a thing, and hire mathematicians to run it.
What is happening now is that some aspects of mathematics are turning into essentially an exercise in software engineering. It is well known that proofs and computer programs have an isomorphism, and I think the eventual merger is more or less inevitable.
That's not to say that there isn't an infinite amount of work remaining for mathematicians to do. There are only so many problems that are going to be amenable to this approach.
I think people sometimes overestimate what proving something means. For example the 4 color problem was a computer assisted proof from a long time ago. I doubt that many people actually read the proof even though leafing through the book is kind of fun if you can find a copy. Another example is the Kepler conjecture on sphere packing. The peer reviewers said they couldn't vouche for its correctness. While reviewers are volunteers with busy schedules, a lack of interest surely was part of the issue. This motivated Hales to formaly check his own proof (Gonthier formalized the four color problem earlier). But you can be sure that no one had been waiting for the Kepler conjecture to be proved in order to pack their spheres. Otoh, at the opposite extreme, the proof of FLT greatly advanced the field because the modularity theorem behind Wiles' proof is at the heart of a great section of math theory.
> there are still diamonds (in the rough) in this drop that perhaps should be looked into.
Right, and for some reason that "trillion dollar" company with mathematicians on staff didn't "look into" their own results. Almost like they don't care about engaging with the actual community they're dumping on.
I can't relate to this at all. AI models will surely get better at writing "enjoyable proofs," but for now the situation is what it is. You're passionate about this problem, right? But you don't want to do the work to understand the result? Fine. There's a new generation of younger, hungry mathematicians that are highly interested in figuring out why the result is true and I am sure they'd be happy to wade through it and spoon-feed you the answer instead. Maybe they should be running things.
So OpenAI should be able to flood the world with AI pollution and ask scientists and mathematicians to wade through it all and tell us if there is any sense in it, then sit back and wait for them to report in?
Yes, they should. They have invented a magic button that can tell you the long-awaited answers to the burning mathematical questions that you've spent your life researching. The caveat is that the technology is still new, so the explanations "are not fun to read" like set theory papers usually are (lol). If you don't think that's a worthwhile tradeoff, that's your call, but it sure as hell isn't everyone's.
I don't think the claim is "they definitely have an oracle that solves the problem, and I reject it because it's hard to read". The claim is "OpenAI claims to have used an oracle to solve the problem. The proof is very difficult to read, and to even know if it does or not, we have to go through it with a fine-toothed comb, but they're going around claiming they definitely solved the problem (or at least getting press that claims that which they aren't pushing back against) and this might convince the people who sign grants even if it isn't true"
People who are invested in the idea that we've invented a general intelligence, now, which includes all these companies that are literally financially invested in this claim they are making, will tend to believe that its results can already be trusted in domains like this. Some mathematicians seem to believe some of the proofs written by their models, and some, like this one, don't. I do think it's valid for an expert to push back against the claim that the best use of their time right now is to verify the poorly written work of everyone who's claimed to solve the problem
Over the year, nothing they've released with a lean proof attached has turned out false (That's sort of the entire point. It's not impossible but it's really difficult). There's a reason most mathematicians, including the ones vehemently against OpenAI's dumping are not arguing the results are secretly false or have a high potential to be. And indeed, if that were the case, it would quickly become apparent and all this worry about grant signers would vanish into the wind. It's very easy to ignore nonsense. The problem is that it isn't nonsense.
I really don't know enough about it to know whether you're right or not, nor do I know whether or not you know enough to make the claim you're making, so I won't make an argument one way or another because it's non-sequitur to what I said anyway. The fact that you or I or Sam Altman or Terrence Tao believe the claim is irrelevant to whether this obligates the person who wrote the blog post to believe the claim, and it sounds like he's willing to consider the possibility that it is right, and would read the paper if it reached a threshold of comprehensibility expected of people making that kind of claim.
I's not a non sequitor because it cuts right to the point. He's under no obligation to read it sure, but that doesn't mean Open AI isn't justified in claiming to have proved it. The justification isn't Sam Altman's belief or Tao's or anyone else's authority. Results accompanied with lean-verified proofs whose formal statements match the problem at hand have arguably stronger justification than the vast majority of human math publications.
> Over the year, nothing they've released with a lean proof attached has turned out false
That's not true. [0]
> On July 25, Ramana Kumar published a repository containing a sorry-free "disproof" of the Collatz conjecture, produced with AI assistance. It is not a valid proof because it exploits a bug in the kernel's handling of nested inductive types.
Even in this dump we're talking about, it hasn't been true. [1]
> In “Algebraicity of Weil classes on split abelian eightfolds” a sign error invalidates a stabilization-trace cancellation argument and the construction used by two dependent papers.
1. That wasn't from Open AI or any major lab. I don't know what random people are getting to. It's curious also that it had no natural language proof attached. Usually these labs have a NL proof then translate to lean. Lot less possibility of lean maxxing.
2. None of the results Open AI retracted had an attached lean proof
The problem isn't just that the papers aren't fun to read. The problem is that a lot of the research that goes into solving these issues leads to other discovers, new fields to explore and people have to develop new approaches to solve them. The other part of it is, the quality and the enjoyment of working on these problems leads people to find new and other interest problems to work on.
If you just strip mine the answers and Sam Altmans magic button solves 100/100 problems, what's next? Who is left to come up with a new interesting question for the magic button to solve?
Lastly, life and the present moment is all there is, if there is no enjoyment in anything we do, then what's the point of all the "living for ever" Altman et al want to achieve.
We will live forever to read boring papers generated by LLMs? Literally sounds like an eternal hell.
Nuclear fusion is already proven by the universe to be a viable energy source by the fact that the sun exists but people still work on understanding and taking it and developing new approaches to accomplish it. People didn't stop experimenting with and developing programming languages because technically they're all turning complete and the first one was "enough." Y'all will be fine - every JavaScript framework that exists is someone looking at a theoretically correct and complete solution and deciding actually it sucks and they could do better. "I want to understand xyz but the proof is trash and I think it's ugly" will be plenty motivation for a lot of people to work on it.
You don't have to wade through the slop. The point is a giant star in the sky figured it out and that doesn't demotivate you from figuring it out yourself
He's tracking the community progress on sub-n log n multiplication. OpenAI started with 1 - 1.63e-55. The result has been now improved on 115 times, and the current record is "rohanarun"'s 1 - 9.87e-5. I'm sure by tomorrow it'll have improved again.
Does this look like people aren't having fun? Does it look like they aren't discovering stuff? It looks like it's spurred a cascade of interesting community activity. It doesn't really seem much different from what happened with the twin primes conjecture. Isn't that supposed to be the point of all this?
People are already finding stuff in the release to get excited about, and as the models get better at distilling proofs to make them more coherent, this will only amplify. Of all the things to worry about, human curiosity and the ability to run with new ideas probably aren't at stake.
In fact, I can't remember a time when I was more excited about the future of science. This could herald an end to the replication crisis, and kill off bullshit science completely. The danger of course is that we end up with two companies effectively dominating cutting edge research in every field, but it remains to be seen if that's even possible given the pace of improvement in open weight models.
I don't understand. Why is this the onus of scientists and PhDs to review whatever results OpenAI had dumped out? If OpenAI had produced incomprehensible papers, surely any journals would just reject it, or demand the author to do a complete rewrite? Unless we are talking about a race to solve problems, which PhDs are afraid that they had been scooped up on?
What else should they do? See these models get smarter and smarter, somewhat-solve things but only to the tune of 90% what mathematicians (or experts in any other field) would deem acceptable, and then gate keep the findings for the next few years going through peer review and paywalled journals? I for one welcome the flood, bring on more in every possible industry and see where all that progress lands up. Sure it will upset a lot. A lot of things also upset the luddites.
Are you going to be doing the work to verify the results ? Will you just expecting other people to wade through the flood and reap the benefits later on?
Nobody is forced to verify the results. I am honestly not expecting anything other than AI to get smarter and smarter and people who are motivated and interested enough to pick up after it; and potentially reap all the long term benefits ahead of those who aren’t (seeing this happening with software development in my own field). But in the end; if you don’t like it, you’re not forced to do anything.
> this is an opportunity for people to pick up where the model left off and run with it
Like people enjoy racing in front of a stopped train? As soon as they turn on the engine again, they will run you over. The questions that remain will be only the low value ones, not worth the effort to vacuum up.
So no, the smarter, hungrier people are not the ones that are going to swoop in. It will be the most desperate.
> Some output is going to be wrong or incomplete
This is a very human take on the situation. No, the Lean proof is not going to be wrong, and it will be incomplete only in the sense that OpenAI didn’t try to push the results further.
They should certainly be allowed to share their findings. No one is forcing scientists and mathematicians to review the findings in general. It’s just the case that the findings are of such such a quality that it would not make sense to ignore them wholesale
> No one is forcing scientists and mathematicians to review the findings in general.
This is like "no one is forcing software engineers to use AI tooling" or "no one is forcing you to show your ID in the airport" or "no one is forcing you to own a car in your small midwestern city" - there can be no law requiring something and the practical consequences of not doing so can be so painful that you're effectively forced anyway.
That is exactly what I’m trying to say. The findings are of such such a quality that it would not make sense to ignore them wholesale.
That’s why it doesn’t make sense to present AI companies as dumping or burdening the scientific community into doing labor for them; the scientific community is self motivated to do so.
It's not self motivated. The motivation is not "this is doing amazing things for us", it's "if we don't review this, the bullshit headline complex and the bullshit-spewing (sorry, marketing) departments of tech giants are going to misinterpret/misrepresent everything and our grant money will be taken away".
You know, I kinda relate to the feeling of not wanting look at those outputs if I think of it from a layman's perspective.
I just imagined that instead of math papers, they released 700+ feature length films, and the only way to tell if one of them is any good is to watch it in its entirety.
That feels pretty unappealing to me.
I know it's the same for human made films, so what's the difference right? But those are good enough most of the time that it's a decent bet, and the people that made them had real skin in the game.
Contrast that with something made by a nondeterministic slop machine with no skin in the game where small details can be off in a way that's jarring. Right out the gate I have an aversion to committing that much time to something that very well may waste it.
That's actually a really interesting thought. Given Sora, and the amount of funding they have, they could have created started their own film festival and dropped 700+ feature length films, had they wanted to go in that direction. But they didn't. Hmm.
The best part is that when the academics fix OAI’s issues, the model gets better and OAI shareholders get richer and more powerful!
As someone who uses LLM tech occasionally, this is why I prefer using open local models. If I’m making myself obsolete, at least I’m not making some asshole richer and their closed model better.
OpenAI in fact didn't know what to do with results and didn't want to flood the world, so they asked mathematicians. Mathematicians recommended OpenAI to release them. My guess is it would have been better for OpenAI if they didn't release them. OpenAI is basically doing this as a goodwill.
People seem to have very misguided ideas about why OpenAI is doing this at all. It is not to brag or to torture mathematicians. It is an eval. OpenAI is known to be willing to pay large amount of money to get a good eval, think FrontierMath. FrontierMath is now saturated, so they need a replacement eval for math. Open math problems are actually a fairly good eval, although a proper eval is better (eg FrontierMath has known difficulty and have tiers from 1 to 4).
Mathematicians would prefer if OpenAI didn't use open math problems as an eval, but OpenAI is not obliged. I actually think OpenAI wouldn't point AI to open math problems if unsaturated FrontierMath Super Duper is available, as it just angers mathematicians, but such eval is not in fact available. Given OpenAI used open math problems as an eval, they could just throw out the result (this is in fact better as an eval since it will keep problems useful longer), but mathematicians preferred to see the result. So OpenAI released them.
As I said, a proper eval is better, but it measures something real that gives a good training signal, and there is lack of good alternatives for math eval. Since the result is 372/4000, it is also unsaturated.
I think this is right - it can be seen as trying to get free feedback from the community. In that way it’s reasonably described as exploitative, since it’s not a good faith effort. The problem of ai slop being submitted to conferences to get publication counts is similar.
"Mathematicians recommended OpenAI to release them."
This is a misleading characterisation of the mathematicians' position.
The very first paragraph of the AGMAI recommendations explicitly states:
"we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models." You appear to have acknowledged this by saying “Mathematicians would prefer if OpenAI didn't use open math problems as an eval…”.
The mathematicians did not ask OpenAI to produce these results. They explicitly asked AI labs to stop producing them in this manner. Their subsequent recommendations concern what labs should do if they have already produced significant results, not an endorsement of the practice.
Furthermore, the recommendation was not simply to release the results, but to responsibly release already existing results. Section 2.B, Step I, explicitly recommends "...labs that have AI mathematical output that is not understood by the people who prompted the AI systems", to search the literature for relevant prior work, provide appropriate attribution, and improve the exposition of AI-generated proofs before releasing them, rather than leaving this work to mathematicians afterwards.
OpenAI published the results on GitHub while still exploring repositories that meet the committee's guidelines. So they followed some of the recommendations, but not all of them and hence, did not release the results as requested by the mathematicians.
I do not think it is a settled matter whether this was done out of goodwill. This is because releasing these results as they were can benefit OpenAI more than releasing them according to the AGMAI recommendations. AGMAI recommended in section 2.B, Step 1.5 that "Each time a solution to a problem is released, it should be clearly documented how exactly AI came to be used on that particular problem. If many results are released at once, then in addition to the results themselves a further document should be written and made public that references all of the released results and explains how many other problems of comparable difficulty the models tried and failed to solve, as well as how the problems were chosen." If the results are released, it is easy to expect that the media will discuss the capabilities of the AI used in the work, as indeed happened. If this AGMAI recommendation was followed, the media would plausibly have also discussed the number of failed attempts and then the overall attitude would not be as favourable to OpenAI as it is now when it comes to the capabilities of the AI that was used. OpenAI did release on GitHub that approximately 4,000 problems were attempted and resulted in 719 manuscripts (after 3 containing suspected errors were removed by OpenAI) across 372 families of problems, but this does not give a calculable number of problems it failed to solve. I do not claim to know OpenAI's intentions or reasoning when these results were released and am not arguing that it was done with improper intentions, only that whether it was done out of goodwill is not a settled matter.
AGMAI's October 6 statement explicitly clarified that its advisory role should not be interpreted as an endorsement of OpenAI's process, and that it was up to the mathematical community to assess how successfully its recommendations had been followed.
Recommending how to responsibly handle the outcomes of something you oppose is not the same as asking for it to happen.
It is also worth adding that the Association for Human Mathematics published a statement (which Tao reposted on his blog) in which they explicitly say the following:
> Mathematicians did not ask for this work to be done. The Advisory Group on Mathematics and Artificial Intelligence, from whom OpenAI has claimed to derive its legitimacy, opened their initial advisory statement by saying that frontier AI corporations should not test advanced mathematical problems on internal models. In ignoring the central premise of the Advisory Group’s position, OpenAI has indicated total disregard for the norms of scientific research — norms that guarantee that mathematics remains trustworthy, ethically researched, and in the public interest.
>> I do not claim to know OpenAI's intentions or reasoning when these results were released and am not arguing that it was done with improper intentions, only that whether it was done out of goodwill is not a settled matter.
That's right honourable of you but for me it is very clear that the only incentive in AI companies' effort to produce mathematical results is to advertise their technology. There is no reason at all to assume they have any other motive; certainly not any kind of interest in mathematics as such.
The problem here is that while you can call what OpenAI does "mathematics", I would hesitate to call it science. Science as a process of acquiring and developing knowledge within a domain involves a lot more than just dumping unfinished work on the scientific community. Among other things, it involves developing frameworks and understanding of the domain, formulating questions, creating results in a fashion suitable for verification/testing/replication, and relating these results to and integrating them with that edifice.
Kinda sounds like computer science vs developing - in the sense that people with a master's in CS and are dedicated to the craft will write wonderfully artistic software, while it doesn't actually take a love for the process to write code and get hired at some tech company (even less so now with agentic development).
Not what I am getting at. In fact, the problem I am getting at is the abolition of existing scientific and engineering principles and processes without a replacement and applies to programming with agents as well. This is not about artistry, but about building durable things.
Have you looked around recently? Because there are plenty of mathematicians that are excited to read and learn about all of the new results. They've improved on the sub-n log n result and have even made a web site to track progress on it: https://beyond-n-log-n.netlify.app/
It looks like people are enjoying themselves, having fun with the new results, and generally doing all of the things you say "science" is supposed to be about. So what's the problem?
I didn't say it is useless. Consider Ramanujan, whose work gave rise to a lot of interesting math, even (and sometimes, especially) the parts that lacked proofs or had other gaps (probably because it was obvious to him unlike us mere mortals). But a singular genius, whether a person or a machine, does not science make.
The question is why the situation is what it is. Did OpenAI publish a large volume of unreadable proofs because that was their best attempt to contribute to the field of mathematics? Or does OpenAI feel that it’s more profitable for them if people come to see mathematics as something that’s less focused on understanding and more focused on using AI to generate proofs?
Yes, I guess it matters a lot whether this was quite close to the best they could do or the best they could do given specific resource constraints or whether they just didn’t bother to try doing better (eg to invest more tokens into readable papers)?
> I guess it matters a lot whether this was quite close to the best they could do
As the author of the post points out, there is no way this is “the best they could do”. It’s a write up that didn’t involve someone with the math + communication skills required to clearly explain the result.
Even with humans, the first publication is rarely the best expression of the lesson. Getting published does more than establish priority; it also frees up the community to build on the result. Nobody expects a human to wait until the proof is comprehensible to anyone except themselves and the referees.
I don’t think that’s true? I never did research math myself, but the people I’ve known who do would definitely invest time in making a proof clearer and better even if they had the idea basically correct. There’ve been multiple recent stories of researchers saying “we’re going to publish this lame proof that isn’t up to our standards because the AI companies let us know they’re going to scoop us if we don’t”.
Mathematicians care a lot about the exposition of their ideas, invest a lot of time in giving talks, writing, and don't publish too frequently, compared to other sciences.
They normally don't feel like they are in some kind of race to publish the results ASAP and claim priority. Cases like that are very rare (but they get media coverage because they are so unusual).
OpenAI did a publicity stunt, their motivation is not to make a good contribution to the field, which has very different standards and culture, compared to the AI labs.
You don't know what you're talking about. Just because it was created by an LLM, and verified in Lean, does not make it true. The whole point of writing a proof is for it to be understandable.
Suppose I directed some llm agents to factor the primes from network logs of your machine. I then publish the private and public key in full. You would not change your keys of course because its not true right?
Knowing a few PhD candidates and non-tenure track postdocs, the “younger, hungry mathematicians” are more worried about finding a damn job in the poorest job market - both academic and in industry - in a decade.
There are plenty of mathematicians who think the results are interesting. Not only do they have the "patience" you speak of, some even seem to be enjoying exploring the new results.
No idea why you're downvoted, my intial thought was that I would like to see some of this supposedly widespread sentiment as well.
I even think it's plausible a lot of mathematicians are excited by it, but the sweeping confidence of the comment you replied to without anything to back it up leaves some to be desired
Checking someone else's work carries a lot of opportunity cost, and is only fruitful if one can learn new methods which apply to the own work. This is pretty risky, especially without tenure!
There's a "Silicon Valley-ism" for you. We offer a thing in whatever form we want and people "who are passionate" will gobble it up, should gobble it up, 'cause they're "passionate".
I think this might feel about the same as getting a PR from Claude that purports to solve some issue that it deems exists, but it doesn’t conform to the contribution guide, isn’t clear in its objectives and looks quite likely to be utter bullshit. I close them without comment and lock the issue.
"AI models will surely get better at writing "enjoyable proofs,"" why? Why is that surely true? they've increased in all other capacities at shocking rates while still writing awful, slippery, turgid prose. Very silly to assume that this will just go away.
Every single day now, for close to 4 years, ever since ChatGPT 3.5 was released - there's been people dismissing AI progress. Every step of the way.
It is entirely possible that one day progress just stops or slows down, but with current evidence, I don't find that too likely - at least not in the near future. The sheer amount of resources being put into this (AI) race is mind-boggling.
So while past performance does not guarantee future results, I'm just going to kick back, and assume that many of the current issues will be fixed with future models.
The progress is clearly not uniform though, it’s shaped by how amenable things are to collecting data and verifying computation. I think writing clear and cogent prose is harder to quantify than formally verifiable logic, and this is why progress in coherent communication has been slower than raw problem solving ability
Maybe this is just a matter of what model developers choose to invest training resources in, but I don’t think it’s inevitable unless clarity is made a higher priority
I think it’s silly to ignore the sheer rate of progress that’s occurred in the last 5 years. Compare gpt 2 output to latest foundation models and then re-read your comment. You’ll realise just how out of touch you sound.
> There's a new generation of younger, hungry mathematicians that are highly interested in figuring out why the result is true
The problem is that right now mathematicians don't have the economical incentive to read these AI generated results. Even if you love mathematics and all that, it's always more important to get a job, and for that it doesn't seem like a good idea to invest time around problems that AI touches because you can't compete with it and you don't know if tomorrow they'll improve by x10 the sota.
I am a mathematician, and I do have the incentive to read the results. Two results in the drop were two major life goals of mine, and all I got was three lousy citations. :) But a third question I've spent a lot of time on is a not-so-hard consequence of one of the lemmas in there. So, yeah, I do have the incentive.
Of course, it's not clear at this point whether reporting such a result even matters, but still. In its own right, it's a very cool result.
Of course, everyone is curious about these results, we're mathematicians after all. But by incentives I refer to those things that really sustain the system: obtaining permanent positions, grants for future projects, ideas for PhD theses, etc.
I think quite a few mathematicians would be interested in using AI to figure out the answer.
But providing the answer in gibberish along with a certificate is not that, it's at best a cruel way to do it, but I'm leaning towards the idea that it's a fundamental misunderstanding of what it means to do math and what it means to communicate a result.
If you think sending an answer in gibberish is acceptable just because it's true then SSdtIG5vdCBzdXJlIHdoYXQgdG8gdGVsbCB5b3UsIGJ1dCB3ZSBkaXNhZ3JlZSBvbiB0aGF0.
Very sensible comments. It is along the lines of the fury I get when I am confronted with an 11 page dump of an issue analysis created by an AI agent that makes no sense but I have to go through because customer shared it.
If you did’t bother to write it, I shouldn’t be bothered to read it.
Perhaps AI agents can have their own publications and magazines where they are the chairs and associate editors and reviewers.
Nobody came to him and forced him to read it, just like nobody is forcing you to read code I had chatgpt write.
If AI can solve such grand, outstanding math problems, and mathematicians argue these pure math problems are important, what’s the problem with them needing to read the output if they want to understand it?
I’m not sure OpenAI shouldn’t have released math papers because some guy made a wager about one of the problems that loosely socially obligated him to read material about a solution to that problem.
The absolutely funny thing is, he is obligated precisely because the proof is very likely to be correct.
The alternative you are proposing implicitly is even crazier. OpenAI should not release a proof that is most likely correct so that it doesn’t burden others. What? It’s not about that guy dude. It’s about the society TM. One can’t delay progress because a guy may be burdened.
“Guys plz don’t release this thing that is absolutely correct but I’m kinda busy with other things ok?”
> The alternative you are proposing implicitly is even crazier. OpenAI should not release a proof that is most likely correct so that it doesn’t burden others.
The alternative is that they do the work to properly present the results. They spend billions of dollars in AI training and inference but can't afford to even cite the literature properly? They're doing the bare minimum because they're inly interested in doing a PR stunt.
they will do it. in a year, the same mathematicians will cry about (O)AI making their lives hard by not only solving more problems, but also presenting them with "enjoyable proofs". They'll still cry because the current excuse is a veil.
Bare minimum is still _solving_ the open problem standing there for years. Nobody owns math. Nobody owns giving enjoyable proofs to someone else.
If you don't like to engage with OAI proof dumbs in current state, don't. Maybe others will. Or maybe _these_ mathematicians are afraid that _other_ mathematicians will do it. Just elitism and gate keeping.
> but also presenting them with "enjoyable proofs".
> Or maybe _these_ mathematicians are afraid that _other_ mathematicians will do it. Just elitism and gate keeping.
Ok, I don't see the point of discussing with you. It's clear that you decided what to believe in and no evidence will convince you that reality is more complex. The proof is that you ignored all the nuances expressed here by simply sticking to your simplistic interpretation, without any explanation of why such nuances are invalid.
Whats the nuance here? Your post implied that the problem was unreadable proof and we are saying that this is not central to the discussion. Unreadable proofs are actually very very irrelevant to this whole drama
I think you’re out of touch with the discourse in the mathematical community, it’s quite relevant because clear communication is an important aspect of intelligence
Proofs can be unreadable for more than one reason. Are these ones unreadable because the math is super advanced or because current agents suck at clear writing? Maybe a bit of both?
Fine and we are saying that this is not central to the discussion because even if models wrote it nicely, there'd be even more outrage.
Do you disagree with this? For example, if openai had provided really readable proofs with utmost care but still dropped 400 at once, would there have been less outrage?
> if openai had provided really readable proofs with utmost care but still dropped 400 at once, would there have been less outrage?
Less criticism, yes
There’s always going to be outraged people, but outrage isn’t the word I would choose to describe the positions of the mathematicians I’ve read on this topic, including TFA. There’s a lot of optimism mixed with frustration that something important is missing
Absolutely. If they cared about properly citing the bibliography, sharing how their models work, etc, it would be so much easier to make an assessment of the situation, of what will be math in the future, and all those deeper questions. What we got instead? from those 400 papers, there are already 4 that have been found to be a copy of recent publicly-available papers. They spend millions for "the good of science", but they can't afford basic scientific ethics?
Why should it change when there's good chance it's just slop? That's exactly the standard that human mathematicians are held to -- and it's their job to refine and polish the work to make it understandable by the community. And standards have evolved that way because otherwise there's too much slop to wade through (eg. you'd be surprised how many papers the typical theoretical physicist gets claiming to have proven Einstein wrong -- from absolute crackpots who don't understand the basics of the subject). This will just multiply now with AI, and doesn't change just because the prompter happens to be on OpenAI payroll.
We'll see what the final slop rate is, but the three papers they retracted yesterday were for a trivial sign error. If they didn't catch that, that means OpenAI isn't bothered to put in the minimum effort of sifting through their own garbage and making sense of it.
It's not like they needed to hurry out this release before carefully vetting. They're just "hacking" the math system and disrupting the work of thousands of researchers to create a gigantic RL dataset for themselves.
Really, what's the bloody hurry?
They could have released 1-5 papers, worked with researchers to understand what methods work and what don't, how to prompt the models better, how to build better guardrails for reasoning, etc. And give those researchers access to latest models and empower then to solve thousands of problems!
Instead OpenAI wants to piss all over the city to claim territory and now human mathematicians have to go around cleaning up that slop, only so that OpenAI can made some bullshit statement like: math is solved [mistakes are next].
--
Imagine someone gave you a million line PR claiming to have vibecoded the operating system of the future (or whatever your application domain). Would you drop all your other work to focus on this? And they generate enough PR that your manager and company leadership and public all start pressing you to accept it quickly? Guess what, it's your lucky day! You have not one, but 700 breakthrough PRs!
Imagine someone gave you a million line PR claiming to have vibecoded the operating system of the future (or whatever your application domain). Would you drop all your other work to focus on this?
If it verifiably works and solves important outstanding issues, then quite possibly yes.
And if three of their 700 PRs were retracted within a day because of unresolvable bugs? Clearly it's not all verified.
Not to mention the math paper (on HN yesterday) which pointed out that OpenAI might have verified the wrong thing, in their Navier Stokes proof.
It's really not as cut and dried as you (and software/AI folks more generally) think it is.
PS: If a human mathematician had to retract three of their papers a day after posting publicly, they'd lose all credibility and their mathematical career would be all but finished. That social incentive structure is the field's immune system against slop. You're basically asking them to turn off their immune system, and for unclear gains (other than OpenAI's grandstanding).
If a human mathematician had to retract three of their papers a day after posting publicly, they'd lose all credibility and their mathematical career would be all but finished.
If a human published 700 papers and only 3 (or 30) ended up having significant errors, I'd call that a pretty good batting average, considering that the typical rate of errors may be around a third (https://lamport.azurewebsites.net/pubs/statistics.pdf). But for some reason people hold AI output to an absurd standard where if it's not 100% perfect then it's useless.
You're basically asking them to turn off their immune system
I'm not asking "them" to do anything. They can do whatever they want, including rejecting obviously useful tools. But then they shouldn't be surprised when they're quickly surpassed by others who don't share their ideological blinders.
A lot of the comments are claiming that "no one is forcing them to engage with AI proofs" and that's not the case, as explained in the article. The author is forced to engage with the public by the very nature of being a prominent researcher on this problem. The public is drowning him in messages regarding this result. So yes, he is being forced.
Academics have always been required to engage with hacks and cranks to some extent; the deluge of AI proof writing has only exacerbated the problem.
I am not a native speaker but in my experience it is an accurate use of the word “forced”.
English speakers generally use this word in a very broad sense “and now Netflix is forcing ads on paying users”, “because there was no sink, I was forced to drink the whole thing”. It is only when you are literally describing a crime where this word has this strict meaning you are alluding to.
"not at the Lean code, since I know very little of the actual usage of Lean, and that code was enormous"
This part I don't understand. Not that anyone should read the entire Lean code of any proof, but if the statement of the theorem to be proven in lean seems to be correct, then I would think there would be at least some interest if in fact there was a formal proof (which might or might not correspond to the written proof) of something I was working on. That to me would be interesting. Or you are saying you doubt the validity of the formal proof, which would also be interesting. But saying it is of no consequence doesn't make any sense to me.
I think TFA’s point is that it’s interesting - it’s just not feasible to do what follows after “it’s interesting”, which is to try to make heads or tails of the stack of writing that we’ve been given. Engaging with a well-written proof of a similar scope is enough of a task already.
Some of the lean proofs are apparently incomprehensible.
It would be like trying to look at a completed video game's assembly code, being told that it was call of duty, and then being asked questions about the high level code architecture.
AI models are perhaps unsurprisingly good at low level translation (see the progress being made for decomp games)
These models have surpassed human capabilities at math/machine code, but they can't "simplify" yet - in part because they don't have the same need to due to their comparative lack of cognitive constraints. AI Slop code is getting better, but it takes time. At the moment, its embarrassing frankly. It will come eventually, but right now OpenAI is not handling this with the care, respect, or concern that it deserves.
Have you ever wrote some code/algo that seemed "simple/obvious" to you yet to someone else, it seemed incomprehensible?
If you have a 20-40 IQ points gap with another developer, this happens a lot.
The baseline of "simplify" is wildly different based on your IQ points. That's precisely why exceptional students are usually bad in teaching. They try to break things down, simplify, but things still go over the head of normies.
However, we can intervene/train the models. So it should be possible to focus on the simplification, and as you said, it will come eventually.
I've been mostly reading here but created an account to disagree with this statement: The smartest people I know were always amazing at explaining. This was true for me as undergrad and graduate student where the smartest peers and the most renowned professor were always also the best at explaining, and it is true now at research level in a related field.
When I am not sure if I really understood something to the core, I try to find a colleague who knows very little about it; if they understand my explanation well, that's a good sign.
Those who really understand a topic are usually also able (and great at) explaining it in very clear and "simple" terms. This may be part of my personal bias; I see theory builders as those who advance the field the most, and these are usually also amazing at explaining it. On the other hand, those who mostly "grind" through problems (approach them as complicated puzzles) with effort/time were often bad at explaining.
I observed the same for programming: the "architects" usually explain very well, the "debuggers" often don't.
LLMs very much remind of the grind/puzzle approach. It does makes sense that RLVR, which in my understanding enables a lot of these results, would lead to a more mechanical approach.
Of course, I can't make any predictions on whether it will stay that way. But I strongly suspect that we need different ways of training for LLMs to write better text and explain better (I suspect the vagueness of LLM language is the result of RLHF as vague expression is less often incorrect).
I don't disagree with you. However, you are ignoring one key point that I made: the difference of base intelligence.
Let's assume you have an IQ of 135. You are not that far from even the crazy smarts at 150 and as long as they have some social skills, it will work out. I, myself, never had issue with understanding the smart people.
Now let's assume you have an IQ of 110. Unless 140/150 IQ people are spending inordinate amount of time formulating how to simplify things for you, things will fly above your head.
This does not mean just because someone is 150, and someone is 110, the 110 is always the problem.
This also does not mean an 120 explaining something to 140, and failing is 120's fault or the other ones'.
Every combination is possible as we are talking about personalities/ability to articulate oneself.
However, I was talking about a specific scenario, which I have faced myself and saw it happen a lot of times (to other people), where someone's simplification level is still higher than someone's max understanding level simply because the deviation between their caliber is too great. In most cases, this can be worked around by the explainer spending inordinate amount of time.
I agree, and I disagree. There are plenty of published mathematical papers that are just as poorly written as OpenAI's. Nobody says nothing because the authors are big names. In some cases, the proofs are not even correct, but everybody has a feeling the result are true nonetheless, so they pretend not to see it.
So I agree that OpenAI should have done a better job of writing down the results, probably by paying working mathematicians like Anthropic did.
But I disagree that this low-quality writing is somehow a good reason to be angry at OpenAI specifically, otherwise you would have to be angry at a lot of people.
> There are plenty of published mathematical papers that are just as poorly written as OpenAI's. Nobody says nothing because the authors are big names.
Can you provide some evidence of this claim? "Nobody says nothing" probably works on reddit but I generally expect higher quality discourse on hackernews.
I am a working mathematician. A problem that I cared about greatly (and probably spent > 3000 hours working on) was on their list. I looked at the paper, and I have to say it is more clearly written than about 30% of the papers I typically referee. I don't want to name poorly written papers, but I agree that "there are plenty of published papers that are just as poorly written as OpenAI's".
And which papers do you typically referee? Without that information this claim is meaningless. For example, if you referee free-for-all papers that today are likely written by LLMs as well then sure I can understand that. But if you referee papers from grad students then that's more concerning.
I have only twice (knowingly) refereed AI slop. I'm mostly talking about refereeing in the period 2010 - 2020. (However, looking here https://proofsandprompts.com/2026/10/08/100-reactions-to-100... it seems that many other consider many of the papers poorly written. I just wanted to give one datapoint.)
The response by some in the field of mathematics to this repo is ... I guess not unexpected; but it's quite disappointing.
I sympathize with those who've worked on some problem for years and now don't have something to work on; it's been a part of their identity. I also especially sympathize with those whose career tracks and plans were thrown in disarray.
That being said, I absolutely cannot understand how one can't be excited and happy and enthused about these advances in one's field. Assuming just that the ones with formal lean proofs are actually true, these are reportedly huge advances. Even if folks don't understand it YET.
It is because in fact MOST mathematicians (and researchers) don't really just want "advances" in their field. They also want to have some significance and play a role in it. What's the point of doing research, which takes a lot of work for much less pay than industry, except to have some of the glory?
I didn't go into math so I could read the results of others all day. I went into it to contribute meaningfully, and I certainly don't consider interpreting the results of a machine to be a meaningful contribution.
Your perspective is just the perspective of a consumer of things. In that case, it doesn't matter where they come from.
> But, back to the Partition Principle. I took a brief look at the preprint released by OpenAI (not at the Lean code, since I know very little of the actual usage of Lean, and that code was enormous). It sucked. It is unclear, muddled, and has a strange structure.
Incredible lack of curiosity. The Lean artifact shows that there is a proof. Maybe the natural language writeup sucks (maybe it doesn't even correspond to the Lean proof!) but the proof is there and if he were really interested in the problem he would try to understand it. Rather, his revealed preference is that what he's really interested in is good style in academic papers.
the OP's point is that if another academic had submitted that paper to a journal with their Lean code, the paper would likely have been rejected by reviewers regardless of the Lean code
it's just holding OpenAI to the same standard as everyone else
> The Lean artifact shows that there is a proof.
how are you sure, if it can't be explained properly?
update: to me this feels like the equiv of dumping an enormous PR that probably has some great stuff in the code but is poorly explained and documented, and expect the maintainer to try to make sense of it and see if it's valid merge
Maybe on your end because that's the exact point the author explained - these companies need to adhere to the scientific community rather than expecting the opposite.
If the write up - the natural language part - sucks - why are you expecting humans to waste their time on understanding the proof???
Because there is a proof! If you are interested in the conjecture, then you would be interested in the proof. If you are not interested in the proof, then you weren't really interested in the conjecture. You were interested in the good style of articles about the conjecture, or the friends you made along the way, etc.
It is of course possible that either (1) the statement of the theorem in Lean is busted or (2) there is a bug in Lean. But both of these seem to me to be lower probability than that the proof is correct. It's not just an e-mail from a crackpot. For someone who is actually interested in the problem, the probability that the proof is correct is high enough to warrant effort to understand it, or at least to learn enough Lean to check the theorem statement.
>these companies need to adhere to the scientific community rather than expecting the opposite
Why? The pre-existing community doesn't own science.
Considering how hostile the community sometimes is to outsiders, how they frequently demand form over function and think connections are often more important than correctness of argument, perhaps it's good for them to be confronted with a new approach to science that does away with those things and returns to the real cornerstones: proof and empiricism.
If you don't want to engage with proofs that's your call, many of us are happy to see progress being made and don't need you specifically for the confirmation
At first I thought this was going to be more Luddite babble, but it makes a good point. OpenAI isn't contributing if they are make unreadable papers. They should use a little more of their compute to nail interpretability. The difficulty will only get worse as AI plow deeper into the frontier and produce increasingly alien looking output. I suspect it's a workflow issue. If not, it's a bad oversight if the current generation of models are capable of making mathematical breakthroughs but can't explain how they build on existing frameworks.
Work this abstract almost certainly has no value outside the community that is (was) interested in the result. OpenAI should engage with the community to realize the value (beyond PR).
I can't 'ignore' OpenAI's paper any more than you can verify it unless you are an unusual person with a very uncommon skill set. Very few people with the skill set to look at these are getting slammed with difficult to read artifacts, which is confusing because they have internal AIs so smart that they can solve hard math problems.
Is it just me or is this line of thinking fundamentally dishonest? It may be true that the papers are hard to read but the papers have extremely high signal towards the proof.
All of these comments seem to suggest that if papers are not 100% abiding by readable books their standards, they are basically the same as literal noise. This is highly dishonest.
I’m just a dude and even I’m able to understand the paper after using ChatGPT to help me through it.
There's certainly a conceit in saying if the math is not accessible to my community then it doesn't even count as math, and the last few days this has been a common ideological talking point.
OpenAI is publishing the math because they want mathematicians to read along and verify the work they are doing. The author complained that even the citations don't make a lot of sense. This is completely consistent with my experience of using AI for coding. When I task my agent a hard problem and it comes back with a 5000 line PR that I don't follow I reject it and work it into a better shape. I certainly don't slam it into production because it has a 'high signal' of the program I asked for. The mathematician's situation is a lot worse because they have no view over the process that created the artifact like we do with coding agents.
> When I task my agent a hard problem and it comes back with a 5000 line PR that I don't follow I reject it and work it into a better shape. I certainly don't slam it into production because it has a 'high signal' of the program I asked for.
??? That's literally what everyone does. You are creating a hypothetical that is nonsensical. Here's your hypothetical:
1. Agent gives you 5000 lines of slop
2. You reject it and just do it yourself
This is reality
1. Agent gives you 5000 lines of slop
2. you realise that it has done a lot of research and is mostly in the correct direction and you ask it nicely to refine it
3. verify that you understood it and push it to prod
Are mathematicians babies that they need a completely different approach?
The mathematician in this case didn't ask an AI for the proof, they were handed a badly written artifact and other people are messaging him asking for analysis. He has no visibility into the process that produced the paper because it came from an internal OpenAI model. The bad behavior here isn't coming from the complaining mathematician. OpenAI hasn't sufficiently developed their workflow to create cutting edge research AND publish it intelligibly, which is odd because the latter should be easier. OpenAI has perverse incentive to do this because spamming badly written papers can still give them priority credit for solving open problems at the expense of mathematicians (who are currently indispensable because they decide whether claims are credible). What Mr. Karagila is trying to enforce are the old standards that review should come after refined explanation of the solution. If that was the case before AI, I can't see a reality in which that isn't the standard now that papers can be pumped out at faster-than-human time scales.
It doesn't seem like the paper is really hard to read, but rather that the terminology is "a bit off", with "unexpected theorems" and references that are correct but that could point to more useful targets (read: papers that will increase the citation count of the researcher being cites...). None of which are valid reasons to be dismissive of a good result.
> I’m able to understand the paper after using ChatGPT to help me through it.
How could you possibly know that? If you don't have the knowledge to understand the paper without using a chatbot, you don't have the knowledge to verify that what the chatbot said is the same thing as what's in the paper.
Maybe the reason the AI could find this proof is exactly what the author is complaining about: that it left the beaten path of theorems expected in a paper like this and went off in an unexpected direction.
I'm just amazed at how quickly we moved from "AI is just a parrot and can't do anything useful" to "AI can solve toy problems but not anything of value" to "when AI pushes the state of the art, it can't quite get the proper citations in its preprints".
Except that the proofs are very likely valid. This is a country of Ramanujans in a data center. You're free to ignore them because they don't follow your style guide, the rest of us will enjoy seeing humans and AIs build on the results.
> the rest of us will enjoy seeing humans and AIs build on the results.
That would be nice, but the rest of the world wasn't interested in the results before and they won't be interested after.
I wonder what this means long term. Maybe mathematicians will keep plodding on as usual except sporadically when an AI company need a marketing boost so they spend millions of dollars to dunk on them. Because the mathematicians sure don't have that kind of money.
the rest of the world wasn't interested in the results before and they won't be interested after
I can't speak for the rest of the world, but I think it's amazing that multiplication and 3SUM are sub-quadratic despite that being "obviously" impossible.
I wonder what this means long term.
Long term, stronger models than what OpenAI used will be available to everyone. Some mathematicians will take advantage of those tools and do great things; others will continue their whiny gatekeeping and become irrelevant.
Mathematicians struggle with the same problem as software engineers; you can let AI generate the artifact, but to understand fully what is going on is challenging. Perhaps even more for mathematicians.
Do you rely on the tests/Lean to accept correctness or not…
I would have expected the math community to celebrate all of this since they got into math for the love for mathematics rather than the love for tenure. Their own world will not end - more mathematics will require more mathematicians to make sense of it all. And if all of it leads to some form of superabundance they are looking at the best possible future: they’ll be able to do math forever without having to worry to get paid for it.
> they got into math for the love for mathematics rather than the love for tenure
They got into math for the love for mathematics but if they couldn't make a living from it (tenure) then they would have done something else just like everybody else who is not a starving artist. So, no, they won't celebrate it and neither would you if you were honest. (note also that starving has a short time limit before you die so "if all of it leads to some form of superabundance" won't work in that circumstance)
> Their own world will not end - more mathematics will require more mathematicians to make sense of it all.
Someone tell the administration to restart funding mathematicians then [1] -- because the exact opposite of "more mathematicians" is happening right now, both at the post/graduate level [2], as well as the undergraduate level (for much longer) [3].
>OpenAI drops some hundreds of "solutions", incomprehensibly written "solutions", and we are all expected to jump on them and what? Appreciate their contributions? Why?
...
>So, no, I will not be sending Sam Altman a bottle of whisky anytime soon, nor I am planning on spending my time reading through that paper and trying to make sense of it.
Think about a hypothetical circumstance where we get radio communication with some aliens on another planet. They send over tons of math to help us advance our tech, we know the math they're sending us is correct, but their explanations are really hard to work through because they aren't humans and the math is so different from anything we've done. Should we whine about the results they sent to us and refuse to engage with it?
In this hypothetical case funding agencies would probably happily pay for “alien maths translation grants” and you could write papers about advances in that field and get jobs etc. So arguably for better or worse it would create a specialised academic cottage industry that somehow interfaces with the rest of maths.
Love your metaphor here. I guess there will always be people who try to keep going and follow their current projects and say that only human math is real math and we should only use alien technology for proofreading mails and such but not their highly established and esteemed professions.
Or just keep your hubris low, update your priors, work towards progress.
As a whole, the math community would want to simply ignore the distractions from AI companies … they are using math problems as medium to promote their commercial interests.
> not at the Lean code, since I know very little of the actual usage of Lean, and that code was enormous
Dismissing results on the basis that Lean code is too long disqualifies this opinion. It is not hard at all to read the Lean result statement, even with very superficial Lean knowledge.
There is no nice way to tell someone that you’ve scooped them, and this is industrial scale scooping.
A few papers have been retracted, but it looks like many are withstanding intense scrutiny. Lean is making the results more likely to be correct, but I think making them harder to understand.
The world has changed and you’ll know a math department is making a serious attempt to adapt when it teaches a required Lean course in freshman year.
Why do you believe that learning Lean is a better use of time when the AI is clearly better at writing and interpreting Lean than it is writing quality papers? At this point, Lean is for autoformalization, no one is really supposed to read it.
That's fair, but I would argue that reading the specification is very easy by comparison. A quick one hour tutorial is usually enough judging from my students' experiences.
Indeed! Almost none of these most math folks are likely to read. The only thing to read is the statement, which is only a handful of lines in both cases. The point of Lean is that if that statement compiles and is validated by hand to be equivalent to the natural language statement, then it is true. That was what I was trying to say here.
> This is why I generally avoid using AI for mathematics (I am happy to ask LLMs to consolidate information for me, or to generate a useful infographic, or to proof read an email, etc.)
In other words, the author is OK with using LLMs to replace data analysts (that could consolidate information), to replace graphic designers (that could generate infographics), and to replace editors (that could proofread an email). But don't you dare use LLMs in their mathematics.
I don't know about this person, but for me and I'm guessing the average professional in basically any field, before LLMs, they simply ignored information that needed summarizing, made their own crappy slides, and didn't proofread their emails.
People seem to be very likely to say that LLMs are good at things that they don't want to do. That's also why you see so much of "well of course they can't do x, but they're great at y" but everyone has a different x and y.
openai will take expert responses like this and improve the next set of papers
it won't be long before there's no more low hanging fruit like this to complain about, and the writing / explanations of the results are superhuman as well
separately, i really liked the author's denial-of-service analogy. super useful practical framing
>That's not what you'd expect from a serious preprint claiming to solve a problem.
...
>It seems to me, that the AI tech companies would like us to conform to their standards, rather than spend the time and energy to conform to ours.
The entire argument is ugly and unfamiliar methods are used, and nobody's got time to drop everything and understand all that, meanwhile OpenAI didn't even ask us about this, therefore we should reject the paper as "unreadable" and you should be "very angry at OpenAI" today.
No, it's the press where you need to direct your anger. That is the nature of the beast. And being arrogant at the time of mass-disruption is a great way to lose control of everything. Maybe Sam should be sending the bottle of whisky to you.
I think the problem is that they mixed problems that some solutions in Lean (and I consider they solved, assuming someone checked the formalization of the statement) with some solution in English.
I expect most of the solutions in English to be correct or mostly correct, but we all have made and read mistakes mixed with the bla bla bla. In the Mythbuster scale I'd classify them as "plausible" instead of "confirmed".
Very well written summary of the current situation
Imagine being someone who is working on one of these problems. You have no good guarantee that the problem was solved, but you will have the horrible homework of reading the AI slop. Also, if you do have something interesting to say about the problem, people will have less enthusiasm about it now
The gradient into formalisms is steep and the bottom is deep. These are the worst results anyone will ever see again. The bottom of the internet is also deep, but won't be around much longer.
I am no mathematician, but I can definitely relate to the "horrible homework of reading AI slop". We've started allowing a non-developer to submit AI generated code into our codebase (in droves).
When I am tasked with reviewing that code. First, I don't know if it's valid. The person who wrote it doesn't know if it's valid. In order to validate it I must step back and understand the full problem space. Then, when I ask for revisions or clarifications, it's seen as either
A - Slowing progress, being resistant to change... or...
B - Thanks for catching that (Claude fix PR 532 with the review comments)
It's 100% removed the enthusiasm.
Perhaps, there is a point where we just "give up" the understanding and accept AI output as the ground truth, because the sheer amount of generation is too much for our puny human minds to comprehend, and a lot of the times it IS right, even if a little wonky.
"Er, can you check this for free?
...
It would be great for my share price if you could, would really give the investors a badly needed shot of confidence!"
Here’s an idea I would love to see play out. Have one of the labs train a new model, using cutting edge architecture, on a whole lot of data and math papers from before 1905. Then see if it can come up with, or even understand, Einsteins theory of relativity.
Special relativity is one of those unusual theories that requires very little in the way of math or concepts to understand. You take the data from the 1887 Michelson-Morley experiment (which measures the same speed of light, despite different reference frames). You take the idea that the laws of physics work in any reference frame. You take some high-school math, et voilà, special relativity.
If there is a formal proof, and it is the proof of your precise statement (which I imagine is easy to check, otherwise I think that mathematicians would not accept the Navier Stokes result so quickly), then there is no way to ignore the result, however badly it is written.
This was always the essence of mathematics, and it will stay this way whichever statement by whomever is made.
As much as I personally despise altmans, "darios", and their bootlickers, this is one aspect which is undoubtedly "good for the mathematical community" as a whole.
The fact that the validity of your statement does not depend any more on an expert opinion of some person with grants, but as it always should have had been, just on the validity of the chain of deductions.
It certainly feels that we are witnessing the early days of LLMs and math, similar to the early days of LLMs and code. I feel this entire post sounds so similar to angry computer engineers posts from 1 or 2 years ago. If we continue on this path, eventually many of these posts will ripen like milk in a desert.
It’s fine to dislike AI slop and not engage with it. But I would claim there is a difference between a human submitting a sloppy paper to a journal vs producing one with AI. The former is lazy and unprofessional, the latter is an interesting experiment. I appreciate not wanting to engage in an experiment you didn’t sign up for, but therein lies the difference between this with an attitude of embracing new technology and those wanting to stick to the old. I think Terrence Tao has had a very interesting attitude to AI recently and had also had some fruitful outcomes from it.
I would admit seeing more mathematicians irritated is a good popcorn show.
For what's worth it, this batch of results are not very tight intentionally by OpenAI and promising mathematicians already started to consume them and improve the results, while someone is still complaining on it.
> OpenAI drops some hundreds of "solutions", incomprehensibly written "solutions", and we are all expected to jump on them and what? Appreciate their contributions? Why? I am trying to finish several papers, I am supervising a number of Ph.D. students, and I have a lot of active research of my own to do. When am I supposed to sift through a badly written paper? Why should I bother, when they don't bother to communicate better?
The mathematicians who don't approve of the deliverables should boycott the proofs, that is the only way OpenAI will get what they deserve on this one.
I see an argument about the paper being poorly written, which I believe, but I don’t see how that relates to honest final paragraph about mathematicians being chefs or whatever.
It’s clear from the author’s tone about lean, emails, infographics, etc. that he thinks automating those away is fine. Why should math be any different?
Is a proof that cannot be understood worthless? How would this be framed philosophically?
In fact, the academic system is a kind of worldview created by humans. And as it is shared and the community grows, the problem will gradually become more complex. Because when a discipline develops sufficiently, just as in a mine where rich veins are easy to extract early on but become very hard to extract once much has been dug out... in that sense, as things gradually become more complex, once a certain threshold is reached, won't scholarship surpass the limits of human understanding? Of course, scholarship is entirely for humans, but at some point the system itself may face its limits, and then wouldn't it again reduce the existing normalized minimum within that discipline and establish a new normalization of a new logical system?
In my view, perhaps for very complex work like today, AI will do it, and then there will be work that normalizes and further simplifies the results of that AI. Then, coming back to the human fold, if humans create the initial skeleton, the LLM will learn that again and it will become complex work again, and won't this create a continuing cycle?
I think verification and understanding can be separated. If the proof targets a correctly formalized proposition and passes a reliable proof checker, isn't it valuable? We have obtained knowledge justified as true, but there is simply no new theory that understands that knowledge.
As was the case with the Four Color Theorem...
I am always curious what shape the newly compressed new discipline will take. At that time, I hope even people like me, who are intellectually behind, will be able to learn that discipline.
I've seen my entire profession vanish overnight due to AI... well, not vanish. But, yeah, software development is WAYYYY different. And I couldn't be happier. I think it's amazing. I see the productivity boost. Even if it means I can't add nearly as much value as I used to.
Sorry, but I don't get why mathematicians are so upset. Like, just accept the knowledge and insights and acceleration in your field! If it isn't "fit for human consumption" because an AI produced, okay... it soon will be explained ELI5 by even better models.
Most software developers (or people in any field, working for anyone) rarely own anything they do at work.
And the entrepreneurs running their own companies (which there's an explosion of atm largely because of AI) do indeed "own" the higher-level products and things they're producing, even if AI writes the code.
Did you... Enjoy software development beforehand? Like there is a huge chunk of people that really enjoyed writing software and are bummed that that part of their job is being subsumed by models.
The same could go for mathematics or any field; there are lots of people who enjoy the process and aren't satisfied by being handed and opaque final result
I don’t get why you aren’t upset. In general I am actually quite disappointed by how meekly the software engineers handed our industry over to AI. But I am particularly revolted by those who exercise their free will to rise from the foetid swamp and say “come on in, the water’s great!”.
> OpenAI drops some hundreds of "solutions", incomprehensibly written "solutions", and we are all expected to jump on them and what? Appreciate their contributions?
This is a strawman. OpenAI didn't say they are expecting all mathematicians to read the solutions, incomprehensible or not.
So mathematicians are upset with OpenAI for solving "their" math problems. Software engineers are even more affected by AI, yet mathematicians seem to be reacting more strongly. I don't get why.
It looks like the spent $20 on writing the actual papers.
If they actually wanted to do good for the world, they wouldn't have released these as the slop grenades they are.
In their current state, they are actively damaging the mathematics community.
It shows a lack of respect and care for the impact that their technology has.
It shows that they cannot be trusted for things like private data, AI safety, and company partnerships.
In math/science, repeatability and review are critical to the process.
The right way to handle this would have been to work with the mathematics community to co-develop and create meaningful proofs rather than slop grenades.
If they proceed in the current state, we'll just get a bunch of spaghetti math that won't do anything for helping people build an understanding.
Maybe some day, we won't need people to understand things, but that's certainly not the case at the moment, and likely won't be for several more years.
Perhaps this is projection, and the staff at OpenAI doesn't understand their work anymore? Not a great sign regardless.
> The right way to handle this would have been to work with the mathematics community to co-develop and create meaningful proofs rather than slop grenades.
The mathematics community can finish the job OpenAI started. Or are you saying the community has no incentive to do that because there is no reward/recognition for doing that?
They didn't just "start" it though - they released slop papers.
That's finishing it - not starting it as far as scientific publishing is concerned.
> Or are you saying the community has no incentive to do that because there is no reward/recognition for doing that?
Not exactly - but that is part of it.
I think what would have went over better is:
1. Immediately announce a solution has been found.
2. Do not publish the solution.
3. Put out an open request for anyone with experience in the area who wants to get involved to help collaborate on a construction and human-comprehensible paper. Accept anyone who can demonstrate potentially useful work/experience in the field/problem. Share the solution with them after they sign some kind of NDA that they won't independently publish or share the solution/work.
4. Work with people until a paper is ready (I mean actually ready - not the kind of slop that they released).
5. Publish. Include names of everyone who made meaningful contributions to the paper (not just the proof).
EDIT: Notice the incentive with my proposed second path is that it gives OpenAI an incentive to improve the interpretability of its proofs. This is a good thing! The maths community would be thrilled to actually gain understanding from such releases, and OpenAI would be happy because they could more quickly and independently publish their results. At the moment the "value" of their mathematics research "product" is low because of the lack of this interpretability, and this current approach is simultaneously destroying the opportunity value of the community as well as the incentive for OpenAI to ever improve on what's missing.
OpenAI is doing science the right way. Making the information available as widely as possible so that anyone can check and verify it.
That is the scientific process working exactly as it should.
If that "damages the mathematics community" then all it means is the mathematics community is not doing science and should be ignored.
I cannot stress this enough. If you are complaining about "how they released it" or calling this stuff "slop cannons", you're an unscientific hack that's dragging down humanity.
Engage with the actual claims. Prove them or disprove them. Nothing else matters here.
This is great - I think we're getting close to the core disagreement here because I disagree with almost all of that, and I think its because we have different conceptions of what science is.
To me, science is not a database of facts and data, and the scientific process is not just the process of adding to this database.
It is more than that.
Science is building the database of facts and data, but it is also sharing/disseminating that database so that others can write to it too.
We both agree that "openness" is critical for this, but to me, the scientific process includes things like the actual words that a scientist uses to communicate their ideas/data, the paper they wrote, the collaborations they had, the work that they built upon. Science is not like a database only - its more like a distributed system, and the scientific community is like a living, breathing thing. It is the health of that system which has been damaged in my opinion.
So I have no problem that OpenAI updated the database.
They didn't just make the information available though.
They updated the knowledge database, then released low quality slop papers.
They made it clear that they don't give a damn about sharing/disseminating information.
They are trying to fundamentally say "we can do science without prioritizing sharing.
We don't think its important. We'll release the proof because it makes our stock go up, but we're gonna keep going and building and not spend any time on communication or comprehension because we think that it doesn't matter, and your contributions don't matter, and AI is gonna do all of this independently anyway. We spent $20m on the proof because it made our stock go up $20b, but we'll spend 20 min on the paper. Priorities clear.
I don't disagree that having comprehensible papers is useful. There is undoubtedly value in having well written, easy to understand documents that make it easier for third parties to validate the work presented.
The thing is, that is a "nice to have", not a condition to make science.
If people believe that such things are useful, then the appropriate response is to reproduce, simplify and otherwise make the proofs provided into more easily digestible pieces.
The scientific process is fundamentally indifferent to whatever shibollets and cultural hangups that its presumed participants may have. An alien with fundamentally different norms from us, who experiments to determine the boiling temperature of water is doing science just as much as a human, even if he never submits a paper to a journal.
The issue you're having is that you're under the mistaken belief that the scientific process is the "community" and "norms" that happen to have developed over the past few centuries by this "scientific" priesthood.
It is not. It has never been. It will never be.
This "scientific" priesthood has simply used their power and reputation to enforce their own self-aggrandizing norms, much like the religious priesthood they replaced. They should not be taken seriously, and to go along with these norms when they are unnecessary is frankly ridiculous and shameful for anyone who actually prizes results.
I would urge you to reconsider your stance, and in particular to put deep thought into why you believe these norms are so important.
> Indeed, at least two people asked if I plan on sending Sam Altman a bottle of whisky, as promised in my Problems page. The answer to that is no. And let me explain to you why
> I took a brief look at the preprint released by OpenAI. It sucked. It is unclear, muddled, and has a strange structure.
If OpenAI actually spent time writing good proofs that are readable, mathematicians would show *even more outrage*. So let’s not pretend this has anything to do with readability. It has all to do with people protecting their status in society.
> if this was an academic paper submitted to a journal, it should be issued a desk rejection for the quality
If Mr Tao had submitted a shitty, poorly-written proof of a famous outstanding problem, no journal would reject it. That extends to anyone with sufficient credibility. They might ask him to keep at it and fix it up, but nobody would begrudge him putting his shitty (but ultimately correct) draft of arXiv while he did so.
We're all getting disrupted, we all have feelings about it, but from the perspective of a software engineer who's been dealing with all of this for several years now, this post is just cope.
This whole "oh no this problem is solved now who would ever want to work on it" thing is quite funny. Mathematicians, let me introduce you to something called bike shedding. Folks proved 1s and 0s are turing complete decades ago and yet we have 27 new JavaScript web frameworks every week (or every day or minute now with ai). Y'all will be fine. Thinking a solution is ugly and that you could do better is also a perfectly good motivator and most of the sciences and engineering are "ok, so we know xyz to be true about the world because like, I'm looking at it, but wtf is going on." The theorems were true false or otherwise before some random openai model solved them. And if you don't understand the proof nothing of significance has changed except you've got a bit of a hint now.
Cathedral and Bazaar. Maths professors are used to working diligently behind closed doors before releasing artifacts of high quality whose authorship they guard jealously. The researchers at OpenAI, who come from a software background, are used to working in the open, releasing anything to anyone and expecting nothing but also guaranteeing nothing.
The amount of time the OP attacks OpenAI for errors of form and not substance is unfortunate. Do they also attack amateurs who try to contribute like this?
I think it can both true that 1) OpenAI is being inconsiderate/harmful/<pick-whatever-adjective> with their math releases, and 2) there is now a treasure trove of mathematical results ready for the taking.
Yes, the situation sucks overall and mathematics as a whole is in a turbulent time now.
But it also sucks when mathematicians, who are considered experts on a particular problem, refuse to engage with breakthrough results about that problem. #2 above is still true regardless of where it came from or how hard it can be to absorb.
While reading "The Mathocalypse" post [0] by Scott Aaronson, Scott described his wife Dana's reaction to one of the newly solved results in her primary domain of expertise, on which she'd been working for decades.
After her initial shock, and annoyance with the format/style, she decided to start using Astra - for the first time - to help her understand the new result. And he reported in the comments that she had made a lot of progress understanding it in one day, and may be even excited to give a talk about it!
That seems like a much healthier attitude towards these new results.
Yes, everything else sucks about this messy period. But there are still diamonds (in the rough) in this drop that perhaps should be looked into. If the author is too busy, perhaps one of their students can take a look? Someone will, eventually.
[0] https://scottaaronson.blog/?p=10169
Why are you two-siding this. There is only one misaligned agent here and that is the company OpenAI. What OpenAI is doing sucks, and people are calling OpenAI out for sucking. However mathematicians (the main victims of OpenAI’s lousy behavior) behave is not the issue here.
If some mathematicians complain about OpenAI sucking, that is fine actually, and if others are more “mature” about it, that that is fine too. Neither of these reactions should be at put as an equivalence to the blame OpenAI deserves for this stunt.
doing math is not "misaligned"
write down "LLM's will cure cancer" on sticky note, put it on your desk and recite that every day, a hundred times and in the bathroom if you have to
you're really crashing out about this huh
I know I am, though maybe not as visibly so as your parent. Us anti-AI Luddites are arguing against a massive propaganda machine with trillions of dollars on the line. It is not easy, and it is mentally draining. Especially since AI-companies are behaving in an obviously anti-human, anti-science, and anti-industry ways for their own short term profits.
And worse yet, we have seen this before (albeit on a lesser scale). GMO was supposed to solve global hunger, Carbon capture was supposed to solve climate change, etc. etc. These were obvious lies and propaganda back then as much as the lies where AI are supposed to help humanity today. It gets tiring, and at some point we simply crash out.
unfortunately technology will never fix humanity's tendency to ruin itself (see: war, climate change, etc; all of which we know the solutions for ie 'stop killing people' 'stop using oil' etc) but that is not a fault of the technology itself (gmo, carbon capture, crispr, llms, etc)
This is simply not true. Plenty of technology exists which has helped humanity unambiguously. But there is a difference between good and helpful technology and a grift. GMO could have been helpful but it was mostly just a grift, GMO solving world hunger was always an obvious marketing stunt, and it worked, even though many of us (particularly on the political left) saw it as such and called it out as such. Ditto Carbon capture.
AI is currently the grift of choice.
Also I fundamentally disagree that there is any tendency for humanity to ruin it self. There are perverse intensives, which rewards undesirable behavior, there are unfair, undemocratic, or otherwise hostile powerstructures which allows one class of peoples to exploit another class of peoples. And some technology sometimes plays a pivotal role in maintaining or furthering these hostile power structures/perverse intensives.
Are they fully “doing math” in an aligned way though? They are finding mathematical proofs, which is important, but contextualizing results in the prior literature and clearly communicating the approach and implications is just as much part of mathematical research.
You might argue that these aspects of math are less important in the new AI accelerated math world, because agents will inevitably be smarter than humans, but I think clear framing and communication is even more important than before because with this technology we can choose to augment our intelligence instead of defer it
> Are they fully “doing math” in an aligned way though?
No matter what you think "aligned" means here, publicly releasing advancements in a scientific field shouldn't be gatekept or be perceived as misaligned in any possible way.
If anything, the true misalignment comes from people trying to prevent these advancements from happening or being disclosed, or putting research behind BS paywalls.
I don‘t think this is gatekeeping. OpenAI is filthy rich, they can afford to pay mathematicians to go over these papers and produce results consistent with the scientific method. The fact they don‘t is evidence of their misalignment. Their motives are obviously ulterior and have nothing to do with advancing knowledge. If they were they would adhere to the scientific method, hire experts, and produce reproducible results.
> Putting research behind BS paywalls.
Yes, that is gatekeeping. But there are more then one way to be misaligned. And the practice of publishers is not being discussed here. No need for whataboutism.
Wait, what has OpenAI done that is not “consistent with the scientific method”?
https://www.sciencebuddies.org/science-fair-projects/science...
Do other researchers or research institutions “pay mathematicians (other than those in their employ) to go over these papers” or do they simply publish their work for review as part of producing “reproducible results”?
AI may be taking over Mathematicians’ jobs, along with everyone else’s, and it’s OK to hate that, but the scientific method says nothing about that, or hiring experts, or ulterior motives, or “being filthy rich.”
Lmao, now solving mathematical problems and releasing the solution to the public is "misaligned". The anti-AI hysteria is getting really funny.
Dredging every inch of the lake to catch every single fish and dumping them all in one big pile the town square is just “solving fishing and releasing the fish to the public”, this anti-trawler hysteria is getting really funny.
Paraphrasing Hardy, 'exposition is for second-rate minds'.
I don't believe this myself. But I do believe that if you've formed your very ideas about what is good and desirable on the basis of a culture that has held certain values dear for hundreds of years, and have fought against every doubt and difficulty in life for decades to mold yourself into that image, that it does not 'suck' that you are unable to adapt to a new reality overnight.
Very few people that would love to be craftsmen would love to be factory foremen. It is far too insensitive to the human experience to expect people to just deal.
> Paraphrasing Hardy, 'exposition is for second-rate minds'.
Forgive me if I have no sympathy for current mathematicians who think this way. It's a pretty ugly kind of arrogance.
Some people told themselves they were the pinnacle, the first-rate mind, as opposed to all the second-rater. Well guess what, now your first-rate mind is a commodity and exposition is more valuable. They'd better learn to live with it.
It turns out your an expendable commodity, and my computer system is predicting that society will profit from your annihilation, really sorry about that and I hope there's no hard feelings.
> Forgive me if I have no sympathy for current mathematicians who think this way. It's a pretty ugly kind of arrogance.
Thankfully, very few mathematicians share Hardy's opinion, just as very few share his opinion that "mathematics is a young man's game" (and indeed we now have prizes like the Abel Prize with no age limit).
In fact, many of the greatest mathematicians throughout history have taken exposition very seriously, e.g. Euclid, Euler, Lagrange, Cauchy, Dirichlet, Kolmogorov etc. all wrote textbooks. Many mathematicians today carry on that tradition of taking exposition seriously and write books and freely share their lecture notes.
So we should not take Hardy's opinion as representing the opinion of all mathematicians or even most mathematicians. In fact, Hardy's statement is somewhat self-contradictory since he himself wrote several expository books (e.g. "A Course of Pure Mathematics").
His statement was self-deprecating (in a not so endearing way), as he was referencing his young self as one of those first-rate minds but his current old self doing exposition as second-rate.
I know that it was self-deprecating, but that doesn't redeem it much to me. I also still find it pretty contradictory/ironic because he wrote A Course of Pure Mathematics when he was in his early thirties.
Is that how he himself felt when trying to figure out proofs by Ramanujan? Nevertheless what a shitty view
>> Some people told themselves they were the pinnacle, the first-rate mind, as opposed to all the second-rater.
Wait, who are those people? Who is that Hardy and what did he really say? What did he really mean? Who else said or meant the same things?
Who are you criticising, exactly?
FALSE.
A first rate mind communes with mathematics not matter terrestrial. An Uber-Erdos entering to lecture is indifferent to the auditorium as well as it's contents. He's at the board silently communicating new mathematics. Not for himself, not for you, but for mathematics.
> But it also sucks when mathematicians, who are considered experts on a particular problem, refuse to engage with breakthrough results about that problem.
It might be a shock for you but they are very few in numbers. Most of researchers I know are always busy with something. They cannot just drop other responsibilities for something like this. They will take their own time getting through the proofs (if they want to).
> That seems like a much healthier attitude towards these new results.
Another thing to consider is not all mathematicians are from US or with good funding. The PI or graduate students cannot afford to pay 200/month.
I think the role of specific _human_ mathematicians at OpenAI should not be understated.
TFA was about a niche topic that OpenAI doesn't have in-house expertise in.
Otoh Aaronson is the co-author on Lijie Chen's (reasoning lead at OAI) top cited paper. OAI have deployed their resources more effectively against UGC that some of their staff are already familiar with
https://scholar.google.com/citations?user=T_OhvOsAAAAJ
https://finance.biggo.com/news/B0eMxZsBy4YEFZDUVPWh
That's true. OpenAI has some of the best talents. In my experience, domain experts get the most benefits from the models. They can work much faster, catch false positives, and stir the model in right direction.
I wish they take a bit of more time to communicate the findings effectively.
> I wish they take a bit of more time
There are good reasons not to delay publishing at all:
> They should release all their results immediately. (Imagine working on one of the problems they already solved.)
This is the most popular answer to a question regarding AI advisory group and immediate access on a popular website for professional mathematicians: https://mathoverflow.net/a/515442/473286
The whole debate regarding the behaviour of OpenAI is a red herring. Mathematics need to redefine their profession and how they work (like us software developers too). There are very good reasons to believe mathematics has an important role to play. If they could just stop talking about OpenAI and get back to work - they are very much needed, in particular now!
> If they could just stop talking about OpenAI and get back to work - they are very much needed, in particular now
This kind of phrasing sounds particularly empty. We are not in WWII researching the nuclear bomb. What are they so urgently needed for to drop everything and work on understanding openai's proof on partition principle and axiom of choice?
Where did I say they should "drop everything"? I hope to read more about how they envision their future, instead of all this regretting and whaling how one company (that I very much dislike too) published a large amount of proofs. They will get more of them, very soon - if they like it or not -, and I wish that would be the primary subject of the discussion. And yes, I'd also hope they engage with the published proofs. The more raw these proofs are, the better. If AI companies start selecting mathematicians to write nice expositions of their proofs, this is doomed to become a very elitist science.
> The more raw these proofs are, the better.
Says who? The professional mathematician writing the article disagreed. Why should people start dancing the tune that openai wants to play for their own reasons and interests? And I do not see how taking the time and effort to write a proper exposition makes it "a very elitist science" when this exact effort and time is needed to actually get other experts understand and build on a result. Unless you equate spending time and effort learning math as "elitism", which is the ai-shilling moto some time now with everything time and effort related. I cannot see how spending time and effort to understand a field and then spend time and effort to make a proper exposition so that other people can also understand it as "elitist" vs throw everything out there "in raw form".
>> The more raw these proofs are, the better.
This was my opinion. But anyone arguing to publish "results immediately" is likely to imply something like it. I guess in chemistry we have the situation you envision - for different reasons: Laboratories holding back their data, until their scientists have published their papers or developed their products. There is a real danger AI companies will do something similar too.
Who do you expect they will select for the exposition?
On which basis do you want a (likely US based) AI company to decide who is to untangle a proof that their latest internal model has just spit out?
Do you expect this to fall to an aspiring, but still unknown mathematician at - say - the mathematics department of Nairobi university?! This is what I meant with my rather unclear "elite" reference: The first publication will always show the name of a mathematician already known to the field, more likely than not to come from the same country as the company ("Our message ... is, you’re a great American company, but you’ve got to hire great American workers"). Are you not worried at all? Don't you think it would be good, if anyone in mathematics had a chance to write that first paper on a new proof?
> the more raw these proofs are, the better
There are parts of mathematics where the result is the important part. That's not what we're seeing here. Knowing whether the partition principle implies the axiom of choice doesn't meaningfully shape downstream knowledge and decisions. For these more foundational problems, clever proof techniques and the exposition around them are literally the point. Without that, neither humans nor AI can take this slop and derive anything useful.
And if AI can do that then great. I don't care about being elitist or not, and I'm fine with AI taking over math. That's not what it's done though, at least not yet.
> Mathematics need to redefine their profession and how they work (like us software developers too). There are very good reasons to believe mathematics has an important role to play. If they could just stop talking about OpenAI and get back to work - they are very much needed, in particular now!
The root issue is OpenAI et al.'s thoughtlessness in their engagement with a field.
OpenAI has resources.
That they fail to allocate enough of those to cleaning up pre-print papers (that seem to be a corporate PR priority for them to release) so they can be consumed and engaged with by the field they're targeting is... acting like a jackass?
It's the same "Meta / Alphabet can't vs won't hire more human reviewers" problem.
OpenAI could, at an immaterial salary level to them, pay a ton of PhD students and mathematicians just to clean up their proofs and papers.
Not doing so is a leadership and financial choice.
I'd prefer AI companies don't decide who's "cleaning up pre-print papers". This should remain the job of mathematicians at universities, which I am happy to pay with my taxes. Ideally there was something like a Bermuda Principles declaration for mathematics (https://en.wikipedia.org/wiki/Bermuda_Principles). This gave mathematicians even at poor universities and beyond the chance to participate in mathematical progress.
What do you think is gained, if AI companies manage "cleaning up"? Tax money?
If OpenAI researchers are publishing shoddy papers, why in the world would anyone else be responsible for editing and cleaning up their shoddy papers?
No one is "responsible". If the paper is read, depends on the interest of the individual scientist. Many mathematicians interested in a theorem proven/disproven in a new AI paper will be curious, even if it is utterly cumbersome to extract the relevant line of thought. But don't you agree that the job can only be done by a mathematician anyway - be she/he paid by the company or by a university?!
What you are arguing for is somewhat like the "proprietary period" in astronomy (e.g. see https://www.scientificamerican.com/article/nasas-plan-to-mak... ). But there is no mathematician who asked for the run of AI, i.e. there is no mathematician, who can be regarded as the owner of the result, even if ownership was temporary. Complex proofs might take years to explain in detail (think of the ternary Goldbach problem, https://en.wikipedia.org/wiki/Goldbach%27s_weak_conjecture ). It's just unfair, if AI companies held back with their data until a proof has been put into a nicely readable article, and at the same time mathematicians elsewhere are spending all their time trying to solve it.
> which I am happy to pay with my taxes
I'm not. To the extent my taxes are used for maths research (which is minute as a proportion of them) I want them to be used for the development of human mathematical understanding, culture and education, not trawling through a mountain of slop mechanically generated by a Silicon Valley startup that's about to IPO for trillions in the hope there might be some nuggets of insight hiding in there. If the latter is an activity with some value the startup in question can pay for it.
If you can't fight them, join them.
If you can't leave them, love them
> But it also sucks when mathematicians, who are considered experts on a particular problem, refuse to engage with breakthrough results about that problem.
For me the problem is that rigth now the structure of incentives that has been built (e.g. you publish more = you get a grant; good exposition < solving a conjecture) is now broken. So, for instance, you would be very irresponsible if you throw your student into one of those AI papers, it's too much the risk. This part is mathematician's responsability, they need to change this incentives structure.
In any case, OpenAI is being a dickhead here. They throw millions of dollars at these problems, but they can't afford basic literature reviews (the drafts barely cite previous work)? Or checking that Lean's formalizations really correspond to what they claim to prove (even for Navier-Stokes they made this mistake)? It's obvious that for them this is just a PR stunt.
Let us not forget that there's much more to this than just OpenAI being lazy and incompetent and willingly ignoring the high standards that researchers usually holds themselves to.
There's also the case of ethical violations, straight up scientific misconduct, as when OpenAI steals results of others (their customers) and present them as their own.
One particularly bad one came yesterday: https://arxiv.org/abs/2610.10072
> The result is also contained in a paper [8] released by OpenAI on October 6, 2026, in which the proof strategy and specific choices of notation are identical to a preliminary version of the present paper that was uploaded to ChatGPT on September 8, 2026.
Of course it's hard to say what to make of that without knowing what exactly went into the machine, but it certainly looks bad. And there's obviously a non-zero probability that it is indeed another instance of plagiarism, given that that's how they operate.
In this case, the author is a grad student, so what we're looking at is a company willing to steal from a student, ignoring whatever impact that could have on their career prospects, for a tiny piece of marketing material.
At this point I would be surprised if internal sandboxes are not trivially by-passed and that openai's agents do not (at the very least) have complete read access to all user accounts, chat histories and uploaded documents. Orthonogally, openai could still be wholesale lying about not training on this user data, of course.
So... If you use ChatGPT for anything of value, including abstract stuff like obscure maths problems, you should assume that at some point in the future OpenAI will include that in their training set and sell it onto other people.
Yes assume this.
But this is even worse, because there is no way that OpenAI "trained" on this data between September 8, 2026, the date Chenglong Ma uploaded the paper to ChatGPT; and October 6, 2026, the date that OpenAI released a paper with "identical proof strategy and specific choices of notation" (Ma). That's one month, that's not the timescale for model training.
So this implies _not_ that OpenAI is training on user input, in the conventional sense of adjusting weights; but rather that they are *straight-up channeling ideas from user input*, and with a very short lag. You would think there would be about a million controls to prevent this.
This is next-level alarming. I would be very interested in knowing whether Chenlong activated the privacy (do not train, etc) options in ChatGPT, and any other details of their setup (which plan, etc). Also, note that "do not train" might be, in a lawyerly sense, considered by OpenAI to be strictly about weights, and not covering "we hoover up your results and regurgitate them".
This is complete malarkey - OpenAI is _not_ "straight-up channeling ideas from user input"
Yeah that is my strong baseline prior. But it's colliding with what Ma reports. How do you reconcile those things?
Impossible to know without either more information about the actual occurrence or a deep understanding of the person making the claims. Just as OpenAI could easily be acting improperly, researchers who (understandably) feel deeply attacked by their problem getting solved right from under them might not be the most unbiased, either.
AFAIK OpenAI has not directly responded to the claims. If they ignore it, I don’t see what choice you have except to assume the worst.
You are significantly overrating researchers and the quality of work they produce.
> specific choices of notation are identical to a preliminary version of the present paper
I don't see how it's related to quality at that point. Call your technique "the banana method" and wait to see OpenAI invent "the banana method"
>Yes, the situation sucks overall
No it doesn't. Hundreds of open problems in a STEM field getting solved at once does not suck at all.
You would have to be deeply jaded and cynical to conclude that.
> Part of me would like a little longer in the world where the problem is still open and I am still looking for its solution. But reaching a summit, even by someone else’s route, comes with a view. From here I can see new mountains, and I look forward to climbing them with my students, collaborators and the machines.
https://dakshitakhurana.substack.com/p/classical-at-heart
"Climbing" is not the same as taking a trip on a funicular up the top of a mountain.
I think eventually there will be sort of a centralized more or less automated repository for ingesting and sorting ai-generated lean proofs and making them searchable and re-usable. On some level it kind of doesn't matter if mathematicians can ingest the results, if coding agents can just search for them online and use them in their own proofs.
I actually think it would be very smart for the big AI labs to get together to fund an independent organization to manage such a thing, and hire mathematicians to run it.
What is happening now is that some aspects of mathematics are turning into essentially an exercise in software engineering. It is well known that proofs and computer programs have an isomorphism, and I think the eventual merger is more or less inevitable.
That's not to say that there isn't an infinite amount of work remaining for mathematicians to do. There are only so many problems that are going to be amenable to this approach.
I think people sometimes overestimate what proving something means. For example the 4 color problem was a computer assisted proof from a long time ago. I doubt that many people actually read the proof even though leafing through the book is kind of fun if you can find a copy. Another example is the Kepler conjecture on sphere packing. The peer reviewers said they couldn't vouche for its correctness. While reviewers are volunteers with busy schedules, a lack of interest surely was part of the issue. This motivated Hales to formaly check his own proof (Gonthier formalized the four color problem earlier). But you can be sure that no one had been waiting for the Kepler conjecture to be proved in order to pack their spheres. Otoh, at the opposite extreme, the proof of FLT greatly advanced the field because the modularity theorem behind Wiles' proof is at the heart of a great section of math theory.
> there are still diamonds (in the rough) in this drop that perhaps should be looked into.
Right, and for some reason that "trillion dollar" company with mathematicians on staff didn't "look into" their own results. Almost like they don't care about engaging with the actual community they're dumping on.
I can't relate to this at all. AI models will surely get better at writing "enjoyable proofs," but for now the situation is what it is. You're passionate about this problem, right? But you don't want to do the work to understand the result? Fine. There's a new generation of younger, hungry mathematicians that are highly interested in figuring out why the result is true and I am sure they'd be happy to wade through it and spoon-feed you the answer instead. Maybe they should be running things.
So OpenAI should be able to flood the world with AI pollution and ask scientists and mathematicians to wade through it all and tell us if there is any sense in it, then sit back and wait for them to report in?
Nice idea.
Yes, they should. They have invented a magic button that can tell you the long-awaited answers to the burning mathematical questions that you've spent your life researching. The caveat is that the technology is still new, so the explanations "are not fun to read" like set theory papers usually are (lol). If you don't think that's a worthwhile tradeoff, that's your call, but it sure as hell isn't everyone's.
I don't think the claim is "they definitely have an oracle that solves the problem, and I reject it because it's hard to read". The claim is "OpenAI claims to have used an oracle to solve the problem. The proof is very difficult to read, and to even know if it does or not, we have to go through it with a fine-toothed comb, but they're going around claiming they definitely solved the problem (or at least getting press that claims that which they aren't pushing back against) and this might convince the people who sign grants even if it isn't true"
People who are invested in the idea that we've invented a general intelligence, now, which includes all these companies that are literally financially invested in this claim they are making, will tend to believe that its results can already be trusted in domains like this. Some mathematicians seem to believe some of the proofs written by their models, and some, like this one, don't. I do think it's valid for an expert to push back against the claim that the best use of their time right now is to verify the poorly written work of everyone who's claimed to solve the problem
Over the year, nothing they've released with a lean proof attached has turned out false (That's sort of the entire point. It's not impossible but it's really difficult). There's a reason most mathematicians, including the ones vehemently against OpenAI's dumping are not arguing the results are secretly false or have a high potential to be. And indeed, if that were the case, it would quickly become apparent and all this worry about grant signers would vanish into the wind. It's very easy to ignore nonsense. The problem is that it isn't nonsense.
I really don't know enough about it to know whether you're right or not, nor do I know whether or not you know enough to make the claim you're making, so I won't make an argument one way or another because it's non-sequitur to what I said anyway. The fact that you or I or Sam Altman or Terrence Tao believe the claim is irrelevant to whether this obligates the person who wrote the blog post to believe the claim, and it sounds like he's willing to consider the possibility that it is right, and would read the paper if it reached a threshold of comprehensibility expected of people making that kind of claim.
I's not a non sequitor because it cuts right to the point. He's under no obligation to read it sure, but that doesn't mean Open AI isn't justified in claiming to have proved it. The justification isn't Sam Altman's belief or Tao's or anyone else's authority. Results accompanied with lean-verified proofs whose formal statements match the problem at hand have arguably stronger justification than the vast majority of human math publications.
> Over the year, nothing they've released with a lean proof attached has turned out false
That's not true. [0]
> On July 25, Ramana Kumar published a repository containing a sorry-free "disproof" of the Collatz conjecture, produced with AI assistance. It is not a valid proof because it exploits a bug in the kernel's handling of nested inductive types.
Even in this dump we're talking about, it hasn't been true. [1]
> In “Algebraicity of Weil classes on split abelian eightfolds” a sign error invalidates a stabilization-trace cancellation argument and the construction used by two dependent papers.
[0] https://leodemoura.github.io/blog/2026-8-24-postmortem-for-t...
[1] https://github.com/openai/math/blob/main/history.md
1. That wasn't from Open AI or any major lab. I don't know what random people are getting to. It's curious also that it had no natural language proof attached. Usually these labs have a NL proof then translate to lean. Lot less possibility of lean maxxing.
2. None of the results Open AI retracted had an attached lean proof
The problem isn't just that the papers aren't fun to read. The problem is that a lot of the research that goes into solving these issues leads to other discovers, new fields to explore and people have to develop new approaches to solve them. The other part of it is, the quality and the enjoyment of working on these problems leads people to find new and other interest problems to work on.
If you just strip mine the answers and Sam Altmans magic button solves 100/100 problems, what's next? Who is left to come up with a new interesting question for the magic button to solve?
Lastly, life and the present moment is all there is, if there is no enjoyment in anything we do, then what's the point of all the "living for ever" Altman et al want to achieve.
We will live forever to read boring papers generated by LLMs? Literally sounds like an eternal hell.
Nuclear fusion is already proven by the universe to be a viable energy source by the fact that the sun exists but people still work on understanding and taking it and developing new approaches to accomplish it. People didn't stop experimenting with and developing programming languages because technically they're all turning complete and the first one was "enough." Y'all will be fine - every JavaScript framework that exists is someone looking at a theoretically correct and complete solution and deciding actually it sucks and they could do better. "I want to understand xyz but the proof is trash and I think it's ugly" will be plenty motivation for a lot of people to work on it.
How is the journey to understand fusion related to not wanting to spent your limited time on earth wading through AI slop?
You don't have to wade through the slop. The point is a giant star in the sky figured it out and that doesn't demotivate you from figuring it out yourself
Here's a guy who's made a "beyond n log n" tracker: https://x.com/aurel_pr/status/2108214135179944096
He's tracking the community progress on sub-n log n multiplication. OpenAI started with 1 - 1.63e-55. The result has been now improved on 115 times, and the current record is "rohanarun"'s 1 - 9.87e-5. I'm sure by tomorrow it'll have improved again.
Does this look like people aren't having fun? Does it look like they aren't discovering stuff? It looks like it's spurred a cascade of interesting community activity. It doesn't really seem much different from what happened with the twin primes conjecture. Isn't that supposed to be the point of all this?
Some people are having fun doesn't negate the rest of the issues with this sort of thing.
then the rest of them can pound sand.
People are already finding stuff in the release to get excited about, and as the models get better at distilling proofs to make them more coherent, this will only amplify. Of all the things to worry about, human curiosity and the ability to run with new ideas probably aren't at stake.
In fact, I can't remember a time when I was more excited about the future of science. This could herald an end to the replication crisis, and kill off bullshit science completely. The danger of course is that we end up with two companies effectively dominating cutting edge research in every field, but it remains to be seen if that's even possible given the pace of improvement in open weight models.
I don't understand. Why is this the onus of scientists and PhDs to review whatever results OpenAI had dumped out? If OpenAI had produced incomprehensible papers, surely any journals would just reject it, or demand the author to do a complete rewrite? Unless we are talking about a race to solve problems, which PhDs are afraid that they had been scooped up on?
What else should they do? See these models get smarter and smarter, somewhat-solve things but only to the tune of 90% what mathematicians (or experts in any other field) would deem acceptable, and then gate keep the findings for the next few years going through peer review and paywalled journals? I for one welcome the flood, bring on more in every possible industry and see where all that progress lands up. Sure it will upset a lot. A lot of things also upset the luddites.
Are you going to be doing the work to verify the results ? Will you just expecting other people to wade through the flood and reap the benefits later on?
Nobody is forced to verify the results. I am honestly not expecting anything other than AI to get smarter and smarter and people who are motivated and interested enough to pick up after it; and potentially reap all the long term benefits ahead of those who aren’t (seeing this happening with software development in my own field). But in the end; if you don’t like it, you’re not forced to do anything.
Wondering if you read the article?
Exactly.. this is an opportunity for people to pick up where the model left off and run with it. Like you don't have to, nobody is forcing you to.
However, there are always smarter, hungrier people out there and this is a buffet.
Some output is going to be wrong or incomplete. I am willing to bet even those have nuggets that can be used elsewhere.
> this is an opportunity for people to pick up where the model left off and run with it
Like people enjoy racing in front of a stopped train? As soon as they turn on the engine again, they will run you over. The questions that remain will be only the low value ones, not worth the effort to vacuum up.
So no, the smarter, hungrier people are not the ones that are going to swoop in. It will be the most desperate.
> Some output is going to be wrong or incomplete
This is a very human take on the situation. No, the Lean proof is not going to be wrong, and it will be incomplete only in the sense that OpenAI didn’t try to push the results further.
They should certainly be allowed to share their findings. No one is forcing scientists and mathematicians to review the findings in general. It’s just the case that the findings are of such such a quality that it would not make sense to ignore them wholesale
> No one is forcing scientists and mathematicians to review the findings in general.
This is like "no one is forcing software engineers to use AI tooling" or "no one is forcing you to show your ID in the airport" or "no one is forcing you to own a car in your small midwestern city" - there can be no law requiring something and the practical consequences of not doing so can be so painful that you're effectively forced anyway.
add "no one is forcing you to own a smartphone"
That is exactly what I’m trying to say. The findings are of such such a quality that it would not make sense to ignore them wholesale.
That’s why it doesn’t make sense to present AI companies as dumping or burdening the scientific community into doing labor for them; the scientific community is self motivated to do so.
It's not self motivated. The motivation is not "this is doing amazing things for us", it's "if we don't review this, the bullshit headline complex and the bullshit-spewing (sorry, marketing) departments of tech giants are going to misinterpret/misrepresent everything and our grant money will be taken away".
It is definitely self-motivated. People want to read and interpret these results.
No, they don't have to ask. Those interested enough will jump at the opportunity, even if it's just to be "one of the first to get it".
You know, I kinda relate to the feeling of not wanting look at those outputs if I think of it from a layman's perspective.
I just imagined that instead of math papers, they released 700+ feature length films, and the only way to tell if one of them is any good is to watch it in its entirety.
That feels pretty unappealing to me.
I know it's the same for human made films, so what's the difference right? But those are good enough most of the time that it's a decent bet, and the people that made them had real skin in the game.
Contrast that with something made by a nondeterministic slop machine with no skin in the game where small details can be off in a way that's jarring. Right out the gate I have an aversion to committing that much time to something that very well may waste it.
That's actually a really interesting thought. Given Sora, and the amount of funding they have, they could have created started their own film festival and dropped 700+ feature length films, had they wanted to go in that direction. But they didn't. Hmm.
From the perspective of "is this just a garbage dump to impress the market?", it's way too easy for investors to figure out 700 films are garbage.
The best part is that when the academics fix OAI’s issues, the model gets better and OAI shareholders get richer and more powerful!
As someone who uses LLM tech occasionally, this is why I prefer using open local models. If I’m making myself obsolete, at least I’m not making some asshole richer and their closed model better.
OpenAI in fact didn't know what to do with results and didn't want to flood the world, so they asked mathematicians. Mathematicians recommended OpenAI to release them. My guess is it would have been better for OpenAI if they didn't release them. OpenAI is basically doing this as a goodwill.
(See https://agmai.org/general-sep29/ for the recommendation in question.)
People seem to have very misguided ideas about why OpenAI is doing this at all. It is not to brag or to torture mathematicians. It is an eval. OpenAI is known to be willing to pay large amount of money to get a good eval, think FrontierMath. FrontierMath is now saturated, so they need a replacement eval for math. Open math problems are actually a fairly good eval, although a proper eval is better (eg FrontierMath has known difficulty and have tiers from 1 to 4).
Mathematicians would prefer if OpenAI didn't use open math problems as an eval, but OpenAI is not obliged. I actually think OpenAI wouldn't point AI to open math problems if unsaturated FrontierMath Super Duper is available, as it just angers mathematicians, but such eval is not in fact available. Given OpenAI used open math problems as an eval, they could just throw out the result (this is in fact better as an eval since it will keep problems useful longer), but mathematicians preferred to see the result. So OpenAI released them.
How is solving unsolved problems a good eval? Once a problem is solved and released, you can’t evaluate a future model’s ability to solve it.
As I said, a proper eval is better, but it measures something real that gives a good training signal, and there is lack of good alternatives for math eval. Since the result is 372/4000, it is also unsaturated.
I think this is right - it can be seen as trying to get free feedback from the community. In that way it’s reasonably described as exploitative, since it’s not a good faith effort. The problem of ai slop being submitted to conferences to get publication counts is similar.
Any idea what sort of percentage the sentence
> supported by a clear plurality of respondents
was referring to?
"Mathematicians recommended OpenAI to release them."
This is a misleading characterisation of the mathematicians' position.
The very first paragraph of the AGMAI recommendations explicitly states:
"we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models." You appear to have acknowledged this by saying “Mathematicians would prefer if OpenAI didn't use open math problems as an eval…”.
The mathematicians did not ask OpenAI to produce these results. They explicitly asked AI labs to stop producing them in this manner. Their subsequent recommendations concern what labs should do if they have already produced significant results, not an endorsement of the practice.
Furthermore, the recommendation was not simply to release the results, but to responsibly release already existing results. Section 2.B, Step I, explicitly recommends "...labs that have AI mathematical output that is not understood by the people who prompted the AI systems", to search the literature for relevant prior work, provide appropriate attribution, and improve the exposition of AI-generated proofs before releasing them, rather than leaving this work to mathematicians afterwards.
OpenAI published the results on GitHub while still exploring repositories that meet the committee's guidelines. So they followed some of the recommendations, but not all of them and hence, did not release the results as requested by the mathematicians.
I do not think it is a settled matter whether this was done out of goodwill. This is because releasing these results as they were can benefit OpenAI more than releasing them according to the AGMAI recommendations. AGMAI recommended in section 2.B, Step 1.5 that "Each time a solution to a problem is released, it should be clearly documented how exactly AI came to be used on that particular problem. If many results are released at once, then in addition to the results themselves a further document should be written and made public that references all of the released results and explains how many other problems of comparable difficulty the models tried and failed to solve, as well as how the problems were chosen." If the results are released, it is easy to expect that the media will discuss the capabilities of the AI used in the work, as indeed happened. If this AGMAI recommendation was followed, the media would plausibly have also discussed the number of failed attempts and then the overall attitude would not be as favourable to OpenAI as it is now when it comes to the capabilities of the AI that was used. OpenAI did release on GitHub that approximately 4,000 problems were attempted and resulted in 719 manuscripts (after 3 containing suspected errors were removed by OpenAI) across 372 families of problems, but this does not give a calculable number of problems it failed to solve. I do not claim to know OpenAI's intentions or reasoning when these results were released and am not arguing that it was done with improper intentions, only that whether it was done out of goodwill is not a settled matter.
AGMAI's October 6 statement explicitly clarified that its advisory role should not be interpreted as an endorsement of OpenAI's process, and that it was up to the mathematical community to assess how successfully its recommendations had been followed.
Recommending how to responsibly handle the outcomes of something you oppose is not the same as asking for it to happen.
It is also worth adding that the Association for Human Mathematics published a statement (which Tao reposted on his blog) in which they explicitly say the following:
> Mathematicians did not ask for this work to be done. The Advisory Group on Mathematics and Artificial Intelligence, from whom OpenAI has claimed to derive its legitimacy, opened their initial advisory statement by saying that frontier AI corporations should not test advanced mathematical problems on internal models. In ignoring the central premise of the Advisory Group’s position, OpenAI has indicated total disregard for the norms of scientific research — norms that guarantee that mathematics remains trustworthy, ethically researched, and in the public interest.
https://www.ahmath.org/
>> I do not claim to know OpenAI's intentions or reasoning when these results were released and am not arguing that it was done with improper intentions, only that whether it was done out of goodwill is not a settled matter.
That's right honourable of you but for me it is very clear that the only incentive in AI companies' effort to produce mathematical results is to advertise their technology. There is no reason at all to assume they have any other motive; certainly not any kind of interest in mathematics as such.
The problem here is that while you can call what OpenAI does "mathematics", I would hesitate to call it science. Science as a process of acquiring and developing knowledge within a domain involves a lot more than just dumping unfinished work on the scientific community. Among other things, it involves developing frameworks and understanding of the domain, formulating questions, creating results in a fashion suitable for verification/testing/replication, and relating these results to and integrating them with that edifice.
Kinda sounds like computer science vs developing - in the sense that people with a master's in CS and are dedicated to the craft will write wonderfully artistic software, while it doesn't actually take a love for the process to write code and get hired at some tech company (even less so now with agentic development).
Not what I am getting at. In fact, the problem I am getting at is the abolition of existing scientific and engineering principles and processes without a replacement and applies to programming with agents as well. This is not about artistry, but about building durable things.
Have you looked around recently? Because there are plenty of mathematicians that are excited to read and learn about all of the new results. They've improved on the sub-n log n result and have even made a web site to track progress on it: https://beyond-n-log-n.netlify.app/
It looks like people are enjoying themselves, having fun with the new results, and generally doing all of the things you say "science" is supposed to be about. So what's the problem?
I didn't say it is useless. Consider Ramanujan, whose work gave rise to a lot of interesting math, even (and sometimes, especially) the parts that lacked proofs or had other gaps (probably because it was obvious to him unlike us mere mortals). But a singular genius, whether a person or a machine, does not science make.
why are AI sites so hard to read... is it the terrible contrast? or something? any designers here?
There’s so much information, but it has no idea where to put the focus. Why is the number of PRs the same size as the actual best result?
Except for the fact that academia today is bs:
https://www.youtube.com/watch?v=LKiBlGDfRU8 https://www.youtube.com/watch?v=shFUDPqVmTg
The question is why the situation is what it is. Did OpenAI publish a large volume of unreadable proofs because that was their best attempt to contribute to the field of mathematics? Or does OpenAI feel that it’s more profitable for them if people come to see mathematics as something that’s less focused on understanding and more focused on using AI to generate proofs?
Yes, I guess it matters a lot whether this was quite close to the best they could do or the best they could do given specific resource constraints or whether they just didn’t bother to try doing better (eg to invest more tokens into readable papers)?
> I guess it matters a lot whether this was quite close to the best they could do
As the author of the post points out, there is no way this is “the best they could do”. It’s a write up that didn’t involve someone with the math + communication skills required to clearly explain the result.
Even with humans, the first publication is rarely the best expression of the lesson. Getting published does more than establish priority; it also frees up the community to build on the result. Nobody expects a human to wait until the proof is comprehensible to anyone except themselves and the referees.
I don’t think that’s true? I never did research math myself, but the people I’ve known who do would definitely invest time in making a proof clearer and better even if they had the idea basically correct. There’ve been multiple recent stories of researchers saying “we’re going to publish this lame proof that isn’t up to our standards because the AI companies let us know they’re going to scoop us if we don’t”.
Mathematicians care a lot about the exposition of their ideas, invest a lot of time in giving talks, writing, and don't publish too frequently, compared to other sciences.
They normally don't feel like they are in some kind of race to publish the results ASAP and claim priority. Cases like that are very rare (but they get media coverage because they are so unusual).
OpenAI did a publicity stunt, their motivation is not to make a good contribution to the field, which has very different standards and culture, compared to the AI labs.
You don't know what you're talking about. Just because it was created by an LLM, and verified in Lean, does not make it true. The whole point of writing a proof is for it to be understandable.
"true" != "understandable"
It's a necessary condition
Suppose I directed some llm agents to factor the primes from network logs of your machine. I then publish the private and public key in full. You would not change your keys of course because its not true right?
My argument is that LLM generation and a correlated Lean verification are not sufficient conditions. Both are falliable.
So is peer review. False proofs have been published before and will again. Does that mean journals are publishing slop?
How about the scenario where they are already better if you give them more time/tokens/…? Would it be a more relatable concern in that case?
Knowing a few PhD candidates and non-tenure track postdocs, the “younger, hungry mathematicians” are more worried about finding a damn job in the poorest job market - both academic and in industry - in a decade.
They don’t have any more patience for this.
There are plenty of mathematicians who think the results are interesting. Not only do they have the "patience" you speak of, some even seem to be enjoying exploring the new results.
Your point would be stronger with links to such discussion.
No idea why you're downvoted, my intial thought was that I would like to see some of this supposedly widespread sentiment as well.
I even think it's plausible a lot of mathematicians are excited by it, but the sweeping confidence of the comment you replied to without anything to back it up leaves some to be desired
Seems like an excellent opportunity to learn how these new proofs work and be at the frontier.
Checking someone else's work carries a lot of opportunity cost, and is only fruitful if one can learn new methods which apply to the own work. This is pretty risky, especially without tenure!
I realize that, but when you’re out of a job, you may as well chew on something like this.
When you are out of a job that's probably the last thing on someone's mind. Will I be able to pay rent for next month is probably their first thought
Right, so instead of adapting to the biggest revolution in their (and everybody else’s) chosen discipline, their best option is to ignore it?
You're passionate about this problem, right?
There's a "Silicon Valley-ism" for you. We offer a thing in whatever form we want and people "who are passionate" will gobble it up, should gobble it up, 'cause they're "passionate".
I think this might feel about the same as getting a PR from Claude that purports to solve some issue that it deems exists, but it doesn’t conform to the contribution guide, isn’t clear in its objectives and looks quite likely to be utter bullshit. I close them without comment and lock the issue.
"AI models will surely get better at writing "enjoyable proofs,"" why? Why is that surely true? they've increased in all other capacities at shocking rates while still writing awful, slippery, turgid prose. Very silly to assume that this will just go away.
Every single day now, for close to 4 years, ever since ChatGPT 3.5 was released - there's been people dismissing AI progress. Every step of the way.
It is entirely possible that one day progress just stops or slows down, but with current evidence, I don't find that too likely - at least not in the near future. The sheer amount of resources being put into this (AI) race is mind-boggling.
So while past performance does not guarantee future results, I'm just going to kick back, and assume that many of the current issues will be fixed with future models.
The progress is clearly not uniform though, it’s shaped by how amenable things are to collecting data and verifying computation. I think writing clear and cogent prose is harder to quantify than formally verifiable logic, and this is why progress in coherent communication has been slower than raw problem solving ability
Maybe this is just a matter of what model developers choose to invest training resources in, but I don’t think it’s inevitable unless clarity is made a higher priority
Doesn't address my point at all and I did not dismiss AI progress.
Yes it does address your point. Maybe actually read what he wrote.
OP is not dismissing AI progress, they're saying there's been no progress in writing better.
Which is false. Any of the sota models write far better than chstgpt that was released 3 years ago.
I think that would be very difficult to say. Good or bad writing is subjective.
Compare gpt-2 to todays foundation models. If you think writing hasn't objectively improved, you are willfully ignorant or a moron.
The tech industry is chock full of straight up magical thinking, and nobody shows it more than AI boosters.
I think it’s silly to ignore the sheer rate of progress that’s occurred in the last 5 years. Compare gpt 2 output to latest foundation models and then re-read your comment. You’ll realise just how out of touch you sound.
> There's a new generation of younger, hungry mathematicians that are highly interested in figuring out why the result is true
The problem is that right now mathematicians don't have the economical incentive to read these AI generated results. Even if you love mathematics and all that, it's always more important to get a job, and for that it doesn't seem like a good idea to invest time around problems that AI touches because you can't compete with it and you don't know if tomorrow they'll improve by x10 the sota.
I am a mathematician, and I do have the incentive to read the results. Two results in the drop were two major life goals of mine, and all I got was three lousy citations. :) But a third question I've spent a lot of time on is a not-so-hard consequence of one of the lemmas in there. So, yeah, I do have the incentive.
Of course, it's not clear at this point whether reporting such a result even matters, but still. In its own right, it's a very cool result.
Thanks for your perspective, genuinely.
What do you do if those results suck like in the article?
Of course, everyone is curious about these results, we're mathematicians after all. But by incentives I refer to those things that really sustain the system: obtaining permanent positions, grants for future projects, ideas for PhD theses, etc.
Like all the junior devs getting opportunities left and right.
I think quite a few mathematicians would be interested in using AI to figure out the answer.
But providing the answer in gibberish along with a certificate is not that, it's at best a cruel way to do it, but I'm leaning towards the idea that it's a fundamental misunderstanding of what it means to do math and what it means to communicate a result.
If you think sending an answer in gibberish is acceptable just because it's true then SSdtIG5vdCBzdXJlIHdoYXQgdG8gdGVsbCB5b3UsIGJ1dCB3ZSBkaXNhZ3JlZSBvbiB0aGF0.
SXQncyBub3QgYSBtaXN1bmRlcnN0YW5kaW5nLCB0aGV5IGp1c3QgZG9uJ3QgY2FyZSBhYm91dCBtYXRoIGF0IGFsbC4gSXQncyBhIG1hcmtldGluZyBzdHVudCB0aGF0IGEgc2hvY2tpbmcgbnVtYmVyIG9mIHBlb3BsZSBoZXJlIHNlZW0gdmVyeSBlbW90aW9uYWxseSBpbnZlc3RlZCBpbiBib29zdGluZy4=
Very sensible comments. It is along the lines of the fury I get when I am confronted with an 11 page dump of an issue analysis created by an AI agent that makes no sense but I have to go through because customer shared it.
If you did’t bother to write it, I shouldn’t be bothered to read it.
Perhaps AI agents can have their own publications and magazines where they are the chairs and associate editors and reviewers.
Nobody came to him and forced him to read it, just like nobody is forcing you to read code I had chatgpt write.
If AI can solve such grand, outstanding math problems, and mathematicians argue these pure math problems are important, what’s the problem with them needing to read the output if they want to understand it?
Can't you read the article and understand how he kind is obligated to engage with it ?
I’m not sure OpenAI shouldn’t have released math papers because some guy made a wager about one of the problems that loosely socially obligated him to read material about a solution to that problem.
The absolutely funny thing is, he is obligated precisely because the proof is very likely to be correct.
The alternative you are proposing implicitly is even crazier. OpenAI should not release a proof that is most likely correct so that it doesn’t burden others. What? It’s not about that guy dude. It’s about the society TM. One can’t delay progress because a guy may be burdened.
“Guys plz don’t release this thing that is absolutely correct but I’m kinda busy with other things ok?”
> The alternative you are proposing implicitly is even crazier. OpenAI should not release a proof that is most likely correct so that it doesn’t burden others.
The alternative is that they do the work to properly present the results. They spend billions of dollars in AI training and inference but can't afford to even cite the literature properly? They're doing the bare minimum because they're inly interested in doing a PR stunt.
they will do it. in a year, the same mathematicians will cry about (O)AI making their lives hard by not only solving more problems, but also presenting them with "enjoyable proofs". They'll still cry because the current excuse is a veil.
Bare minimum is still _solving_ the open problem standing there for years. Nobody owns math. Nobody owns giving enjoyable proofs to someone else.
If you don't like to engage with OAI proof dumbs in current state, don't. Maybe others will. Or maybe _these_ mathematicians are afraid that _other_ mathematicians will do it. Just elitism and gate keeping.
> but also presenting them with "enjoyable proofs". > Or maybe _these_ mathematicians are afraid that _other_ mathematicians will do it. Just elitism and gate keeping.
Ok, I don't see the point of discussing with you. It's clear that you decided what to believe in and no evidence will convince you that reality is more complex. The proof is that you ignored all the nuances expressed here by simply sticking to your simplistic interpretation, without any explanation of why such nuances are invalid.
Whats the nuance here? Your post implied that the problem was unreadable proof and we are saying that this is not central to the discussion. Unreadable proofs are actually very very irrelevant to this whole drama
I think you’re out of touch with the discourse in the mathematical community, it’s quite relevant because clear communication is an important aspect of intelligence
Proofs can be unreadable for more than one reason. Are these ones unreadable because the math is super advanced or because current agents suck at clear writing? Maybe a bit of both?
Fine and we are saying that this is not central to the discussion because even if models wrote it nicely, there'd be even more outrage.
Do you disagree with this? For example, if openai had provided really readable proofs with utmost care but still dropped 400 at once, would there have been less outrage?
> if openai had provided really readable proofs with utmost care but still dropped 400 at once, would there have been less outrage?
Less criticism, yes
There’s always going to be outraged people, but outrage isn’t the word I would choose to describe the positions of the mathematicians I’ve read on this topic, including TFA. There’s a lot of optimism mixed with frustration that something important is missing
Ok agree to disagree. I think readable proofs are smallest of problems and a distraction to a bigger problem which is loss of meaning.
> Do you disagree with this?
Absolutely. If they cared about properly citing the bibliography, sharing how their models work, etc, it would be so much easier to make an assessment of the situation, of what will be math in the future, and all those deeper questions. What we got instead? from those 400 papers, there are already 4 that have been found to be a copy of recent publicly-available papers. They spend millions for "the good of science", but they can't afford basic scientific ethics?
>If you did’t bother to write it, I shouldn’t be bothered to read it.
You don't think this changes when the thing in question is a proof of a STEM problem no human has ever been able to solve?
Why should it change when there's good chance it's just slop? That's exactly the standard that human mathematicians are held to -- and it's their job to refine and polish the work to make it understandable by the community. And standards have evolved that way because otherwise there's too much slop to wade through (eg. you'd be surprised how many papers the typical theoretical physicist gets claiming to have proven Einstein wrong -- from absolute crackpots who don't understand the basics of the subject). This will just multiply now with AI, and doesn't change just because the prompter happens to be on OpenAI payroll.
We'll see what the final slop rate is, but the three papers they retracted yesterday were for a trivial sign error. If they didn't catch that, that means OpenAI isn't bothered to put in the minimum effort of sifting through their own garbage and making sense of it.
It's not like they needed to hurry out this release before carefully vetting. They're just "hacking" the math system and disrupting the work of thousands of researchers to create a gigantic RL dataset for themselves.
Really, what's the bloody hurry?
They could have released 1-5 papers, worked with researchers to understand what methods work and what don't, how to prompt the models better, how to build better guardrails for reasoning, etc. And give those researchers access to latest models and empower then to solve thousands of problems!
Instead OpenAI wants to piss all over the city to claim territory and now human mathematicians have to go around cleaning up that slop, only so that OpenAI can made some bullshit statement like: math is solved [mistakes are next].
--
Imagine someone gave you a million line PR claiming to have vibecoded the operating system of the future (or whatever your application domain). Would you drop all your other work to focus on this? And they generate enough PR that your manager and company leadership and public all start pressing you to accept it quickly? Guess what, it's your lucky day! You have not one, but 700 breakthrough PRs!
Imagine someone gave you a million line PR claiming to have vibecoded the operating system of the future (or whatever your application domain). Would you drop all your other work to focus on this?
If it verifiably works and solves important outstanding issues, then quite possibly yes.
And if three of their 700 PRs were retracted within a day because of unresolvable bugs? Clearly it's not all verified.
Not to mention the math paper (on HN yesterday) which pointed out that OpenAI might have verified the wrong thing, in their Navier Stokes proof.
It's really not as cut and dried as you (and software/AI folks more generally) think it is.
PS: If a human mathematician had to retract three of their papers a day after posting publicly, they'd lose all credibility and their mathematical career would be all but finished. That social incentive structure is the field's immune system against slop. You're basically asking them to turn off their immune system, and for unclear gains (other than OpenAI's grandstanding).
If a human mathematician had to retract three of their papers a day after posting publicly, they'd lose all credibility and their mathematical career would be all but finished.
If a human published 700 papers and only 3 (or 30) ended up having significant errors, I'd call that a pretty good batting average, considering that the typical rate of errors may be around a third (https://lamport.azurewebsites.net/pubs/statistics.pdf). But for some reason people hold AI output to an absurd standard where if it's not 100% perfect then it's useless.
You're basically asking them to turn off their immune system
I'm not asking "them" to do anything. They can do whatever they want, including rejecting obviously useful tools. But then they shouldn't be surprised when they're quickly surpassed by others who don't share their ideological blinders.
A lot of the comments are claiming that "no one is forcing them to engage with AI proofs" and that's not the case, as explained in the article. The author is forced to engage with the public by the very nature of being a prominent researcher on this problem. The public is drowning him in messages regarding this result. So yes, he is being forced.
Academics have always been required to engage with hacks and cranks to some extent; the deluge of AI proof writing has only exacerbated the problem.
That is an incredibly loose use of the word "forced"
I am not a native speaker but in my experience it is an accurate use of the word “forced”.
English speakers generally use this word in a very broad sense “and now Netflix is forcing ads on paying users”, “because there was no sink, I was forced to drink the whole thing”. It is only when you are literally describing a crime where this word has this strict meaning you are alluding to.
Those are perfect examples because those are also hyperbolic and unserious uses of the word “forced”.
Merriam Webster seems to agree with me: https://www.merriam-webster.com/dictionary/force#dictionary-...
> forced; forcing
> transitive verb
> 1 :to compel by physical, moral, or intellectual means
> A player was forced out of bounds; They forced the CEO to resign; I forced myself to finish.
Hyperbolic, perhaps, but how does that make them "unserious"?
Because nobody is being seriously coerced or forced!
it takes two keystrokes to type NO
Just ignore the emails. Academics are already very good at that.
"not at the Lean code, since I know very little of the actual usage of Lean, and that code was enormous"
This part I don't understand. Not that anyone should read the entire Lean code of any proof, but if the statement of the theorem to be proven in lean seems to be correct, then I would think there would be at least some interest if in fact there was a formal proof (which might or might not correspond to the written proof) of something I was working on. That to me would be interesting. Or you are saying you doubt the validity of the formal proof, which would also be interesting. But saying it is of no consequence doesn't make any sense to me.
I think TFA’s point is that it’s interesting - it’s just not feasible to do what follows after “it’s interesting”, which is to try to make heads or tails of the stack of writing that we’ve been given. Engaging with a well-written proof of a similar scope is enough of a task already.
Isn't it obvious that the feasible way to respond is to start developing AIs that rewrite proofs for human understanding
Some of the lean proofs are apparently incomprehensible.
It would be like trying to look at a completed video game's assembly code, being told that it was call of duty, and then being asked questions about the high level code architecture.
AI models are perhaps unsurprisingly good at low level translation (see the progress being made for decomp games)
These models have surpassed human capabilities at math/machine code, but they can't "simplify" yet - in part because they don't have the same need to due to their comparative lack of cognitive constraints. AI Slop code is getting better, but it takes time. At the moment, its embarrassing frankly. It will come eventually, but right now OpenAI is not handling this with the care, respect, or concern that it deserves.
Have you ever wrote some code/algo that seemed "simple/obvious" to you yet to someone else, it seemed incomprehensible?
If you have a 20-40 IQ points gap with another developer, this happens a lot.
The baseline of "simplify" is wildly different based on your IQ points. That's precisely why exceptional students are usually bad in teaching. They try to break things down, simplify, but things still go over the head of normies.
However, we can intervene/train the models. So it should be possible to focus on the simplification, and as you said, it will come eventually.
I've been mostly reading here but created an account to disagree with this statement: The smartest people I know were always amazing at explaining. This was true for me as undergrad and graduate student where the smartest peers and the most renowned professor were always also the best at explaining, and it is true now at research level in a related field. When I am not sure if I really understood something to the core, I try to find a colleague who knows very little about it; if they understand my explanation well, that's a good sign.
Those who really understand a topic are usually also able (and great at) explaining it in very clear and "simple" terms. This may be part of my personal bias; I see theory builders as those who advance the field the most, and these are usually also amazing at explaining it. On the other hand, those who mostly "grind" through problems (approach them as complicated puzzles) with effort/time were often bad at explaining.
I observed the same for programming: the "architects" usually explain very well, the "debuggers" often don't. LLMs very much remind of the grind/puzzle approach. It does makes sense that RLVR, which in my understanding enables a lot of these results, would lead to a more mechanical approach.
Of course, I can't make any predictions on whether it will stay that way. But I strongly suspect that we need different ways of training for LLMs to write better text and explain better (I suspect the vagueness of LLM language is the result of RLHF as vague expression is less often incorrect).
I don't disagree with you. However, you are ignoring one key point that I made: the difference of base intelligence.
Let's assume you have an IQ of 135. You are not that far from even the crazy smarts at 150 and as long as they have some social skills, it will work out. I, myself, never had issue with understanding the smart people.
Now let's assume you have an IQ of 110. Unless 140/150 IQ people are spending inordinate amount of time formulating how to simplify things for you, things will fly above your head.
This does not mean just because someone is 150, and someone is 110, the 110 is always the problem.
This also does not mean an 120 explaining something to 140, and failing is 120's fault or the other ones'.
Every combination is possible as we are talking about personalities/ability to articulate oneself.
However, I was talking about a specific scenario, which I have faced myself and saw it happen a lot of times (to other people), where someone's simplification level is still higher than someone's max understanding level simply because the deviation between their caliber is too great. In most cases, this can be worked around by the explainer spending inordinate amount of time.
I agree, and I disagree. There are plenty of published mathematical papers that are just as poorly written as OpenAI's. Nobody says nothing because the authors are big names. In some cases, the proofs are not even correct, but everybody has a feeling the result are true nonetheless, so they pretend not to see it. So I agree that OpenAI should have done a better job of writing down the results, probably by paying working mathematicians like Anthropic did. But I disagree that this low-quality writing is somehow a good reason to be angry at OpenAI specifically, otherwise you would have to be angry at a lot of people.
These takes forget that OpenAI is doing all this as a PR stunt. They're utilizing the field without worrying about any consequence to it.
> There are plenty of published mathematical papers that are just as poorly written as OpenAI's. Nobody says nothing because the authors are big names.
Can you provide some evidence of this claim? "Nobody says nothing" probably works on reddit but I generally expect higher quality discourse on hackernews.
I am a working mathematician. A problem that I cared about greatly (and probably spent > 3000 hours working on) was on their list. I looked at the paper, and I have to say it is more clearly written than about 30% of the papers I typically referee. I don't want to name poorly written papers, but I agree that "there are plenty of published papers that are just as poorly written as OpenAI's".
And which papers do you typically referee? Without that information this claim is meaningless. For example, if you referee free-for-all papers that today are likely written by LLMs as well then sure I can understand that. But if you referee papers from grad students then that's more concerning.
I have only twice (knowingly) refereed AI slop. I'm mostly talking about refereeing in the period 2010 - 2020. (However, looking here https://proofsandprompts.com/2026/10/08/100-reactions-to-100... it seems that many other consider many of the papers poorly written. I just wanted to give one datapoint.)
Random questions--
Were you satisfied with the paper?
Having read the paper, do you understand "what you missed" in those 3000 hours?
The response by some in the field of mathematics to this repo is ... I guess not unexpected; but it's quite disappointing.
I sympathize with those who've worked on some problem for years and now don't have something to work on; it's been a part of their identity. I also especially sympathize with those whose career tracks and plans were thrown in disarray.
That being said, I absolutely cannot understand how one can't be excited and happy and enthused about these advances in one's field. Assuming just that the ones with formal lean proofs are actually true, these are reportedly huge advances. Even if folks don't understand it YET.
It is because in fact MOST mathematicians (and researchers) don't really just want "advances" in their field. They also want to have some significance and play a role in it. What's the point of doing research, which takes a lot of work for much less pay than industry, except to have some of the glory?
I didn't go into math so I could read the results of others all day. I went into it to contribute meaningfully, and I certainly don't consider interpreting the results of a machine to be a meaningful contribution.
Your perspective is just the perspective of a consumer of things. In that case, it doesn't matter where they come from.
This is the key bit.
> But, back to the Partition Principle. I took a brief look at the preprint released by OpenAI (not at the Lean code, since I know very little of the actual usage of Lean, and that code was enormous). It sucked. It is unclear, muddled, and has a strange structure.
Incredible lack of curiosity. The Lean artifact shows that there is a proof. Maybe the natural language writeup sucks (maybe it doesn't even correspond to the Lean proof!) but the proof is there and if he were really interested in the problem he would try to understand it. Rather, his revealed preference is that what he's really interested in is good style in academic papers.
the OP's point is that if another academic had submitted that paper to a journal with their Lean code, the paper would likely have been rejected by reviewers regardless of the Lean code
it's just holding OpenAI to the same standard as everyone else
> The Lean artifact shows that there is a proof.
how are you sure, if it can't be explained properly?
update: to me this feels like the equiv of dumping an enormous PR that probably has some great stuff in the code but is poorly explained and documented, and expect the maintainer to try to make sense of it and see if it's valid merge
>>Incredible lack of curiosity
Maybe on your end because that's the exact point the author explained - these companies need to adhere to the scientific community rather than expecting the opposite.
If the write up - the natural language part - sucks - why are you expecting humans to waste their time on understanding the proof???
Because there is a proof! If you are interested in the conjecture, then you would be interested in the proof. If you are not interested in the proof, then you weren't really interested in the conjecture. You were interested in the good style of articles about the conjecture, or the friends you made along the way, etc.
It is of course possible that either (1) the statement of the theorem in Lean is busted or (2) there is a bug in Lean. But both of these seem to me to be lower probability than that the proof is correct. It's not just an e-mail from a crackpot. For someone who is actually interested in the problem, the probability that the proof is correct is high enough to warrant effort to understand it, or at least to learn enough Lean to check the theorem statement.
>these companies need to adhere to the scientific community rather than expecting the opposite
Why? The pre-existing community doesn't own science.
Considering how hostile the community sometimes is to outsiders, how they frequently demand form over function and think connections are often more important than correctness of argument, perhaps it's good for them to be confronted with a new approach to science that does away with those things and returns to the real cornerstones: proof and empiricism.
If you don't want to engage with proofs that's your call, many of us are happy to see progress being made and don't need you specifically for the confirmation
At first I thought this was going to be more Luddite babble, but it makes a good point. OpenAI isn't contributing if they are make unreadable papers. They should use a little more of their compute to nail interpretability. The difficulty will only get worse as AI plow deeper into the frontier and produce increasingly alien looking output. I suspect it's a workflow issue. If not, it's a bad oversight if the current generation of models are capable of making mathematical breakthroughs but can't explain how they build on existing frameworks.
If you don't think they have contributed anything then you can safely ignore them.
Exactly what the author is doing.
Work this abstract almost certainly has no value outside the community that is (was) interested in the result. OpenAI should engage with the community to realize the value (beyond PR).
I can't 'ignore' OpenAI's paper any more than you can verify it unless you are an unusual person with a very uncommon skill set. Very few people with the skill set to look at these are getting slammed with difficult to read artifacts, which is confusing because they have internal AIs so smart that they can solve hard math problems.
Is it just me or is this line of thinking fundamentally dishonest? It may be true that the papers are hard to read but the papers have extremely high signal towards the proof.
All of these comments seem to suggest that if papers are not 100% abiding by readable books their standards, they are basically the same as literal noise. This is highly dishonest.
I’m just a dude and even I’m able to understand the paper after using ChatGPT to help me through it.
There's certainly a conceit in saying if the math is not accessible to my community then it doesn't even count as math, and the last few days this has been a common ideological talking point.
That’s right and readability argument is a distraction
https://news.ycombinator.com/item?id=50016860
OpenAI is publishing the math because they want mathematicians to read along and verify the work they are doing. The author complained that even the citations don't make a lot of sense. This is completely consistent with my experience of using AI for coding. When I task my agent a hard problem and it comes back with a 5000 line PR that I don't follow I reject it and work it into a better shape. I certainly don't slam it into production because it has a 'high signal' of the program I asked for. The mathematician's situation is a lot worse because they have no view over the process that created the artifact like we do with coding agents.
> When I task my agent a hard problem and it comes back with a 5000 line PR that I don't follow I reject it and work it into a better shape. I certainly don't slam it into production because it has a 'high signal' of the program I asked for.
??? That's literally what everyone does. You are creating a hypothetical that is nonsensical. Here's your hypothetical:
1. Agent gives you 5000 lines of slop
2. You reject it and just do it yourself
This is reality
1. Agent gives you 5000 lines of slop
2. you realise that it has done a lot of research and is mostly in the correct direction and you ask it nicely to refine it
3. verify that you understood it and push it to prod
Are mathematicians babies that they need a completely different approach?
The mathematician in this case didn't ask an AI for the proof, they were handed a badly written artifact and other people are messaging him asking for analysis. He has no visibility into the process that produced the paper because it came from an internal OpenAI model. The bad behavior here isn't coming from the complaining mathematician. OpenAI hasn't sufficiently developed their workflow to create cutting edge research AND publish it intelligibly, which is odd because the latter should be easier. OpenAI has perverse incentive to do this because spamming badly written papers can still give them priority credit for solving open problems at the expense of mathematicians (who are currently indispensable because they decide whether claims are credible). What Mr. Karagila is trying to enforce are the old standards that review should come after refined explanation of the solution. If that was the case before AI, I can't see a reality in which that isn't the standard now that papers can be pumped out at faster-than-human time scales.
My point is that
1. the proofs are hard to read
2. but are still sufficiently high enough signal to understand the proof
3. blaming OpenAI for releasing high signal information on important problems is wrong. OpenAI did the right thing here.
4. the mathematicians can take however long they want to digest it and verify it. They have no legal obligation to do it immediately
This is exactly how I'd like society to work. Publish and let information get democratised quickly.
It doesn't seem like the paper is really hard to read, but rather that the terminology is "a bit off", with "unexpected theorems" and references that are correct but that could point to more useful targets (read: papers that will increase the citation count of the researcher being cites...). None of which are valid reasons to be dismissive of a good result.
> I’m able to understand the paper after using ChatGPT to help me through it.
How could you possibly know that? If you don't have the knowledge to understand the paper without using a chatbot, you don't have the knowledge to verify that what the chatbot said is the same thing as what's in the paper.
Maybe the reason the AI could find this proof is exactly what the author is complaining about: that it left the beaten path of theorems expected in a paper like this and went off in an unexpected direction.
That's exactly what happened for several of these theorems.
I'm just amazed at how quickly we moved from "AI is just a parrot and can't do anything useful" to "AI can solve toy problems but not anything of value" to "when AI pushes the state of the art, it can't quite get the proper citations in its preprints".
And it's not 2027 yet. We are so myopic. Ten years from now is going to be insane
In short, we’ve reinvented cranks sending unsolicited, poorly written putative proofs
Except that the proofs are very likely valid. This is a country of Ramanujans in a data center. You're free to ignore them because they don't follow your style guide, the rest of us will enjoy seeing humans and AIs build on the results.
> the rest of us will enjoy seeing humans and AIs build on the results.
That would be nice, but the rest of the world wasn't interested in the results before and they won't be interested after.
I wonder what this means long term. Maybe mathematicians will keep plodding on as usual except sporadically when an AI company need a marketing boost so they spend millions of dollars to dunk on them. Because the mathematicians sure don't have that kind of money.
the rest of the world wasn't interested in the results before and they won't be interested after
I can't speak for the rest of the world, but I think it's amazing that multiplication and 3SUM are sub-quadratic despite that being "obviously" impossible.
I wonder what this means long term.
Long term, stronger models than what OpenAI used will be available to everyone. Some mathematicians will take advantage of those tools and do great things; others will continue their whiny gatekeeping and become irrelevant.
> Long term, stronger models than what OpenAI used will be available to everyone.
Strong and cheap enough to make solving these problems by this approach less than a million-dollar venture?
(Pedantry: We already knew that integer multiplication, at least, was O(n log n). which is subquadratic.)
Mathematicians struggle with the same problem as software engineers; you can let AI generate the artifact, but to understand fully what is going on is challenging. Perhaps even more for mathematicians.
Do you rely on the tests/Lean to accept correctness or not…
I would have expected the math community to celebrate all of this since they got into math for the love for mathematics rather than the love for tenure. Their own world will not end - more mathematics will require more mathematicians to make sense of it all. And if all of it leads to some form of superabundance they are looking at the best possible future: they’ll be able to do math forever without having to worry to get paid for it.
> they got into math for the love for mathematics rather than the love for tenure
They got into math for the love for mathematics but if they couldn't make a living from it (tenure) then they would have done something else just like everybody else who is not a starving artist. So, no, they won't celebrate it and neither would you if you were honest. (note also that starving has a short time limit before you die so "if all of it leads to some form of superabundance" won't work in that circumstance)
> Their own world will not end - more mathematics will require more mathematicians to make sense of it all.
Someone tell the administration to restart funding mathematicians then [1] -- because the exact opposite of "more mathematicians" is happening right now, both at the post/graduate level [2], as well as the undergraduate level (for much longer) [3].
[1] https://www.ams.org/news?news_id=7694
[2] https://www.ams.org/learning-careers/data/impact-report
[3] https://www.ams.org/journals/notices/202310/noti2806/noti280...
>OpenAI drops some hundreds of "solutions", incomprehensibly written "solutions", and we are all expected to jump on them and what? Appreciate their contributions? Why?
...
>So, no, I will not be sending Sam Altman a bottle of whisky anytime soon, nor I am planning on spending my time reading through that paper and trying to make sense of it.
Think about a hypothetical circumstance where we get radio communication with some aliens on another planet. They send over tons of math to help us advance our tech, we know the math they're sending us is correct, but their explanations are really hard to work through because they aren't humans and the math is so different from anything we've done. Should we whine about the results they sent to us and refuse to engage with it?
In this hypothetical case funding agencies would probably happily pay for “alien maths translation grants” and you could write papers about advances in that field and get jobs etc. So arguably for better or worse it would create a specialised academic cottage industry that somehow interfaces with the rest of maths.
Love your metaphor here. I guess there will always be people who try to keep going and follow their current projects and say that only human math is real math and we should only use alien technology for proofreading mails and such but not their highly established and esteemed professions. Or just keep your hubris low, update your priors, work towards progress.
> to help us advance our tech
The major discussion is about whether this will in fact advance our tech.
In the Three Body Problem, the alien race shows us “miracles” in an attempt to discourage us from the pursuit of science. Consider that possibility.
As a whole, the math community would want to simply ignore the distractions from AI companies … they are using math problems as medium to promote their commercial interests.
> not at the Lean code, since I know very little of the actual usage of Lean, and that code was enormous
Dismissing results on the basis that Lean code is too long disqualifies this opinion. It is not hard at all to read the Lean result statement, even with very superficial Lean knowledge.
He's probably talking about understanding the structure of the Lean proof, which is 233,891 lines of Lean (including blank lines).
There is no nice way to tell someone that you’ve scooped them, and this is industrial scale scooping.
A few papers have been retracted, but it looks like many are withstanding intense scrutiny. Lean is making the results more likely to be correct, but I think making them harder to understand.
The world has changed and you’ll know a math department is making a serious attempt to adapt when it teaches a required Lean course in freshman year.
Why do you believe that learning Lean is a better use of time when the AI is clearly better at writing and interpreting Lean than it is writing quality papers? At this point, Lean is for autoformalization, no one is really supposed to read it.
Well, you don't necessarily need to read the proof, but the proof is useless if you don't read the specification.
That's fair, but I would argue that reading the specification is very easy by comparison. A quick one hour tutorial is usually enough judging from my students' experiences.
The Lean proof to Fermat's Last Theorem is 13 mil lines.
The one for the quasi-Riemann Hypothesis is half a million.
Indeed! Almost none of these most math folks are likely to read. The only thing to read is the statement, which is only a handful of lines in both cases. The point of Lean is that if that statement compiles and is validated by hand to be equivalent to the natural language statement, then it is true. That was what I was trying to say here.
> "why you should be very angry at OpenAI"
Please don't tell me how I should feel. Stick with the facts.
Love this!
> This is why I generally avoid using AI for mathematics (I am happy to ask LLMs to consolidate information for me, or to generate a useful infographic, or to proof read an email, etc.)
In other words, the author is OK with using LLMs to replace data analysts (that could consolidate information), to replace graphic designers (that could generate infographics), and to replace editors (that could proofread an email). But don't you dare use LLMs in their mathematics.
I don't know about this person, but for me and I'm guessing the average professional in basically any field, before LLMs, they simply ignored information that needed summarizing, made their own crappy slides, and didn't proofread their emails.
People seem to be very likely to say that LLMs are good at things that they don't want to do. That's also why you see so much of "well of course they can't do x, but they're great at y" but everyone has a different x and y.
openai will take expert responses like this and improve the next set of papers
it won't be long before there's no more low hanging fruit like this to complain about, and the writing / explanations of the results are superhuman as well
separately, i really liked the author's denial-of-service analogy. super useful practical framing
Look forward to seeing an LLM write something well, that will truly be a breakthrough in the field.
>That's not what you'd expect from a serious preprint claiming to solve a problem.
...
>It seems to me, that the AI tech companies would like us to conform to their standards, rather than spend the time and energy to conform to ours.
The entire argument is ugly and unfamiliar methods are used, and nobody's got time to drop everything and understand all that, meanwhile OpenAI didn't even ask us about this, therefore we should reject the paper as "unreadable" and you should be "very angry at OpenAI" today.
No, it's the press where you need to direct your anger. That is the nature of the beast. And being arrogant at the time of mass-disruption is a great way to lose control of everything. Maybe Sam should be sending the bottle of whisky to you.
I think the problem is that they mixed problems that some solutions in Lean (and I consider they solved, assuming someone checked the formalization of the statement) with some solution in English.
I expect most of the solutions in English to be correct or mostly correct, but we all have made and read mistakes mixed with the bla bla bla. In the Mythbuster scale I'd classify them as "plausible" instead of "confirmed".
Very well written summary of the current situation
Imagine being someone who is working on one of these problems. You have no good guarantee that the problem was solved, but you will have the horrible homework of reading the AI slop. Also, if you do have something interesting to say about the problem, people will have less enthusiasm about it now
The gradient into formalisms is steep and the bottom is deep. These are the worst results anyone will ever see again. The bottom of the internet is also deep, but won't be around much longer.
I am no mathematician, but I can definitely relate to the "horrible homework of reading AI slop". We've started allowing a non-developer to submit AI generated code into our codebase (in droves).
When I am tasked with reviewing that code. First, I don't know if it's valid. The person who wrote it doesn't know if it's valid. In order to validate it I must step back and understand the full problem space. Then, when I ask for revisions or clarifications, it's seen as either
A - Slowing progress, being resistant to change... or... B - Thanks for catching that (Claude fix PR 532 with the review comments)
It's 100% removed the enthusiasm.
Perhaps, there is a point where we just "give up" the understanding and accept AI output as the ground truth, because the sheer amount of generation is too much for our puny human minds to comprehend, and a lot of the times it IS right, even if a little wonky.
I read one of these papers (a relatively short one), solving a big problem in a field I used to work in, and it's exposition was above average.
If he doesn’t like the paper’s structure he can just prompt gpt pro to review and amend.
I know the model that produced these proofs is still private, but it’s worth a shot tackling the proofs with the current consumer-available frontier.
LLMs causing problems? The solution, apparently: more LLMs!
How would you know that the new paper contained the same content as the original (let alone the actual Lean code)?
Great article. My take on it is :-
"Er, can you check this for free? ... It would be great for my share price if you could, would really give the investors a badly needed shot of confidence!"
Here’s an idea I would love to see play out. Have one of the labs train a new model, using cutting edge architecture, on a whole lot of data and math papers from before 1905. Then see if it can come up with, or even understand, Einsteins theory of relativity.
Special relativity is one of those unusual theories that requires very little in the way of math or concepts to understand. You take the data from the 1887 Michelson-Morley experiment (which measures the same speed of light, despite different reference frames). You take the idea that the laws of physics work in any reference frame. You take some high-school math, et voilà, special relativity.
If there is a formal proof, and it is the proof of your precise statement (which I imagine is easy to check, otherwise I think that mathematicians would not accept the Navier Stokes result so quickly), then there is no way to ignore the result, however badly it is written.
This was always the essence of mathematics, and it will stay this way whichever statement by whomever is made.
As much as I personally despise altmans, "darios", and their bootlickers, this is one aspect which is undoubtedly "good for the mathematical community" as a whole. The fact that the validity of your statement does not depend any more on an expert opinion of some person with grants, but as it always should have had been, just on the validity of the chain of deductions.
It certainly feels that we are witnessing the early days of LLMs and math, similar to the early days of LLMs and code. I feel this entire post sounds so similar to angry computer engineers posts from 1 or 2 years ago. If we continue on this path, eventually many of these posts will ripen like milk in a desert.
It’s fine to dislike AI slop and not engage with it. But I would claim there is a difference between a human submitting a sloppy paper to a journal vs producing one with AI. The former is lazy and unprofessional, the latter is an interesting experiment. I appreciate not wanting to engage in an experiment you didn’t sign up for, but therein lies the difference between this with an attitude of embracing new technology and those wanting to stick to the old. I think Terrence Tao has had a very interesting attitude to AI recently and had also had some fruitful outcomes from it.
So we should hold LLM slop to a LOWER standard than humans? Great...
> therein lies the difference between this with an attitude of embracing new technology and those wanting to stick to the old
And as we all know, if something is new technology it must necessarily be good. Don't stop to think, just run in the rat race!
I would admit seeing more mathematicians irritated is a good popcorn show.
For what's worth it, this batch of results are not very tight intentionally by OpenAI and promising mathematicians already started to consume them and improve the results, while someone is still complaining on it.
Reading though these comments I truly understand why the hardest part of my career in "IT" has been personalities.
The lack of emotional maturity and empathy is very on par with my experience thus far.
Or there are people who don't define "standing in the way of progress" as a valid item in the list of "emotional maturity/empathy".
> OpenAI drops some hundreds of "solutions", incomprehensibly written "solutions", and we are all expected to jump on them and what? Appreciate their contributions? Why? I am trying to finish several papers, I am supervising a number of Ph.D. students, and I have a lot of active research of my own to do. When am I supposed to sift through a badly written paper? Why should I bother, when they don't bother to communicate better?
Excellent point.
I don't blame the AI for this - I blame OpenAI.
Literal slop grenade (see https://fortune.com/2026/09/17/shopify-tobias-lutke-ai-slop-...)
The mathematicians who don't approve of the deliverables should boycott the proofs, that is the only way OpenAI will get what they deserve on this one.
I see an argument about the paper being poorly written, which I believe, but I don’t see how that relates to honest final paragraph about mathematicians being chefs or whatever.
It’s clear from the author’s tone about lean, emails, infographics, etc. that he thinks automating those away is fine. Why should math be any different?
Is a proof that cannot be understood worthless? How would this be framed philosophically?
In fact, the academic system is a kind of worldview created by humans. And as it is shared and the community grows, the problem will gradually become more complex. Because when a discipline develops sufficiently, just as in a mine where rich veins are easy to extract early on but become very hard to extract once much has been dug out... in that sense, as things gradually become more complex, once a certain threshold is reached, won't scholarship surpass the limits of human understanding? Of course, scholarship is entirely for humans, but at some point the system itself may face its limits, and then wouldn't it again reduce the existing normalized minimum within that discipline and establish a new normalization of a new logical system?
In my view, perhaps for very complex work like today, AI will do it, and then there will be work that normalizes and further simplifies the results of that AI. Then, coming back to the human fold, if humans create the initial skeleton, the LLM will learn that again and it will become complex work again, and won't this create a continuing cycle?
I think verification and understanding can be separated. If the proof targets a correctly formalized proposition and passes a reliable proof checker, isn't it valuable? We have obtained knowledge justified as true, but there is simply no new theory that understands that knowledge. As was the case with the Four Color Theorem...
I am always curious what shape the newly compressed new discipline will take. At that time, I hope even people like me, who are intellectually behind, will be able to learn that discipline.
> Is a proof that cannot be understood worthless?
No.
But its worth a lot less than one that can be understood.
They really should be trying to partner with the mathematics community to add maximum value.
Their current approach is reckless and risks doing more harm than good.
If no one understands a proof, then it is not a proof
Proofs are much easier when one just lists their axioms, leaving the rest an exercise to the reader.
I've seen my entire profession vanish overnight due to AI... well, not vanish. But, yeah, software development is WAYYYY different. And I couldn't be happier. I think it's amazing. I see the productivity boost. Even if it means I can't add nearly as much value as I used to.
Sorry, but I don't get why mathematicians are so upset. Like, just accept the knowledge and insights and acceleration in your field! If it isn't "fit for human consumption" because an AI produced, okay... it soon will be explained ELI5 by even better models.
You will own nothing and you will be happy.
I don't see how this quote is relevant here.
Most software developers (or people in any field, working for anyone) rarely own anything they do at work.
And the entrepreneurs running their own companies (which there's an explosion of atm largely because of AI) do indeed "own" the higher-level products and things they're producing, even if AI writes the code.
What is supposed to have changed?
Once you surrender your cognition to a third party, you stop being a person and become livestock instead.
Today, you are training the machine to accept the poorly specified inputs of the CEO's dumb nephew tomorrow.
Did you... Enjoy software development beforehand? Like there is a huge chunk of people that really enjoyed writing software and are bummed that that part of their job is being subsumed by models.
The same could go for mathematics or any field; there are lots of people who enjoy the process and aren't satisfied by being handed and opaque final result
Loved it... love it even more now. So darn fast.
> I don't get why mathematicians are so upset
Their field is at the stage where the humans are “debugging” the AI slop.
They are still imagining how to escape from having to read the generated code. Hopefully they find a way.
I don’t get why you aren’t upset. In general I am actually quite disappointed by how meekly the software engineers handed our industry over to AI. But I am particularly revolted by those who exercise their free will to rise from the foetid swamp and say “come on in, the water’s great!”.
> OpenAI drops some hundreds of "solutions", incomprehensibly written "solutions", and we are all expected to jump on them and what? Appreciate their contributions?
This is a strawman. OpenAI didn't say they are expecting all mathematicians to read the solutions, incomprehensible or not.
So mathematicians are upset with OpenAI for solving "their" math problems. Software engineers are even more affected by AI, yet mathematicians seem to be reacting more strongly. I don't get why.
OpenAI spent $20m on this at least.
It looks like the spent $20 on writing the actual papers.
If they actually wanted to do good for the world, they wouldn't have released these as the slop grenades they are.
In their current state, they are actively damaging the mathematics community.
It shows a lack of respect and care for the impact that their technology has.
It shows that they cannot be trusted for things like private data, AI safety, and company partnerships.
In math/science, repeatability and review are critical to the process.
The right way to handle this would have been to work with the mathematics community to co-develop and create meaningful proofs rather than slop grenades.
If they proceed in the current state, we'll just get a bunch of spaghetti math that won't do anything for helping people build an understanding.
Maybe some day, we won't need people to understand things, but that's certainly not the case at the moment, and likely won't be for several more years.
Perhaps this is projection, and the staff at OpenAI doesn't understand their work anymore? Not a great sign regardless.
> The right way to handle this would have been to work with the mathematics community to co-develop and create meaningful proofs rather than slop grenades.
The mathematics community can finish the job OpenAI started. Or are you saying the community has no incentive to do that because there is no reward/recognition for doing that?
> the job OpenAI started.
They didn't just "start" it though - they released slop papers.
That's finishing it - not starting it as far as scientific publishing is concerned.
> Or are you saying the community has no incentive to do that because there is no reward/recognition for doing that?
Not exactly - but that is part of it.
I think what would have went over better is:
1. Immediately announce a solution has been found.
2. Do not publish the solution.
3. Put out an open request for anyone with experience in the area who wants to get involved to help collaborate on a construction and human-comprehensible paper. Accept anyone who can demonstrate potentially useful work/experience in the field/problem. Share the solution with them after they sign some kind of NDA that they won't independently publish or share the solution/work.
4. Work with people until a paper is ready (I mean actually ready - not the kind of slop that they released).
5. Publish. Include names of everyone who made meaningful contributions to the paper (not just the proof).
EDIT: Notice the incentive with my proposed second path is that it gives OpenAI an incentive to improve the interpretability of its proofs. This is a good thing! The maths community would be thrilled to actually gain understanding from such releases, and OpenAI would be happy because they could more quickly and independently publish their results. At the moment the "value" of their mathematics research "product" is low because of the lack of this interpretability, and this current approach is simultaneously destroying the opportunity value of the community as well as the incentive for OpenAI to ever improve on what's missing.
or they just release as it is and keep on improving until they can one shot human comprehensible paper.
No, you're just straight up wrong.
OpenAI is doing science the right way. Making the information available as widely as possible so that anyone can check and verify it.
That is the scientific process working exactly as it should.
If that "damages the mathematics community" then all it means is the mathematics community is not doing science and should be ignored.
I cannot stress this enough. If you are complaining about "how they released it" or calling this stuff "slop cannons", you're an unscientific hack that's dragging down humanity.
Engage with the actual claims. Prove them or disprove them. Nothing else matters here.
This. So much this. (and I dont even like OAI)
This is great - I think we're getting close to the core disagreement here because I disagree with almost all of that, and I think its because we have different conceptions of what science is.
To me, science is not a database of facts and data, and the scientific process is not just the process of adding to this database.
It is more than that.
Science is building the database of facts and data, but it is also sharing/disseminating that database so that others can write to it too.
We both agree that "openness" is critical for this, but to me, the scientific process includes things like the actual words that a scientist uses to communicate their ideas/data, the paper they wrote, the collaborations they had, the work that they built upon. Science is not like a database only - its more like a distributed system, and the scientific community is like a living, breathing thing. It is the health of that system which has been damaged in my opinion.
So I have no problem that OpenAI updated the database.
They didn't just make the information available though.
They updated the knowledge database, then released low quality slop papers.
They made it clear that they don't give a damn about sharing/disseminating information.
They are trying to fundamentally say "we can do science without prioritizing sharing.
We don't think its important. We'll release the proof because it makes our stock go up, but we're gonna keep going and building and not spend any time on communication or comprehension because we think that it doesn't matter, and your contributions don't matter, and AI is gonna do all of this independently anyway. We spent $20m on the proof because it made our stock go up $20b, but we'll spend 20 min on the paper. Priorities clear.
Does any of that make sense?
I don't disagree that having comprehensible papers is useful. There is undoubtedly value in having well written, easy to understand documents that make it easier for third parties to validate the work presented.
The thing is, that is a "nice to have", not a condition to make science.
If people believe that such things are useful, then the appropriate response is to reproduce, simplify and otherwise make the proofs provided into more easily digestible pieces.
The scientific process is fundamentally indifferent to whatever shibollets and cultural hangups that its presumed participants may have. An alien with fundamentally different norms from us, who experiments to determine the boiling temperature of water is doing science just as much as a human, even if he never submits a paper to a journal.
The issue you're having is that you're under the mistaken belief that the scientific process is the "community" and "norms" that happen to have developed over the past few centuries by this "scientific" priesthood.
It is not. It has never been. It will never be.
This "scientific" priesthood has simply used their power and reputation to enforce their own self-aggrandizing norms, much like the religious priesthood they replaced. They should not be taken seriously, and to go along with these norms when they are unnecessary is frankly ridiculous and shameful for anyone who actually prizes results.
I would urge you to reconsider your stance, and in particular to put deep thought into why you believe these norms are so important.
> Indeed, at least two people asked if I plan on sending Sam Altman a bottle of whisky, as promised in my Problems page. The answer to that is no. And let me explain to you why
> I took a brief look at the preprint released by OpenAI. It sucked. It is unclear, muddled, and has a strange structure.
https://img.getfn.io/images/e3245d47a159592b06570ffbd64d5af8...
If OpenAI actually spent time writing good proofs that are readable, mathematicians would show *even more outrage*. So let’s not pretend this has anything to do with readability. It has all to do with people protecting their status in society.
Stereotyping mathematicians, well played sir.
You just wait. This will come :)
> if this was an academic paper submitted to a journal, it should be issued a desk rejection for the quality
If Mr Tao had submitted a shitty, poorly-written proof of a famous outstanding problem, no journal would reject it. That extends to anyone with sufficient credibility. They might ask him to keep at it and fix it up, but nobody would begrudge him putting his shitty (but ultimately correct) draft of arXiv while he did so.
We're all getting disrupted, we all have feelings about it, but from the perspective of a software engineer who's been dealing with all of this for several years now, this post is just cope.
Wow talk about sour grapes! No one is forcing you or any other mathematician at gun point to engage with this release at all. Ridiculous drivel.
I just read a childish rant
This whole "oh no this problem is solved now who would ever want to work on it" thing is quite funny. Mathematicians, let me introduce you to something called bike shedding. Folks proved 1s and 0s are turing complete decades ago and yet we have 27 new JavaScript web frameworks every week (or every day or minute now with ai). Y'all will be fine. Thinking a solution is ugly and that you could do better is also a perfectly good motivator and most of the sciences and engineering are "ok, so we know xyz to be true about the world because like, I'm looking at it, but wtf is going on." The theorems were true false or otherwise before some random openai model solved them. And if you don't understand the proof nothing of significance has changed except you've got a bit of a hint now.
Cathedral and Bazaar. Maths professors are used to working diligently behind closed doors before releasing artifacts of high quality whose authorship they guard jealously. The researchers at OpenAI, who come from a software background, are used to working in the open, releasing anything to anyone and expecting nothing but also guaranteeing nothing.
The amount of time the OP attacks OpenAI for errors of form and not substance is unfortunate. Do they also attack amateurs who try to contribute like this?