Last updated: 2026-10-11

U
Undergraduate level

When the AI Becomes the IDE: Mathematics, Programming, and the Economics of Fulfilling Work

11 October 2026

On 6 October 2026, OpenAI published more than seven hundred manuscripts in a single release, organised into 372 result families and claiming solutions across number theory, geometry, operator algebras and a dozen other fields, several of them open for decades. Roughly forty per cent of the headline results carried a Lean formalisation; the rest arrived as conventional manuscripts, resting on the model's own account of its reasoning rather than on anything a proof checker had verified[1]. Within a day the Association for Human Mathematics called the release "a demonstration of power" rather than scholarship, and urged mathematicians to stop working with the company[2]. A month earlier, twenty-five Fields medallists, Terence Tao among them, had already signed a declaration warning that AI labs and mathematicians wanted structurally different things from the same open problems[3]. This essay does not try to referee those 372 claims one by one. It tries to work out what the argument over them is actually about, because the dispute underneath it reaches well past mathematics.

Tao's specific worry, posted a few weeks before the October release, was narrower and more interesting than "AI is cheating": "We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential."[4] He had a live example close at hand. On 8 September 2026, OpenAI announced that an internal model, run as roughly ten thousand coordinating agents over 88 hours, had produced a result on the Navier-Stokes existence and smoothness problem, one of the Clay Mathematics Institute's seven Millennium Prize problems[5]. Hours earlier, the mathematicians Tristan Buckmaster and Levent Alpรถge had released related work of their own, built up over months; Buckmaster later asked publicly whether OpenAI's system had been shaped by seeing the approach the pair had developed[6]. OpenAI, for its part, declined to claim the Millennium Prize for the result, which is itself a plain enough signal of how far even the company rates its own claim from a closed case[5].

None of this is an argument that OpenAI should have been stopped from working on the problem. A correct, useful mathematical result is a public good whichever kind of mind produces it, and treating "a human should get there first" as a value in itself confuses the result with the race to it. The actual failure sits one level down, in how mathematics currently assigns credit. Posing a problem well enough that it becomes tractable is a contribution. Refining a vague question into a sharp conjecture is a contribution. A partial result, a disproof of a special case, a cleaner formalisation of an existing proof: all contributions. So is the checking, explaining, simplifying and applying that happens after a proof exists, work that in this case somebody still has to do by hand, since most of OpenAI's own release remains unformalised[1]. A system built almost entirely around being first to publish the final proof was always going to strain once something could out-race a career's worth of the slower contributions around it, and that strain, not the AI itself, is what actually needs fixing.credit is the currency of the human game

Fulfilment Is Not the Same Question as Employment

A familiar objection follows from the Navier-Stokes case: if AI keeps clearing open problems faster than mathematicians can reach them, won't fewer people bother training as mathematicians at all? Perhaps, for anyone whose only reason to do mathematics was the prospect of being one of the few who could. That motivation was fragile well before any of this; automation has simply made the fragility visible rather than created it. People keep running long after cars made running pointless as transport, keep playing chess against opponents they can never beat, keep making music nobody needs a human to perform. Doing something well, for its own sake, has never depended on being the best available means of getting the result.

What automation does threaten is not mathematical curiosity but mathematical income, and those are different problems with different fixes. Curiosity does not pay rent. Fulfilling work has rarely paid reliably either, which is easy to forget because the exceptions are loud: a handful of musicians, athletes and performers earn enormous sums, and their visibility creates an impression that talent and effort in those fields generally convert to a living. Spotify's own "Loud & Clear" accounting of 2020 royalties shows the more typical shape: the top 0.8 per cent of artists on the platform received around 90 per cent of everything paid out, and only 13,400 of the roughly eight million artists with music on the service earned as much as $50,000 that year[7]. Markets like this reward attention, scarcity and the ability to be sold tickets, subscriptions and merchandise; they do not reliably reward difficulty, skill or social value, and a mathematician proving a theorem that cannot easily be bundled and resold has always sat on the wrong side of that mismatch. None of this is a complaint about musicians. It is a reason not to trust income as a measure of either merit or usefulness, in mathematics or anywhere else.UBI may be needed to decouple income from work

Programming Was Never Mainly Typing Code

The same confusion between an activity's economic footing and its intellectual substance shows up again, more directly, in programming. Good programmers are not made by typing code quickly. They are made by working out what a problem actually is: finding the requirement nobody wrote down, deciding what "finished" would even mean, choosing an abstraction that will still make sense in six months, splitting an unmanageable problem into pieces that are not, noticing the edge case that breaks an otherwise clean design, and explaining the result to a colleague, a user or a manager who each need a different version of the same explanation.

Code was never irrelevant to that work; it just was never the whole of it. A specification that says a system should notify "the user" when a payment fails looks complete until someone has to implement it, at which point building the thing exposes at least three separate questions the sentence never answered: whether "the user" means the customer who made the payment, the merchant whose system declined it, or an administrator watching for fraud; what counts as a failure rather than a delay; and what should happen if the notification itself cannot be delivered. None of those questions were visible in the requirement. All three become visible the moment somebody has to write code that actually runs, because code refuses to stay vague in the way a document can. That forcing function, not the syntax itself, is what implementation was always for.this is why specs fail in practice

When the AI Becomes the IDE

This is the sense in which an AI system, used seriously, can take over an IDE's role as the thing that forces precision, rather than simply taking over the typing. An integrated development environment catches malformed syntax, type mismatches and a handful of other formal inconsistencies; useful, but narrow. A language model pushed to produce something useful can be made to expose the same kind of gap the payment example hides, except in plain language rather than a compiler error: what do you mean by "user" here? Which requirement wins when two of them conflict? What should happen if the external service this depends on is unavailable? Is this explanation meant for the engineer who will maintain the code or the manager who will decide whether to fund it? Is the simplification you are asking for allowed to drop a qualification that actually matters?the IDE checks syntax; the AI checks intent

Calling this "prompt engineering" undersells what is happening. Getting a useful result out of the exchange means stating the intended outcome, the real constraints, which assumptions are safe to make, what would count as success, which edge cases matter, how confident the result needs to be, who it is for, and what evidence would convince a sceptical reader it is actually right. That is specification and modelling, carried out in conversation rather than in a requirements document, and it draws on exactly the audience-awareness and critical judgement that distinguished a good requirements analyst from a mediocre one long before any of this software existed.

Precision Without Syntax

It is tempting to read the move from code to conversation as a move away from precision, since natural language tolerates ambiguity no compiler would accept. That gets the mechanism backwards. Natural language permits ambiguity; it does not require it, and serious work with an AI system creates constant pressure to find and close the gaps a looser conversation would have left open. There are two different kinds of precision in play here, and conflating them is the mistake underneath most of the complaint. Syntactic precision asks whether a statement is well formed in whatever grammar a machine accepts. Semantic and pragmatic precision asks something broader: whether purpose, circumstances, constraints and the conditions for success have been communicated clearly enough that the right result can be produced and then checked. Code historically forced the first kind and left the second to whatever conversation happened to surround it. Conversation with an AI system still needs the first kind wherever formal checking genuinely helps (tests, types, verification and executable contracts have not stopped being useful), but it now forces the second kind too, across a much larger share of the work than before.

The Loop That Builds Expertise

Before generative AI, explaining a solution and implementing it were usually two separate stages, done at different speeds. Increasingly, a sufficiently clear explanation can start the implementation directly: describe, generate, inspect, challenge, clarify, test, revise. That loop sits structurally close to the one programmers already ran with a compiler and a debugger: problem, hypothesis, implementation, failure, diagnosis, revised model. AI can run the new loop faster than the old one ran, but speed alone does not decide whether the person inside it is thinking.

What decides that is whether the loop is being driven or merely watched. Someone who asks an AI system to implement an idea, checks the assumptions behind it, predicts where it will fail, actually runs the tests, challenges a wrong or incomplete answer and revises the underlying model is doing the same work a programmer has always done, just through a different interface. Someone who accepts the first fluent answer gets an artefact and very little else. The risk this creates for learning is not that the code came from a model rather than a keyboard. It is that a learner can quietly stop owning the problem and start mistaking a plausible answer for a correct one, a much older failure than AI wearing a more convincing disguise.fluent !: correct; check the logic

Used the other way, the same tool can strengthen exactly the attention it threatens to erode: trying several competing designs in the time it used to take to build one, getting an explanation pitched at the right level instead of the nearest available textbook's level, generating tests to inspect rather than trusting by default, and practising the specific skill of adapting one explanation for two audiences, a terse account of a race condition for the engineer who will fix it and a plainer one, about two processes both reaching for the same resource at once, for the manager who only needs to know why the fix will take a day rather than an hour. Writing the argument above involved a version of the same loop. An early draft treated the mathematics controversy as the whole subject; testing that draft against what actually happens to mathematicians' careers, and then against programming, is what exposed how much of the real argument was still underneath it.

Abundance Without Security

If mathematics and programming genuinely need much less paid human labour to produce a given result, that need not be decline. It could mean more people able to build real software without years of training first, wider access to mathematical exploration, faster testing of ideas, and more time freed from routine production for whatever someone actually wants to do with a day. None of that follows automatically, and none of it answers the harder questions sitting underneath the productivity gain. Who owns the systems producing it? Who receives the gains once labour's share of the cost falls? Who gets credited for the public mathematics and public code those systems were trained on in the first place? And why should the disappearance of scarce, paid labour reduce anyone's access to the time and security that labour used to buy, rather than simply changing who provides it?

These are not questions this essay can settle, and treating every existing programming or mathematics job as something owed permanent protection would dodge them rather than answer them. A profession does not get to keep its current shape just because its members would prefer that it did. But a transition that quietly lets ownership, credit and income concentrate around whoever controls the AI system, while leaving everyone else to rediscover fulfilling work as an unpaid hobby, is not a neutral outcome either; it is a specific political choice, dressed up as an inevitability. A related page on this site works through the data-centre and automation economics behind that choice, and the early, partial evidence for how an income floor might work, in more detail than fits here.the political economy of the build-out

What Should Actually Be Preserved

Nothing above is an argument against AI doing mathematics, writing code, or doing either faster than any person could. A correct, useful result is valuable whoever or whatever produced it, and pretending otherwise mainly protects the comfort of whoever got there first, not any value a reader should actually care about. The argument is narrower, and more useful: programming was never principally the production of code, and proving a theorem was never principally the moment of first publication. Both were always the slower, harder work of turning an unclear situation into something precise enough to test, built by whoever was doing the posing, the refining, the partial results, the checking and the explaining along the way, long before any single name got attached to the result.

That work has changed its interface. It increasingly happens through conversation with a system that can generate as fast as it can be questioned, rather than through the older, slower negotiation with a compiler, a test suite or a research seminar. What should survive the change is the habit underneath it: stating a problem precisely enough that someone, or something, can be held to a real answer, and then actually checking that answer rather than taking it on trust. What should not survive unexamined is a set of institutions, in mathematics, in software, in how either pays for a living, still built around who typed the first line or filed the first manuscript, in a world where that is no longer the scarce or the interesting part of the work.

References

  1. TechCrunch (2026). "OpenAI's math solutions aren't meeting the field's standards yet." 8 October 2026. techcrunch.com/2026/10/08/openais-math-solutions-arent-meeting-the-fields-standards-yet
  2. Association for Human Mathematics (2026). Statement on OpenAI's 6 October 2026 release of mathematical documents, republished by Terence Tao. What's New, 7 October 2026. terrytao.wordpress.com/2026/10/07
  3. Tao, T. et al. (2026). "A Severe Misalignment of AI in Mathematics." Declaration signed by twenty-five Fields Medal winners. What's New, 11 September 2026.
  4. Tao, T. (2026). Mathstodon post, 8 September 2026, quoted in Lanz, J. A., "AI Is Solving Math's Best Problems Faster Than They Can Be Replaced, Terence Tao Warns." Decrypt, 9 September 2026. decrypt.co/377818/ai-math-best-problems-terence-tao
  5. OpenAI (2026). "Navier-Stokes existence and smoothness." openai.com, 8 September 2026. openai.com/index/navier-stokes-solution
  6. TechCrunch (2026). "OpenAI's feud with mathematicians is only escalating." 11 September 2026.
  7. Spotify (2021). "Loud & Clear." Reporting 2020 royalty data. Spotify Newsroom, March 2021.