Skip to main content
AiGORA
Back to Q&A
Text
12

How good is generative AI at maths?

I am highly sceptical of generative AI and don't use it myself, but I know there's a general impression among students that it 'knows everything'. The discussions I have seen have focused on text (logical enough for 'large language models'), and images (many of which are disastrous), but I cannot see how LLMs can reliably do (complex) mathematics where there are many steps to a derivation, some of which are nuanced or require particular systematic techniques, and the final result should be as general as possible. How can I convince students that they will (probably) waste more time using generative AI than working through the difficulties themselves? Or do I have to resign myself to allowing an extra week or more for students to learn for themselves that there really isn't an alternative to putting in the effort and doing the derivation oneself with pen and paper, no matter how hard it is?

Anon · 19 Jul 2026

Responses from the team

3 perspectives from the community

  • Kirsty Dunnett

    I share your scepticism that a 'next word predictor' (Carbone, 2025) is not likely to be trustworthy when it comes to maths. By their very nature, large language models are not 'capable' to logical reasoning (the increasingly common and powerful 'reasoning models' are still statistical), but how to bring this to students' attention?

    For students doing advanced maths, for example, project work, Carbone's (2025) pre-print may be an interesting and highly relevant read for students as well as teachers and supervisors. It describes how mathematics is incorporated into generative AI (most of generative AI's apparent 'abilities' with maths are feathers unashamedly borrowed by outsourcing many mathematical tasks to specialised deterministic programs and Python libraries), and how generative AI can be used productively for mathematics research (I get the impression this is Carbone's experience). According to Carbone (2025), the main gains of using generative AI in advanced (research) mathematics are for exploring the feasibility of apparently intractable problems, identifying relevant methods from other areas that are needed to address said problems, and for writing the code for non-AI programs to implement (or to be outsourced thereunto).

    Features such as the finite 'memory' of a generative AI session and various settings that can be adjusted are worth knowing about. Practical guidance on prompting, and warnings that one must prompt in small steps, check every line of the output, and that the entire 'problem' needs to completed within a single memory window can be particularly relevant to emphasise. The bottom line is that the time required to do a complex calculation with generative AI is likely to be no less than that required to do the work without generative AI. If one then adds in that any success requires resource greedy 'reasoning' models, and multiple prompts, that true, logical reasoning is completely absent, that the training data includes retracted articles, typographical errors, multiple conventions (i or j may be used for SQRT(-1), among other variations, and we have no real idea what the training data is or any quality control, there seems to be a strong case for students to do the work themselves, with pen and paper, no matter how hard it is, rather than risk wasting a week or more on work that they then have to redo with pen and paper with full knowledge that they are only progressing by logical steps.

    I suspect that LLMs' lack of transparency and traceability of methods that is counter to any sort of scientific practice highlight several pre-exisiting weakness in how maths is often presented or talked about -- or rather not talked about. Being told 'it is obvious [from this equation] that the result is [something not obvious if one cannot identify and do the relevant calculation in 10s]' was bad enough before; now LLMs do that without even allowing one to assume that the derivation has actually been done (for examples of good practice when it comes to not presenting all the steps of a derivation, see Landau and Lifshitz, 1960/1976).

    So much for why advanced mathematics is probably not (yet, maybe never) worth doing any other way than by hand. What about less advanced mathematics?

    Up to the introductory undergraduate level (differential and integral calculus, linear algebra, and various classical physics topics) Generative AI models can perform as well as students, but with a narrower distribution around 70%, the typical grade boundary for a first class degree in the UK (Walker et al., 2025). However, this is with material that is typically pretty easy to find on the internet and is therefore likely to be present in LLMs training data (Carbone, 2025). Despite the small number of research publications, reviews abound, which is useful for gaining an overview of the potential impact of generative AI programs in mathematics education, though no where near conclusive. For example, Walkington (2025) reports that the impact of students using generative AI in learning maths is very dependent on the study, and for creating adaptive or tailored learning paths (note, these capabilities have existed before in curated apps and programs). The additional 'capability' that comes with LLMs is that problems (practice calculations) can be adapted to versions that are contextually relevant to students, but here it can generate nonsensical problems or create problems with blatantly unreasonable numbers, so the most use here is for teachers developing questions that they then review.

    However, although generative AI programs can 'guide' students through solving mathematical problems, that students who use generative AI to support their maths learning can have markedly less confidence in their own abilities (cited in Walkington, 2025) is deeply worrisome. Maths is cognitively demanding and has a reputation for difficulty, yet generative AI may prevent the full development of those abilities in many students (which is quite ironic since all computing, including LLMs, is built on mathematics). The impacts of students potentially becoming reliant on LLMs for maths (learning) may have significant impacts on advanced mathematics and mathematics research (although, as Carbone (2025) writes, there are uses for research progress). While this may be starting to become anecdotally available now, I suspect the full scale and impact will not be apparent for another few years. Maths is the discipline with the second highest 'expectation of brilliance' (only exceeded by philosophy; Leslie et al., 2015), and I think every maths teacher — or acquaintance thereof — should be raising questions about the impact of students using generative AI to support maths learning or solve maths problems on the future state, progress, diversity and inclusivity of mathematics.

    Resources:

    Carbone, L. (2025). Advancing mathematics research with generative AI. arXiv preprint arXiv:2511.07420. https://arxiv.org/abs/2511.07420v2

    Landau, L. D. and Lifshitz, E. M. (1996 [1960/1976]), Course of Theoretical Physics, Volume 1, Mechanics, 3rd Edition. (Any volume, or the 'Shorter Course' will show how they provide very brief explanations of the steps that are not shown.)

    Leslie, S.-J., Cimpian, A., Meyer, M. & Freeland, E. (2015). Expectations of brilliance underlie gender distributions across academic disciplines. Science 347, 262-265 https://doi.org/10.1126/science.1261375

    Walker, B. J., Kalaydzhieva, N., Lameda, B. N., & Reynolds, R. A. (2025). Evaluating undergraduate mathematics examinations in the era of generative AI: a curriculum-level case study. arXiv preprint arXiv:2509.13359. https://arxiv.org/abs/2509.13359

    Walkington, C. (2025). The implications of generative artificial intelligence for mathematics education. School Science and Mathematics, 1–10. https://doi.org/10.1111/ssm.18356

    Read all posts that Kirsty has responded to →

  • Mirjam Glessmer's profile photo

    Mirjam Glessmer

    My expertise is certainly not in mathematics, but I found this recent substack post "Are we just collecting lots of AI-proof-shaped stamps?" by Adam Kucharski very interesting: Even if AI can be used to generate proofs of problems that have puzzled mathematicians for centuries (and there are apparently examples of this happening), what is the actual value of having said proof if nobody can explain how to get there? Can the field still build on it and integrate it in the accepted body of knowledge without being able to explain it in the traditional ways of however people are usually working with proofs? Maybe mathematics is about to face a paradigm shift, who knows. Or maybe AI-proofs are going to be rejected until someone can explain the actual logic behind them, step-by-step. I think these might be interesting questions to discuss with students, and then also in extension what it means for them themselves to use AI to generate proofs...

    Read all posts that Mirjam has responded to →

  • Jan-Fredrik Olsen

    Generative AI is not only good at mathematics, it is good enough to be a valuable help for professional mathematicians. For more on this, see the blog of Timothy Gowers and a recent essay by Terence Tao (links below). Both are highly respected in the mathematical community. However, generative AI is not reliably good. While expert users can use generative AI to prove theorems and solve problems beyond what was possible before generative AI, this makes it a problematic tool for non-expert users. Indeed, since generative AI is probabilistic (it does a little internal lottery for every word it outputs) it will make mistakes every now and then. This is not a big deal for experts, since they will be able to spot (or at least suspect) that something is off. But for novices, they can easily be misled and confused. I am not an expert on the technical reason why generative AI is as good at mathematics as it is, but it is related to the fact that mathematics is a highly structured and logical language -- just like programming languages, which generative AI is even better at. However, a reason generative AI may be a bit worse at mathematics than programming, is that its training data contains many executable programs, while mathematics is often written up in (partially) informal ways (and there is a lot of bad maths online). In recent years, a programming language called LEAN has been developped, and it actually represents mathematical arguments as programs. That is, when you first start using LEAN, you can program axiomatic proofs of simple statements by writing down the axioms you use, line by line, and connecting them with logical statements. You can then store such 'programs' and re-use them when making programs that prove more advanced results. These 'programs' are exactly the propositions, theorems and lemmas of mathematics. When using LEAN, the reliability with which generative AI can do mathematics seems to increase dramatically. But what does this say about teaching mathematics in school? I am actually an optimist. As I said above, the power of generative AI for doing mathematics is (still) linked to whether the AI is collaborating with a human mathematician. So, while I think role and status of mathematics may change in the future, I don't think humans will be taken out of the equation. And the humans will still need to have a solid manual and 'old school' mastery of mathematical concepts, techniques and ideas in order to get the most out of the AI (and vice versa, one could say). From this perspective, I think the more interesting question is: how can we use generative AI to have humans learn mathematics in a better, more robust and interesting way? I am myself teaching introductory calculus at Lund University, allowing free use of generative AI by students (even on homeworks and projects), and getting better results than ever on completely AI-free traditional closed book exams and on one hour theoretical oral exams. I think the trick is that by carefully changing the course structure, and course activites, I have been able to get students to use generative AI to experiment with mathematics, and 'see' much further than what they could before (learning mathematics has always been a bit like walking in a dense fog -- it is very hard to look ahead). In this way, students can become exited about mathematics they do not yet fully understand (just as students can in other subjects - such as in physics and in music), and then make it their business to actually do the hard work necessary to develop the 'manual' mathematical mastery of having the please of understanding and being able to grasp these advanced mathematical concepts and ideas themselves. The exact details of what I do are too long to explain here -- and I hope to publish more about it elsewhere. In the meantime, please contact me for addition details. Below, I provide links to course evaluations of the courses I have taught in this way. They are quite detailed, and at the very end, you can see free-text responses on what students appreciated and what they thought could be improved. Note that the most successful was One Variable Calculus in autumn 2024. The course taught in spring 2025 was a completely new course, which also had some "teething" problems. The evaluation for the same course from spring 2026 will appear soon, and this fall I will be teaching the One Variable course using the same AI-integrated approach.

    One Variable Calculus - Autumn 2024: https://www.maths.lu.se/fileadmin/maths/Matematik_NF/Kursutvaerderingar/HT2024/MATA31HT24.pdf

    Introduction to higher analysis - spring 2025: https://www.maths.lu.se/fileadmin/maths/Matematik_NF/Kursutvaerderingar/VT2025/MATB33VT25.pdf

    Blog of Tim Gowers: https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-are-llms-good-at/

    Essay of Terry Tao: https://arxiv.org/pdf/2608.16753

    Read all posts that Jan-Fredrik has responded to →

Comments

Share your thoughts — comments are reviewed before they appear

Jan-Fredrik Olsen3d ago

Generative AI is not only good at mathematics, it is good enough to be a valuable help for professional mathematicians. For more on this, see the blog of Timothy Gowers and a recent essay by Terence Tao (links below). Both are highly respected in the mathematical community. However, generative AI is not reliably good. While expert users can use generative AI to prove theorems and solve problems beyond what was possible before generative AI, this makes it a problematic tool for non-expert users. Indeed, since generative AI is probabilistic (it does a little internal lottery for every word it outputs) it will make mistakes every now and then. This is not a big deal for experts, since they will be able to spot (or at least suspect) that something is off. But for novices, they can easily be misled and confused. I am not an expert on the technical reason why generative AI is as good at mathematics as it is, but it is related to the fact that mathematics is a highly structured and logical language -- just like programming languages, which generative AI is even better at. However, a reason generative AI may be a bit worse at mathematics than programming, is that its training data contains many executable programs, while mathematics is often written up in (partially) informal ways (and there is a lot of bad maths online). In recent years, a programming language called LEAN has been developped, and it actually represents mathematical arguments as programs. That is, when you first start using LEAN, you can program axiomatic proofs of simple statements by writing down the axioms you use, line by line, and connecting them with logical statements. You can then store such 'programs' and re-use them when making programs that prove more advanced results. These 'programs' are exactly the propositions, theorems and lemmas of mathematics. When using LEAN, the reliability with which generative AI can do mathematics seems to increase dramatically. But what does this say about teaching mathematics in school? I am actually an optimist. As I said above, the power of generative AI for doing mathematics is (still) linked to whether the AI is collaborating with a human mathematician. So, while I think role and status of mathematics may change in the future, I don't think humans will be taken out of the equation. And the humans will still need to have a solid manual and 'old school' mastery of mathematical concepts, techniques and ideas in order to get the most out of the AI (and vice versa, one could say). From this perspective, I think the more interesting question is: how can we use generative AI to have humans learn mathematics in a better, more robust and interesting way? I am myself teaching introductory calculus at Lund University, allowing free use of generative AI by students (even on homeworks and projects), and getting better results than ever on completely AI-free traditional closed book exams and on one hour theoretical oral exams. I think the trick is that by carefully changing the course structure, and course activites, I have been able to get students to use generative AI to experiment with mathematics, and 'see' much further than what they could before (learning mathematics has always been a bit like walking in a dense fog -- it is very hard to look ahead). In this way, students can become exited about mathematics they do not yet fully understand (just as students can in other subjects - such as in physics and in music), and then make it their business to actually do the hard work necessary to develop the 'manual' mathematical mastery of having the please of understanding and being able to grasp these advanced mathematical concepts and ideas themselves. The exact details of what I do are too long to explain here -- and I hope to publish more about it elsewhere. In the meantime, please contact me for addition details. Below, I provide links to course evaluations of the courses I have taught in this way. They are quite detailed, and at the very end, you can see free-text responses on what students appreciated and what they thought could be improved. Note that the most successful was One Variable Calculus in autumn 2024. The course taught in spring 2025 was a completely new course, which also had some "teething" problems. The evaluation for the same course from spring 2026 will appear soon, and this fall I will be teaching the One Variable course using the same AI-integrated approach. One Variable Calculus - Autumn 2024: https://www.maths.lu.se/fileadmin/maths/Matematik_NF/Kursutvaerderingar/HT2024/MATA31HT24.pdf Introduction to higher analysis - spring 2025 https://www.maths.lu.se/fileadmin/maths/Matematik_NF/Kursutvaerderingar/VT2025/MATB33VT25.pdf Blog of Tim Gowers (https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-are-llms-good-at/) Essay of Terry Tao (https://arxiv.org/pdf/2608.16753)