I share your scepticism that a 'next word predictor' (Carbone, 2025) is not likely to be trustworthy when it comes to maths. By their very nature, large language models are not 'capable' to logical reasoning (the increasingly common and powerful 'reasoning models' are still statistical), but how to bring this to students' attention?
For students doing advanced maths, for example, project work, Carbone's (2025) pre-print may be an interesting and highly relevant read for students as well as teachers and supervisors. It describes how mathematics is incorporated into generative AI (most of generative AI's apparent 'abilities' with maths are feathers unashamedly borrowed by outsourcing many mathematical tasks to specialised deterministic programs and Python libraries), and how generative AI can be used productively for mathematics research (I get the impression this is Carbone's experience). According to Carbone (2025), the main gains of using generative AI in advanced (research) mathematics are for exploring the feasibility of apparently intractable problems, identifying relevant methods from other areas that are needed to address said problems, and for writing the code for non-AI programs to implement (or to be outsourced thereunto).
Features such as the finite 'memory' of a generative AI session and various settings that can be adjusted are worth knowing about. Practical guidance on prompting, and warnings that one must prompt in small steps, check every line of the output, and that the entire 'problem' needs to completed within a single memory window can be particularly relevant to emphasise. The bottom line is that the time required to do a complex calculation with generative AI is likely to be no less than that required to do the work without generative AI. If one then adds in that any success requires resource greedy 'reasoning' models, and multiple prompts, that true, logical reasoning is completely absent, that the training data includes retracted articles, typographical errors, multiple conventions (i or j may be used for SQRT(-1), among other variations, and we have no real idea what the training data is or any quality control, there seems to be a strong case for students to do the work themselves, with pen and paper, no matter how hard it is, rather than risk wasting a week or more on work that they then have to redo with pen and paper with full knowledge that they are only progressing by logical steps.
I suspect that LLMs' lack of transparency and traceability of methods that is counter to any sort of scientific practice highlight several pre-exisiting weakness in how maths is often presented or talked about -- or rather not talked about. Being told 'it is obvious [from this equation] that the result is [something not obvious if one cannot identify and do the relevant calculation in 10s]' was bad enough before; now LLMs do that without even allowing one to assume that the derivation has actually been done (for examples of good practice when it comes to not presenting all the steps of a derivation, see Landau and Lifshitz, 1960/1976).
So much for why advanced mathematics is probably not (yet, maybe never) worth doing any other way than by hand. What about less advanced mathematics?
Up to the introductory undergraduate level (differential and integral calculus, linear algebra, and various classical physics topics) Generative AI models can perform as well as students, but with a narrower distribution around 70%, the typical grade boundary for a first class degree in the UK (Walker et al., 2025). However, this is with material that is typically pretty easy to find on the internet and is therefore likely to be present in LLMs training data (Carbone, 2025). Despite the small number of research publications, reviews abound, which is useful for gaining an overview of the potential impact of generative AI programs in mathematics education, though no where near conclusive. For example, Walkington (2025) reports that the impact of students using generative AI in learning maths is very dependent on the study, and for creating adaptive or tailored learning paths (note, these capabilities have existed before in curated apps and programs). The additional 'capability' that comes with LLMs is that problems (practice calculations) can be adapted to versions that are contextually relevant to students, but here it can generate nonsensical problems or create problems with blatantly unreasonable numbers, so the most use here is for teachers developing questions that they then review.
However, although generative AI programs can 'guide' students through solving mathematical problems, that students who use generative AI to support their maths learning can have markedly less confidence in their own abilities (cited in Walkington, 2025) is deeply worrisome. Maths is cognitively demanding and has a reputation for difficulty, yet generative AI may prevent the full development of those abilities in many students (which is quite ironic since all computing, including LLMs, is built on mathematics). The impacts of students potentially becoming reliant on LLMs for maths (learning) may have significant impacts on advanced mathematics and mathematics research (although, as Carbone (2025) writes, there are uses for research progress). While this may be starting to become anecdotally available now, I suspect the full scale and impact will not be apparent for another few years. Maths is the discipline with the second highest 'expectation of brilliance' (only exceeded by philosophy; Leslie et al., 2015), and I think every maths teacher — or acquaintance thereof — should be raising questions about the impact of students using generative AI to support maths learning or solve maths problems on the future state, progress, diversity and inclusivity of mathematics.
Resources:
Carbone, L. (2025). Advancing mathematics research with generative AI. arXiv preprint arXiv:2511.07420. https://arxiv.org/abs/2511.07420v2
Landau, L. D. and Lifshitz, E. M. (1996 [1960/1976]), Course of Theoretical Physics, Volume 1, Mechanics, 3rd Edition. (Any volume, or the 'Shorter Course' will show how they provide very brief explanations of the steps that are not shown.)
Leslie, S.-J., Cimpian, A., Meyer, M. & Freeland, E. (2015). Expectations of brilliance underlie gender distributions across academic disciplines. Science 347, 262-265 https://doi.org/10.1126/science.1261375
Walker, B. J., Kalaydzhieva, N., Lameda, B. N., & Reynolds, R. A. (2025). Evaluating undergraduate mathematics examinations in the era of generative AI: a curriculum-level case study. arXiv preprint arXiv:2509.13359. https://arxiv.org/abs/2509.13359
Walkington, C. (2025). The implications of generative artificial intelligence for mathematics education. School Science and Mathematics, 1–10. https://doi.org/10.1111/ssm.18356