AI工具Score B (51)
OpenAI mistranslated mathematics into code for its Navier-Stokes proof | New Scientist
1 小时前2 viewsSource: newscientist.com
AI-generated proofs are often checked using a process called formalisation Pixels Hunter/Shutterstock OpenAI appears to have made a subtle error when publishing its proofs of the Navier-Stokes problem, a team of mathematicians has claimed. The error doesn’t mean that the proofs are incorrect or that OpenAI hasn’t correctly solved the problem, but it does call into question whether mathematical results generated by AI models can always be relied on. “What has to be done with all of these large language model-generated proofs is that they will have to be read by humans, and this creates an enormous extra burden on mathematicians,” says Anders Hansen at the University of Cambridge. On 8 September, OpenAI announced that it had found a solution to the Navier-Stokes problem , one of the most famous open problems in mathematics. It published the proof in two versions – one written in “natural language”, meaning a combination of English and mathematical symbols, as a human mathematician would write, and another written in the computer code Lean. The Lean proof is meant to be a formalisation of the natural-language version, allowing a computer to mechanically verify that all of its logical statements are true. The problem is, say Hansen and his team, that they don’t match. “This formalisation process is trying to replace peer review,” says team member Fabian Circelli , also at the University of Cambridge. “Peer review would mean that human eyes look at the proofs. But what we’ve shown in this paper is that using this type of AI auto-formalisation can’t serve the same purpose.” Read more The most interesting mathematical discoveries in OpenAI’s 722 new papers To be clear, the researchers aren’t saying that OpenAI has failed to solve the Navier-Stokes problem. It is entirely possible that both the natural-language proof and the Lean one provide a solution, just as there are hundreds of valid proofs of Pythagoras’s theorem . Instead, their point is a more subtle one: that OpenAI’s model has “mistranslated” when converting into Lean. “We are not saying that the natural-language proof is wrong,” says Hansen. “Nor do we say that it is correct.” The issue is that OpenAI presents the two proofs as identical, stating on Github that “This repository contains Lean 4 formalizations of the results presented in [the paper] ‘Finite time blowup for Navier–Stokes’”. This mistranslation occurs because the AI has to produce a Lean proof that “compiles”, meaning that the computer code is fully self-consistent and doesn’t produce an error, says Hansen. If, in the process of auto-formalisation, the AI finds a section of the proof that doesn’t compile, it will attempt to find a workaround even if it means diverging from the proof as written in natural language. Read more OpenAI has dumped 722 maths papers – now it must clean up the mess The team’s specific claim hinges on part of the proofs called Lemma 8.6. In the natural-language proof, an equation in this part requires that a certain value be below m + 4, where m is a whole number. In the Lean proof, the equivalent value is required to be below m + 5, which is mathematically weaker. To understand why, imagine being asked to solve the equation x + 3 = 6, to which the answer is x = 3. It is possible to write a proof that x must be less than 4, and also that x must be less than 5. Both of these are perfectly true mathematical statements, but they say different things. The latter proof allows more possible answers for x , making it mathematically weaker. OpenAI told New Scientist that it is aware of the mismatch between the natural-language proof and the Lean code, and that this doesn’t mean that either proof is invalid. It says it will rectify any errors in the natural-language proof as they are found, and will also continue the process of formalising the 722 maths papers the firm released this week , only some of which are accompanied by Lean proofs, which themselves haven’t been checked by hand. Needle in a proofstack Finding the divergence involved a slightly surreal process of asking ChatGPT to look for potential discrepancies between the natural-language and Lean proofs, then checking them by hand. Many of the discrepancies suggested by ChatGPT turned out, on inspection, to be consistent after all. “Going through all of these things manually was a nightmare,” says Hansen. In all, it took the team about two weeks to identify a true discrepancy, compared with the 88 hours OpenAI said its agents spent generating the proofs. “OpenAI boast about how quickly they were able to generate this result, but that’s only part of the process,” says team member Alexander Bastounis at King’s College London. Answering the question of whether AI models can accurately auto-formalise mathematics is essential if mathematicians are to trust these results. Anders and his team have demonstrated that it is possible for ChatGPT to produce a Lean version of a proof that doesn’t match the original natural language one, with the AI silently altering the logical argument in the process to cover up any errors. “It is trying to help me, but by doing that, it is not helping,” says Hansen. Subscriber-only newsletter Sign up to Lost in Space-Time Sign up to newsletter If this were to happen for a long and complicated proof, it would be very difficult for anyone to notice without inspecting both versions in detail. This issue is only more pressing because of the batch of 722 papers OpenAI just released. If we think ‘all of this is now true, the only thing we need to do now is read the paper’, then that is dangerous, says Hansen. “The purpose of doing science is that mankind should have an understanding of how the world works so we can make educated decisions. If we lose that understanding, what are doing?” Kevin Buzzard at Imperial College London notes that in discussions like these, it is important to distinguish between the statement of a theorem and its proof. For example, the statement of the famous Fermat’s last theorem is that for positive whole number a , b , c and n , aⁿ + bⁿ = cⁿ only if n is 1 or 2. This is easy to convert into Lean and can easily be checked. Once you are happy that a statement in Lean is correct, if the proof of that statement compiles, you can be confident the proof is true. The issue, as Hansen and his team have pointed out, is that this tells you nothing about the natural-language version of a proof published as a PDF document. “I am confident that the Navier-Stokes problem has been correctly resolved,” says Buzzard. “I am far less confident that the proof described in the PDF is correct.” Hansen says he hopes OpenAI will take the team’s work “very seriously” and that more work must be done on developing robust auto-formalisation techniques. “Do we have the solution? Not yet. Is it possible to do this in a controlled way? Yes, it will be, but the optimal and the ultimate way of doing this is completely unknown.” Journal reference: arXiv DOI: 10.48550/arXiv.2610.08144 Topics: Mathematics
Read the full original article:
newscientist.com#OpenAI
