Testing from OmniCalculator suggests Claude and ChatGPT are not the smartestThe report finds Grok 4.2 performs best in logic and problem-solvingClaude still leads in writing quality and tone

ChatGPT is still the most popular AI chatbot around, even with the exodus that’s underway to Claude, but is it the cleverest? A new report from OmniCalculator suggests that ChatGPT might not be the smartest AI around.

When it comes to the quantifiable math ability of these AI chatbots, the smartest free AI model is, rather surprisingly, Grok. xAI’s Grok 4.2 model specifically. That doesn’t mean anything about its writing style and ability, or anything else chatbots can do, but it does suggest that it might have the edge in math prowess.

Omnicalculator AI

(Image credit: Omnicalculator)

military deals, but also by how it composes answers and writes its responses.

Article continues below

You may like

The quality is hard to quantify compared to math skills, but easy to recognize. The OmniCalculator report highlighted Claude 4.6 as the best at it, able to process and respond to long documents without losing coherence and maintaining a consistent voice throughout. For the average person, this is much more important than which AI can make it through complicated logic and math problems.

It even comes out in the facsimiles of personality offered by the AI models. Claude is more willing to acknowledge uncertainty, which can make its answers feel measured rather than overconfident. That tone can create the impression of deeper thinking, regardless of any underlying reasoning.

Omnicalculator AI

(Image credit: Omnicalculator)

Legacy models, including earlier versions of ChatGPT and Claude, were found to revise or second-guess their own answers roughly 60% of the time in complex problem-solving scenarios. That kind of instability does not always show up in casual use, but it becomes obvious when you push these systems through multi-step reasoning tasks where consistency matters.

But Grok 4.2 cuts that instability rate down to 33.1%, meaning it is far less likely to backtrack or alter its conclusions mid-process. That’s great for reasoning and logic, but not much help in mimicking the smooth tones that make other models feel more polished.

Google logo on a black background next to text reading 'Click to follow TechRadar'

Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.

Purple circle with the words Best business laptops in white

The best business laptops for all budgets

Our top picks, based on real-world testing and comparisons