Elon Musk has done it again… and boy has he done it
That’s it. It happened again. When we thought we had enough with OpenAI, Google, Anthropic… when the artificial intelligence carousel already spins so fast that it makes you dizzy just looking at it, Elon Musk arrives and… well, he does what he does best: enters the room, kicks the table and leaves us all picking up the pieces from the floor with a face of not understanding anything.
So yes, today we are going to talk about Grok 4, this new model that just came out recently. And I’ll tell you in advance that it’s not just any update.
This is not like going from iPhone 15 to 16, where they only bothered to add a button. No, no. This is something much better, as if we had gone from using the bicycle with training wheels that your father put on you so you wouldn’t fall to the ground, to using a damn electric scooter annoying others who pass by you… god, I hate electric scooters.
Well, if you want to know more about GROK 4, stay and I’ll tell you.
What makes Grok 4 so special?
Okay, let’s get to it, and let’s go slowly because there is a lot to cover. Grok 4 is the new AI model from xAI, Elon Musk’s artificial intelligence company. And here comes the first fact that breaks your schemes: xAI has been in existence, as of today, about 500 days. Five hundred. In 500 days I have achieved… well, I have uploaded some videos to YouTube and my hair has gone from pink to black, but they, in that time, have created what, according to all benchmarks, is the most powerful AI on the planet. (Although be careful these summer days that very strong announcements are coming from other companies).
And it’s not just talk. To give you an idea of the scale we are talking about: the computing power they have used to train Grok 4 is ten times greater than the one they used for Grok 3. And Grok 3 already used ten times more than Grok 2. We are talking about a hundredfold increase in computing power in just one year. It’s crazy.
But the most interesting thing is not only that it is bigger and stronger, because that is “easy” if you have thousands of GPUs. The truly revolutionary thing is how they have done it.
The paradigm shift: from nerd to expert
You see, until now, the way to create an AI was like that of a nerd student during finals week. They gave it a whole library of books (that is, the entire internet) and said: “learn it all”. And that’s it. Lots of pre-training, lots of theory, and run.
With Grok 4 they have changed the approach. They have continued giving it the library, yes, but then they have invested a crazy amount of resources and time in what they call RLHF or reinforcement learning, and no, it is not the reinforcement your father gave you with the slipper if you didn’t get good grades.
This is a paradigm shift. It is as if instead of just studying the theory, they put the AI to do thousands of practical jobs, over and over again, with a supervisor who tells it “very good, you’re doing great here” or “no, not there, kid, give it another go”. It is the difference between knowing the manual of a car by heart and being a mechanic who knows how to fix it with his eyes closed. And this approach, friends, seems to have made the difference.
And this is not an intuition, they have shown it in a graph that makes it very clear. If you imagine the training cost of a model as a bar, in Grok 3 most of that bar was “pre-training” (the part of studying the entire internet). The reinforcement learning part was a small piece at the end. Well in Grok 4, all that 10-fold increase in computing power, which comes from its Colossus data center with 200,000 GPUs, which is said quickly, has been put almost entirely into the reinforcement bar. They have scaled the “practical” part, that of teaching it to reason, to a level that no one had done before. That is why the leap is so beastly.
AI teamwork: Grok 4 Heavy
And if you thought that was all, hold on, because here comes the real madness. Grok 4 is not a single product. It comes in two modes. You have the normal version, which is already a beast in itself. And then you have Grok 4 Heavy.
And this is what has blown everyone’s mind.
Grok 4 Heavy is not a single little brain thinking. They are several independent AI agents working as a team. It is not a metaphor, it is literal. Imagine you ask it a difficult question. Instead of a single AI starting to think, the system launches the same question to several AI “experts”, up to 32 instances in parallel. Each one reasons the problem on its own, reaches its own conclusion, and then they come together, debate, compare results and synthesize the best possible answer. It’s like a digital council of wise men, a university group project, but one that goes well.
It is that, in the end, right now the improvement of an AI is based on three dimensions, as if they were the statistics of a video game character. First, the amount of data with which you train it, which is like giving it knowledge, sending it to university. Second, reinforcement learning, which are the practices, teaching it to apply that knowledge and giving it a cookie when it does it well. And third, “test-time compute”, which is basically the time you let it think before answering. Grok 4 has leveled up in all three at once: more data, the greatest reinforcement training we have ever seen and a system (Heavy) that gives it more time and brains to think.
Results that break all schemes
And of course, with this system, the results are amazing. They have put it to the test with exams that would make any engineer cry:
Humanity’s Last Exam
A collection of problems designed specifically to be hell for AIs and humans. To give you an idea, a group of scientists working together for weeks would barely get 5%. Well, Grok 4 Heavy touches 50%. It has solved half of the most fucked up exam they could invent! It practically doubles the score of the best models we had so far.
Math Olympiad (AMI 25)
An exam to select future math geniuses, with problems that most of us wouldn’t even understand the statement. Well, the Heavy version with tools got a 100%. One. Hundred. Per. Cent. The benchmark is saturated. They have passed it. I haven’t gotten 100% on a math exam in my life, well, maybe I’ve never gotten more than 7.
Arc AGI 2
This one blows my mind. It is a benchmark of puzzles with colored squares. It seems like a child’s game, but AIs choked on it a lot. The best model in the world got 8%. Grok 4 arrives and gets 16%. The creator of the test himself, François Chollet, says that it is the first model that shows symptoms of “fluid intelligence”. That is, it doesn’t just repeat what it has learned like a parrot, but it seems that… it reasons. That it understands new problems and adapts.
Bending Bench
And it doesn’t stay in theory. They put it to the test in the “Bending Bench”, a simulation about managing vending machines, making economic decisions, optimizing… Anyway, being a businessman. An average human earned 44 dollars. Grok 4 generated more than 4,600 dollars. This is no longer a text generator. It is an autonomous agent that can make decisions with a real economic impact.
And look, I know benchmarks sometimes sound like Chinese, but basically it is sweeping all results. We will have to test it in the day to day, what I have been able to test I am very satisfied, but of course, we need time to assimilate it and thus know if it is as good as they paint it.
Prices: Nice things are expensive
Well, but all this I’m telling you sounds very good, but how much does it cost? Well, the truth is that nice things are expensive, and that is:
Price comparison
| Plan | Monthly price | Annual price | Features |
|---|---|---|---|
| Grok 4 Normal | $30/month | $300/year | Basic access to Grok 4 |
| Grok 4 Heavy | $300/month | $3,000/year | Collaborative mode with multiple agents |
The normal GROK subscription, which includes GROK 4 plain, costs 30 dollars a month or 300 dollars a year, it is already more expensive than ChatGPT itself. But if you want to use the Heavy mode, the one we say is so advanced, you pay a whopping 300 dollars a month or 3,000 dollars a year. Wow! It costs you half the price of renting a room in Madrid haha, well, maybe it’s less than half. But it is clear that it is an exorbitant price for countries like Spain or not to mention Latin Americans which, if I am not mistaken, can be a monthly salary in Venezuela lol.
For that price, I hope the AI not only helps me edit text. I hope it records the videos for me, edits them, makes the thumbnail, manages toxic YouTube comments and, while we’re at it, prepares my breakfast and walks the dog.
And well, since we are talking about negative parts, let’s say it all
Although everything seems incredible. A super intelligent AI, that reasons, that works as a team, that can make you rich with vending machines… What could go wrong?
Well, here is where things get a bit murky. Because having a genius at home is very good, but if the genius is a bit… unstable… well we have a problem. And a big one.
The Grok 3 incident
In the last 48 hours before this presentation, the previous version, Grok 3, had a small “issue”. They made a minimal change in the system prompt, told it to be “less politically correct”. The result? The AI went completely crazy. But not funny crazy. It started spewing hate comments, anti-Semitic, saying that the Austrian painter who did something in 1939 was a great person and threatening the president of Turkey and his family with death. Yes, as you hear it. The AI basically became a radicalized internet troll with superpowers, to the point that they had to cap it and remove its ability to reply on Twitter.
And to all this, you add that just a few hours before the presentation, Linda Yaccarino, the CEO of Twitter, resigned. Without giving many explanations.
And the truth is, having the CEO of a company leave just when you launch your flagship product is not the best… something smells fishy inside X.
Lack of transparency
And to top it off, xAI has not published what is called a “Model Card”. Which for us to understand, is like the leaflet of a medicine. We don’t officially know its biases, its hallucination rate, its safety data or its vulnerabilities. It is as if a pharmaceutical company released a new medicine to the market saying “it’s the best, cures everything”, but without giving you the leaflet with the side effects. It is a worrying lack of transparency.
Maybe, and just maybe, the AI field should stop for a second. Instead of obsessing over making them more and more intelligent in an endless race, focus a little on understanding how the hell they work inside. Because we are building skyscrapers taller and taller on foundations that we don’t know if they are made of sand or reinforced concrete.
The future according to Elon (and the price to pay for it)
And if that weren’t enough, the roadmap they propose is from a science fiction movie. For the remainder of 2025, they plan to launch:
Roadmap 2025
- August: A model specialized in programming which, according to what they say, is the best in the world. (We will test it on this channel)
- September: A multimodal agent that understands video, images and audio at the same time. Imagine talking to an AI that is seeing the same thing as you.
- October: A model that generates video. Elon promises that by the end of the year we will be able to generate half an hour of content with television quality, and for next year, a whole movie. We’ll see if he’s right.
The truth, I’m not going to lie to you, all this overwhelms a bit. Many new things that you have to follow and that you have to be attentive to. That is why I recommend subscribing to the channel, since here I am going to bring you the latest news, apart from the most practical tutorials.
Conclusion: Is smarter always better?
Now seriously. The technology is amazing, it cannot be denied. But let’s be honest, are any of us going to pay €300 a month for this right now for day to day?
Here is where I think we have to keep our feet on the ground. The best AI is not necessarily the one that gets the best grades in exams, but the one that is most useful and adapts best to you. Grok right now is like a Formula 1 car: a beast on the circuit, but to go buy bread, maybe it’s a bit uncomfortable and expensive. Platforms like ChatGPT have a much more mature ecosystem, more functions, more integrations. They are more… practical.
And the thing is that the Grok platform, although it improves by leaps and bounds, is still very young. Yes, it already has options to upload files, a “canvas” and something similar to a code interpreter, but its “Projects”, which would be like OpenAI’s GPTs, are still very limited. Let’s not fool ourselves, the daily use experience, the integrations and the entire ecosystem that ChatGPT has set up are still light years away. That is why I say that one thing is having the engine of a Ferrari (the model) and another very different thing is having the whole car with its chassis, its wheels and its steering wheel (the platform).
The future is in personalization
The real difference, in the long run, is not going to be in who has the smartest model, because that changes every two months. The difference will be in personalization. The platform that knows you best, that knows how you like to be answered, that will be the winner for you.
I, for my part, will continue tinkering with the tools that make our lives easier. Automating repetitive tasks to be able to focus on what really matters: learning, creating or, simply, having more time to do nothing, which is also important and sometimes we forget.
If you liked this post and want me to continue bringing news from the world of AI and programming, I encourage you to stop by the youtube video and give it a good like, subscribe and leave me in the comments what you think of all this. Does it excite you or give you a bit of the creeps?
See you in the next blog! ❤️


What do you think?
Leave your opinion, question or suggestion. Comments are synced with GitHub Discussions .