Party for AI, now what?
I have a strong feeling to write something down this morning.
I thought last week was crazy coming back from a trip in Italy with the new M5 Ultra release and exciting MiniMax H3 for video generation. Obviously, I underestimated the amount of excitement and new things that were coming out this week. First, a new world model. Now, new frontier models from the two frontier labs with amazing multimodal capability. Which makes me think, now what?
I’m definitely going to try out some of these models as they become available and also my local RTX 5090 rig is ready for a video gen splash with my daughter over this long weekend.
At work, I hand-wrote benchmark questions, answers and grading criteria for an AI Agent. What intrigues me is not how exciting that work is and what a great Agent we built based on that. The fact that it took me 2-3 hours per question to answer ~15 questions, and another couple hours to review the context turned into an Agent response these questions and generalize them to similar under 1 minute each turn excited me. I know this is expected, and I shouldn’t be surprised as I’m a part of this field. It’s the norm these days.
What’s our game now?
Residual
The residual concept was introduced by ResNet in 2015 and Transformers are still using this technique in modern LLMs. The idea is that a layer doesn’t learn the whole mapping. It keeps everything it’s given and learns only what to add.
Back to life, the town I went to in Italy is a small town with more than 2,240 years of history. Even one of the ancient Roman roads was discovered there. It was the rich culture and how slow life is there that really surprised me. Things change slowly when you step on a stone that might have been sitting there for a couple hundred, even a thousand, years. In Agentic AI, new things pop up every now and then. We can’t ignore the fact the only constant in today’s world is change. We had a busy agenda during the trip where I woke up at 7am or earlier and did not call it a day until 10pm. However, the surroundings, and how people live their life left something inside me after coming back from the trip.
I think it’s the residual, the ‘unimportant’ stuff, that’s small enough to dismiss and the only thing slightly changed. Sometimes, it’s a good idea to slow it down, listen to the ‘noise’. And maybe a new perspective, new ideas pop up. With LLMs generating tokens way faster than any human would be able to do, our job is not to get the muscle memory to act, but to create white space to Think, to Believe and to Dream.
No playbook is the playbook
Yesterday, my wife was out of town and I had a great time with my daughter after work. She is learning piano and her mom is the parent handling that piece of family affair. My wife is usually very strict about the practice because of the passion for music. Hitting the right tone and expression is definitely important. If you’ve ever played piano, you know it’s all about that kind of small touches passed on by music teachers sometimes helped by the parent to enforce that learning. Following the Playbook is definitely the game in the early stage of learning. I don’t know music that well, but I do enjoy every bit and every punch musicians make.
As you can imagine, my time with my daughter went way off in the traditional music learning experience. While I was following up with all the great AI news over my phone. Something caught my ears, that’s not any tone that I heard from her earlier music practice, it’s something that’s completely unheard of but something familiar. She was altering the original music and making her own. And what I heard is really a girl dancing in the world of music. I was totally shocked! Not by how good she’s playing, but by how much creativity she brought to me.
As we are designing the agentic system, the lowest-hanging fruit is always the thing that has some sort of Playbook. And we are providing context and harness in a way such that the Agent could understand and follow the process either to solve a business problem or to dynamically make decisions or recommendations. Generalization of Agents is what we are looking for in terms of agentic capability. What about beyond that? Is there anything an Agent can create that has never been created before?
When we look at history, most of the great human inventions did not follow the Playbook. I truly believe the Playbook for today is to have no Playbook. Our greatest science and art creation are not something that existed before the same way as the music comes from nowhere and from the fingertips of my daughter. On the side of frontier models, we are using trillions of parameters or some math forms nowadays to capture the variability in the real world. As the truth of the universe is infinite, the model parameters are not. And we are just having the model to look at tiny specks of dust on the ruler of the entire universe. So, what’s next? Can everything be described in parameters of trillions in the deep neural network or other mathematical form? The approximation is close, not the full picture.
And we’re not going to play by the book. Human, go and create.
Who are we playing against
Back in the day when I invested time into quantitative Day Trading, it’s really important to know who you are playing against. I was reading tapes, option Gammas, potential market maker flows to understand how the mechanism works. A major news event can suddenly draw your probability distribution from one direction to another. It was a fun game. When it comes to executing the trade, you really need to know where your Edge is. When you don’t know your edge, you lose.
Same way with Agentic AI. The human brain can generate tokens at a rate of 1-20 tokens per second and is highly bounded by our typing and speaking speed. Thanks to mother nature, we’re 1,000x-10,000x more efficient than modern AI chips in raw token generation and 1 million - 10 million times more efficient than frontier AI models on operational intelligence. And we have about 1,000 to 2,500 Terabytes of comparable VRAM, that’s about 12,500 to 31,250 H100 GPUs. And every one of us is worth $0.375 to $1 billion when that power is fully utilized. And that’s a lot. And even more, the real value of us is the memory, experience and interaction we can bring that no machine could do today.
The next time your executives ask about AI, show them that number. And we have a good edge to win.
It’s under control
I have a passion to develop an Enterprise grade Agentic system that controls a swarm of AI Agents to dynamically take actions and help with business decisions. There are a lot of fruitful thoughts there. A2A, Protocols, vector search. And some are even using coding agent Harness as a master dispatcher. Those are great, but what’s the ultimate Harness? The industry has this idea that when models reach AGI, we will not need Harness anymore. Is it?
The experience I have is that even with the best model and harness combined, the Agent is still an Agent. No direction, no strategy, just matrix multiplications based on billions or trillions of parameters. I can clearly tell what work is done by an Agent and what’s done by a person.
Before it gets too wild, every one of us is still the Harness. The good ones. The ones that are accountable. The ones who don’t generalize our thinking and mind based on a set of reward scores.
Fxxk AI, but since you’re useful, we’re going to be your Harness.