The future is local
Small distilled and quantized AI models are becoming feasible.
We live in a strange world; I feel it has become like a mantra enabling us to feel safe in our ignorance while we watch the world pass us by. Or better yet, feel we can’t understand. And I feel, so many feels in this paragraph, for those people. I’ve spent my entire life trying to be in sync with the world's knowledge, as much as a single person can. I don’t need to understand every nuance and every detail of technology, process, biology, sociology, but to know general strokes. That enables me to focus on finding out more about a subject I'm interested in at any time. … Master of none, but still better than master of one.
And, I’ve finally fully embraced AI. I haven’t seen a line of code in my off time in weeks, months at this point. ( I code for fun too )
I remember talking to a college colleague of mine in 2006 at a nondiscript bar attached to a supermarket about artificial intelligence and how it could emerge spontaneously from the chaos of information we produce. At that time, I thought about my research about how much information the internet can contain at a certain time, which I completed a couple of years ago, way before the AI revolution of 2024.
I was close; it did not self arrange out of chaos at a global scale, so far we know, but did self arange, with guidance, in neural nets.
As I’m writing this, I’m cracking up. Decades of skill summed up in a couple of billion neuron connections. LoL. I know it’s not funny, but it kinda is. The sheer power is intoxicating. Ability to finish projects, experiment with different approaches. The ability to go from idea to working concept is just a prompt away. It truly is a drug.
I could not have predicted it; it was decades away, if ever. But now, here we are. On my modest setup at home, I'm running a 35 billion parameter mixture of experts model, Qwen 35B MoE via LM Studio, that can do tasks that would take me weeks, if not months, to complete, and it does it in minutes. It still breaks my mind thinking about it. Frontier models have over 700 billion parameters and some over 1.5 TRILLION, but, as it turns out, for tasks like coding, they don’t need that much. They need as few as a dozen billion.
I don’t think the future is some global supercomputer; I think that what I wrote in my post about Apple is still correct.
The future is local models. It may seem like it's unattainable, but when, not if, the bubble bursts and memory and compute become affordable, everyone will run their local AI models. Let’s just look at memory. High speed memory this, high speed memory that, but I’m fine if a task takes a couple of minutes longer; that’s just fine. I’m on an AM4 platform with DDR4 memory at 2400mhz, nowhere close to 6000mhz or more that a newer system can achieve, and that’s fine. System memory speed is important, as a large portion of the model I'm running is offloaded to system memory; if not, it would all be on fast GPU vram. And at the end of the day, data is mine, output is mine, control is mine; no one is mining me for data.
I remember a time when 60mb hard drives was a luxury. I remember 5.25 inh drives, not the 3.5’’ disk drives that everyone nowadays thinks of as a save icon. The actual 5.25inch disk that was flimsy as hell, and where the name floppy disk comes from, as you could use it as a freaking fan. And, what’s more to the point, it was not that long ago; I'm not some senile greybeard techno freak; it’s just that we have attention and retention deficits and think that everything that happened 30 years ago is ancient history.
Today, I have a 2 terabyte drive, one of many, in my system, with “modest” 32 gb or system memory. I remember trying to optimize my apps and fighting the bloat for every kilobyte so that the entire Windows app remains bellow 1mb, and feeling ecstatic about it. Now, a freaking Notepad with modern bells and whistles is above 100mb. Such absurdity.
The point is, a 700 billion parameters AI model that requires 1Tb of working high speed memory are not that unattainable. Yes, they are expensive now. Well, if you are well off and have 5k or so to spend, you can get NVIDIA Spark or an AMD equivalent 365 AI with 128 GB of unified memory. Get a couple of those or two Mac Studio Pros with 512 GB each, and you've got yourself a full-blown local equivalent to ChatGPT.
So, yes, it is expensive but not like it's unachievable or that the path towards domestication of that technology is not visible. It is already present.
I’m aware my hardware is not up to snuff to run more productive models, so until that time I’ll just pretend to be happy to spend over 14.000 monopoly bucks on ChatGpt, but that can not last. Such subsidy can not last.
We should all be aware of the ancient Chinese proverb or curse, depending on who you ask: may you live in interesting times. I thought that those times were behind me, but as it turns out, we are in the middle of them right now.
I took a couple of weeks off writing to test and finish a couple of projects I wanted to do, so stay tuned.





