Text me
All posts

Why I Built inference.vip

The goal is not just to host another model. It is to figure out how to make powerful AI cheaper, easier to access, and actually nicer to use.

I built inference.vip because I started noticing something that felt pretty backwards with AI. Everyone was using the biggest, newest model for basically everything. Models like Opus, Fable, Sol, and similar frontier models are incredible, but they can also be 10 to 30 times more expensive than what most people actually need.

So I thought, why not just run the models myself?

Kimi K3 is already more than capable for a huge amount of coding and everyday AI work. By running it on my own infrastructure, I could put a lot of users on the same servers and bring the cost down while still giving people access to a genuinely powerful model.

From there, I wanted to solve the annoying part too. Getting started with AI infrastructure can still feel way harder than it should. With inference.vip, the idea is simple: generate a command, copy it, paste it into your terminal, and you are working.

Then I started building around that idea with things like a built in prompt library, ridiculously fast CLI setup, access from your phone, session resumption from anywhere, and new approaches to make coding with AI feel more seamless.

The goal is not just to host another model. It is to figure out how to make powerful AI cheaper, easier to access, and actually nicer to use.