Nebius, the AI cloud company, announced today that Nebius acquires Inferize, an inference optimization company. Inferize’s technology shortens the time teams need to launch and scale large AI models. As a result, inference workloads become elastic.
Inferize’s technology and team have now joined Nebius Token Factory. This is the managed inference platform that Nebius built for production AI. Therefore, customers running AI at scale stand to gain from the added expertise.
Why Cold Starts Drain GPU Budgets
Cold starts describe the time models need to load before they serve a single request. Teams that run AI models in production at scale face extra time and costs because of them. Specifically, assigned GPUs sit idle at launch time. They also sit idle when demand spikes and new instances spin up. The same problem appears when weights update mid-run, for example during reinforcement learning. Consequently, platforms must hold spare capacity just to hit their service-level targets.
Nebius acquires Inferize technology cuts this “idle GPU tax.” Thus, capacity can scale much more closely with actual usage. In turn, companies achieve higher capacity utilization and better token economics. For any team that pays for GPU time, this shift matters.
Danila Shtan, Chief Technology Officer of Nebius, said: “Running inference well takes more than fast GPUs and optimized models. The whole system needs to respond when demand changes, including how quickly additional capacity is ready to serve customers. Inferize brings technology that accelerates that process and a team with deep expertise in GPU systems. We’re bringing both into Token Factory to make it more responsive to customer demand and get more useful work out of our infrastructure. The team’s contribution will extend well beyond this first integration.”
How Inferize Fits Into the Nebius Inference Stack
Nebius acquires Inferize and adds another layer to the way Nebius runs production inference. Earlier, Eigen AI brought optimization at the model, kernel and system levels into Nebius Token Factory. Likewise, Clarifai’s core team and licensed technology added system-level inference and compute orchestration. Together, these pieces form a broader inference stack.
Guy Bortnikov, co-founder and CEO of Inferize, said: “Keeping spare GPUs running is the price of being ready for demand. Removing that cost is what we built Inferize to do, and Nebius is where it can go straight into the platform. Our team will work across the stack with one objective: serving more customer demand from every GPU.”
The founders of Nebius acquired Inferize and launched Inferize in January 2026. Notably, the team had a working prototype within three months. Inferize’s engineers will now work across Nebius Token Factory. First, they will integrate their technology into the platform. Meanwhile, Nebius did not disclose the terms of the transaction.
Explore IT Tech News for the latest advancements in Information Technology & insightful updates from industry experts!
News Source: Businesswire.com