DeepInfra raises $107M in Series B funding round

DeepInfraCompany
DeepInfra, the Palo Alto‑based AI inference cloud provider, closed a $107 million Series B round on July 9, 2026. The financing will fund its first data‑center outside the United States, a 1.7 MW Toronto facility hosting over 1,000 Nvidia Blackwell B300 GPUs, expanding its global capacity for enterprise AI workloads.
Deal Terms
DeepInfra announced a $107 million Series B funding round on July 9, 2026. The round follows a $18 million Series A in April 2025 and brings the company’s total raised capital to $125 million. While the lead investors were not disclosed, the capital infusion is earmarked for geographic expansion and infrastructure scaling.
Strategic Rationale
The company’s new Toronto data‑center marks its ninth location and its first outside the United States. With 1.7 MW of power and more than 1,000 Nvidia Blackwell B300 GPUs, the site adds a significant inference capacity tier. CEO Nikola Borisov emphasized that enterprises are moving from experimentation to production at unprecedented speed, requiring infrastructure that is both scalable and globally distributed. By locating capacity in Canada’s largest data‑centre market, DeepInfra can serve customers with data residency requirements and lower latency for North‑American workloads.
The Toronto launch also positions DeepInfra to tap into a growing ecosystem of Canadian data‑center operators, including Digital Realty, Equinix, and EdgeConneX. The facility is likely leased from a third‑party provider, allowing the company to scale quickly without heavy cap‑ex. DeepInfra plans to integrate Nvidia’s upcoming Vera Rubin GPUs, further future‑proofing its platform against the rapid evolution of AI model sizes.
From an investor perspective, the Series B underscores confidence in the AI inference niche, where specialized cloud providers compete with hyperscalers for high‑performance, low‑latency compute. The funding will enable DeepInfra to deepen its SaaS‑style offering—charging customers on a usage‑based model—while expanding its global footprint to meet the demand for production‑grade AI workloads.
Why It Matters
DeepInfra’s Toronto expansion gives it a foothold in a market where competitors such as CoreWeave, Lambda Labs, and Run:AI have already announced North‑American sites. By adding capacity outside the U.S., DeepInfra can attract customers with strict data‑sovereignty policies, potentially shifting a slice of enterprise spend away from rivals that are U.S.-centric. The added GPU inventory also raises the ceiling for the company’s expansion revenue, as existing customers can scale workloads without migrating to another provider.
For the broader AI inference ecosystem, DeepInfra’s move signals that demand for dedicated inference infrastructure is outpacing the supply from traditional cloud giants. Operators that can rapidly provision high‑density GPU clusters will likely capture a larger share of the growing $XX billion AI inference market, pressuring rivals to accelerate their own global rollouts.
Key Points
- DeepInfra closed a $107 million Series B round on July 9, 2026.
- The financing will fund a 1.7 MW data‑center in Toronto with over 1,000 Nvidia Blackwell B300 GPUs.
- Toronto is DeepInfra’s first data‑center outside the United States and its ninth overall.
- The company previously raised $18 million in a Series A round in April 2025.
- DeepInfra plans to add Nvidia Vera Rubin GPUs to the platform in the future.
Analysis
The $107 million Series B places DeepInfra on a trajectory that aligns with the rapid scaling of AI inference workloads across enterprises. While the valuation multiple was not disclosed, Series B rounds in the AI infrastructure space typically command 10‑15 times forward‑looking ARR, suggesting a valuation in the high‑hundreds of millions. The Toronto launch expands DeepInfra’s global footprint, addressing latency and data‑residency concerns that are increasingly important as AI models move from proof‑of‑concept to production.
For investors, the round highlights continued appetite for niche cloud providers that can deliver high‑performance GPU compute on a SaaS‑style, usage‑based pricing model. The addition of 1,000 Blackwell GPUs and future Vera Rubin support positions DeepInfra to capture higher‑margin, expansion revenue from existing customers seeking to scale models without the overhead of managing hardware. The move also underscores a broader trend: AI‑focused cloud operators are building a distributed network of data centers to compete with hyperscalers on both performance and geographic coverage. Operators that can efficiently lease capacity and quickly spin up GPU‑dense clusters will likely enjoy superior gross margins and stronger net revenue retention, making them attractive targets for later‑stage investors seeking exposure to the AI infrastructure tailwinds.
