For the last few years, almost every AI infrastructure conversation in India has started on how much compute are we building, and how fast we are building it. From GPU allocations, hyperscale campus announcements to GW-scale roadmaps, the story has been about scale, and understandably so. Training large models is capital-intensive, power-hungry work, and India has been racing to build the core capacity to do it.
But look at what the analysts are reporting through 2026; a quieter shift is underway. According to Gartner’s forecast on AI-optimised infrastructure spending, this year, for the first time, global spending on inference is set to overtake spending on training – $23.3 billion against $19 billion respectively.
Inference already accounts for 55% of AI-optimised infrastructure-as-a-service spend in 2026, and Gartner expects that to climb to 59% by 2027. IDC tells a similar story at the macro level, estimating that worldwide AI infrastructure spending will reach roughly $497 billion in 2026. It has named the ‘faster-than-expected’ scaling of inference workloads as one of the reasons this number keeps getting revised upward.
Simply put, the industry has stopped talking mainly about how models get built, and started talking about how they get used. And this matters more than it sounds.
Training and inference aren’t the same problem
Model training is something you can centralise, but inference doesn’t work that way. Once a model is live, it has to respond in real time, every time. This means scoring a fraud transaction, checking for a defect on a factory line, powering a customer-facing assistant, etc. That workload doesn’t want to sit two hundred kilometres from the user or the machine it’s serving; it wants to be closer.
This is really an infrastructure design question wearing a technology costume. Where you place inference decides how fast, how reliable, and how compliant your AI actually is, especially in India.
India’s advantage is also its infrastructure gap
TRAI’s numbers for early 2026 put the country at over 1.09 billion internet subscribers. Nearly 441 million of them are in rural India, while about 652 million are in urban regions. That’s about as geographically distributed a digital population as exists anywhere in the world. AI adoption here was never going to stay confined to a handful of neighbourhoods in a handful of cities.
And yet that’s roughly where the infrastructure still resides. According to CBRE’s datacenter outlook for India, Mumbai alone accounts for more than half the country’s operational datacenter capacity, with Mumbai, Chennai, Delhi-NCR and Bengaluru together holding close to 90% of it. That’s not a criticism of how the industry has built so far; these cities had the subsea cable landings, the power availability and the enterprise density to serve the demands. But it does mean that India’s AI infrastructure map and India’s AI user map are, right now, two different maps.
The opportunity lies in closing this gap. The same CBRE report already shows early movement. Cities including Ahmedabad, Visakhapatnam, Patna and Bhopal are picking up datacenter interest specifically because of latency, 5G rollout and data localisation requirements.
The rise of the regional AI hub
What’s likely to emerge over the next few years isn’t a wholesale move away from core facilities. Training will stay concentrated, for good reason, but a new tier of regional AI hubs will emerge. This means places where compute, connectivity, cloud on-ramps, internet exchange points, CDN and edge infrastructure reside together, positioned closer to where demand is actually being generated, not just where it’s cheapest to build a shell.
Real-time AI is what will pull this layer into existence faster than anything else. This includes video analytics on a factory floor, fraud scoring inside a banking app, telecom network automation, gaming, industrial IoT; where a few hundred milliseconds of round-trip latency makes all the difference between a system that works and one that doesn’t. A bank scoring a transaction for fraud, a plant flagging a defective part before it moves to the next station, a telecom network rerouting traffic automatically; none of these can afford to wait for a round trip to a datacenter several states away.
The shift in how edge datacenters get built
It also means edge facilities can’t keep being designed the way they were for the CDN and connectivity era. AI-ready edge infrastructure needs meaningfully higher power density per rack, cooling systems built for GPU thermal loads rather than general-purpose IT loads, and connectivity that can carry AI traffic without becoming the new bottleneck. A facility that was adequate for content caching five years ago isn’t automatically adequate for inference today.
Connectivity deserves equal importance here, not just a supporting role. This means direct cloud on-ramps, regional peering, and dense fibers. India’s telecom backbone has done its part; the country’s 5G subscriber base has already crossed 400 million, and now the datacenter layer needs to be built to leverage that network.
Sovereignty is bigger than owning GPUs
There’s a sovereignty dimension here too, and it’s broader than the GPU capacity conversation. EY’s recent findings about Sovereign AI in India frames it well – sovereignty is about the ability to design, develop and process AI on domestic infrastructure and terms, not simply about how many accelerators sit in one campus. This matters because the ambition attached to this is large. NITI Aayog has estimated that AI could add close to $1.7 trillion to India’s economy by 2035. Getting anywhere close to that number would require India to process, store and serve AI workloads reliably across regions, not just in one or two power-rich pockets. This will ensure that the benefits, and the resilience, are genuinely distributed rather than concentrated in the same four cities that already hold most of today’s datacenter capacity.
A different question for enterprises
All of this changes the question enterprises should be asking. It’s no longer about which cloud should host AI. Instead, it’s increasingly about where each part of AI actually runs. Training might occur in one location, inference in another location closer to the customer, data processing and storage somewhere shaped by compliance requirements. These aren’t always the same address anymore, and treating them as one decision is where a lot of AI deployments underperform.
Building the fabric, together
No single player builds this alone. Hyperscalers, neoclouds, telcos, datacenter operators and network providers each hold a piece of this puzzle by bringing compute and connectivity closer to where India’s businesses and users actually are.
India’s AI story will not be decided only by how much compute we build at the core. It will be decided by how intelligently we distribute that compute closer to where AI is actually consumed.
Vipul Kumar, Senior Vice President – Edge & Network Business, CtrlS Datacenters
Vipul is a seasoned telecom and datacenter leader with over two decades of rich experience spanning submarine cable systems, edge datacenters, and network infrastructure ecosystems. He is passionate about building sustainable, compliant, and scalable digital infrastructure that empowers regional enterprises, SMEs, and hyperscale players alike. As Senior Vice President – Edge & Network Business at CtrlS, he leads initiatives that bridge connectivity and compute — from fiber and network deployments to strategic partnerships and business development, building the network foundation and ecosystem partnerships that power CtrlS’s pan-India edge datacenter expansion.