AI Cloud
The demand for “AI Cloud” has skyrocketed in Spain in parallel with the adoption of generative AI among large corporations and public organisations. To train large language models or perform inference at scale, organizations require a platform that combines computing power, data governance, and predictable costs. This article aims to briefly explain what an AI Cloud is, compare public cloud with a private on‑premise AI Cloud, quantify the three‑year total cost of ownership (TCO), and detail the technical components—GPU servers, liquid-cooled racks, and high-speed networking—that meet the demands of cutting-edge enterprises and regulated sectors such as banking, defense, and telecommunications.
Growth of “AI Cloud” in Spain
In the past two years, generative AI projects based on language models (LLMs) have moved from proofs of concept to production deployments in banking, telecoms, healthcare, and public administration. According to the latest reports from the National Observatory of Technology and Society (ONTSI), business use of AI in Spain exceeded 11% in 2024, and the most conservative estimates forecast reaching 18% in 2026 and 25% in 2028.
Meanwhile, the national data center market, currently growing at a 12% annual rate, could surpass €3 000 M in 2027 and exceed €3 600 M in 2029, driven by the “training fever” pushing companies to consume thousands of GPU hours to iterate their models.
In this context, the concept of “AI Cloud” has gained even more traction: cloud environments—public or private—specifically optimized for high-performance training and inference workloads, whose demand is expected to double before 2028.
Speak with a technical advisor and we’ll guide you
What is an AI Cloud?
An AI Cloud is an infrastructure that integrates four technical pillars:
- State-of-the-art GPU servers (NVIDIA H100/H200 and the new Blackwell family from 2025 — B100, B200, GB200 — as well as AMD MI300 or others) with NVLink, NVSwitch, or NVLink‑C2C links to eliminate intra‑GPU bottlenecks.
- 200/400 Gb backbone network—with pilots already at 800 Gb—Ethernet or InfiniBand with ultra‑low latency and zero oversubscription.
- Parallel flash storage capable of delivering hundreds of GB/s to feed the data iterations required for training.
- Advanced cooling, preferably direct liquid cooling (DLC), to maintain energy efficiency with servers consuming 8‑12 kW.
In practice, there are two paths to implement it:
| Option | Location | Billing model | Data governance |
|---|---|---|---|
| Public cloud (hyperscalers) | Provider's data centers (multi‑tenant) | Pay‑per‑use/hour | Data resides outside corporate perimeter |
| Private on‑premise AI Cloud | Own facilities or colocation in Spain | Initial CAPEX + reduced OPEX | Data remains under local jurisdiction |
Public cloud vs Private AI Cloud comparison
Three‑year costs and TCO (≤ €1 M CAPEX scenario) — for the same task
Let's consider the same cluster of 16 NVIDIA H100 GPUs (4 nodes × 4 H100) purchased with a maximum €1 M initial investment. Internal comparative tests show that, thanks to a dedicated 200/400 Gb network and latency under 3 µs, training time is reduced by an average of 15% compared to the same load running on AWS, where inter‑GPU latency between p5 instances is around 30–40 µs.
| Item | Public cloud (AWS) | Private on‑premise |
|---|---|---|
| Instance type | 4 × p5.24xlarge (4 × H100 each) | 4 × Asus ESC N4‑E11 (4 × H100 each) |
| Hourly price per instance | €14.6 | — |
| Equivalent hours/year | 10,074 h (8,760 h × 1.15)* | 8,760 h |
| Total cost over 3 years | €1.80 M | €1.20 M (€1 M CAPEX + €0.20 M OPEX)** |
* AWS requires 15% more hours to complete the same work due to higher latency and network oversubscription.
Result for the same task: the on‑premise cluster saves around €600,000 (~33%) in just three years and completes training faster, translating into shorter innovation cycles.
Security and regulatory compliance
- Public cloud: data is transferred and processed in third-party infrastructure. Compliance with National Security Scheme (ENS) or Organic Law 7/2023 on sensitive data protection may require local zones, end-to-end encryption, and confidential computing controls, which increase costs.
- Private Cloud: data never leaves the corporate perimeter; internal L2/L3 controls are applied and ENS audits are direct, without third-party dependencies.
- Military projects: Law 11/2023 on National Defence, Classified Information Security Manuals (ICN) and NATO STANAG 4774/4778 agreements require all classified assets to remain in “classified clouds” within Spanish territory, operated by accredited personnel with isolated networks. Hyperscalers do not yet offer “Secret”-level regions in Spain, so military AI developments must reside in private infrastructure or Ministry of Defence data centers to ensure compliance.
Performance and latency
Although hyperscalers offer 3,200 Gb aggregated networks, these networks are shared with thousands of tenants. In a private AI Cloud a flat dedicated 200/400 Gb network is deployed, with deterministic training queues and latency below 3 µs, reducing total training time and improving data science team productivity.
Fundamental advantages of private AI Cloud
End-to-end security
Banks, telcos and defense organisations are subject to regulations requiring network segregation, access traceability and end-to-end encryption. Keeping data within national territory simplifies implementing Zero Trust policies and avoids cross-border transfers that require adequacy agreements.
Sector-specific compliance
- Banking: The Bank of Spain demands demonstrable control over third-party providers and minimisation of data geodispersion.
- Defense: Tactical AI projects must run on classified clouds and isolated networks in accordance with Multi-Domain Combat doctrine.
- Telcos: The CNMC heavily penalises call record (CDR) leaks and customer metadata breaches; on‑premise control mitigates the risk.
TCO savings and energy efficiency
Direct liquid cooling can reduce fan consumption, allow higher water temperatures and lower the Power Usage Effectiveness (PUE) from 1.4 to 1.1. This leads to up to 40% lower energy costs and about 20% reduction in TCO.
Essential prior study: the suitability of liquid cooling depends on the technical room, cold water supply capacity and rack server thermal density. For this reason, we evaluate each project with a thermal and flow report: if the environment doesn’t allow it—or node consumption is moderate—optimized air cooling may suffice and be more cost‑effective.
In high‑density scenarios (≥ 8 kW per node), DLC typically pays off in under three years; in lower‑power deployments it’s necessary to compare both strategies to decide which provides the best return.
Dedicated performance
Having a private AI Cloud means the entire cluster works solely for your organization. You don’t share network, GPU or storage with others, so every euro invested converts into actual compute for your models.
- Consistent performance: no “noisy neighbors”, performance and latency remain stable 24×7.
- More control: you choose network speed, topology, and queue policies tailored to your workloads.
- Optimized infrastructure: we tune power profiles for the required training or inference, reducing wait times and improving energy efficiency.
- Higher productivity: internal tests show up to 10× fewer batch wait times compared to multi‑tenant environments, leading to faster iteration and lower cost per experiment.
In summary: a private Cloud eliminates queues and surprises, dedicating all power to your AI team.
Explore our advanced configurations
Key components of a private AI Cloud
GPU Servers: vendor‑agnostic approach and catalogue examples
At Ibertrónica we remain fully vendor-agnostic: we partner with Asus, Gigabyte, ASRock Rack, Supermicro and other global vendors. Each project begins with a technical study—model profiles, training schedules, power and space constraints—and our role is to offer the customer the optimal combination of GPU nodes, networking, storage and cooling system for their specific task.
Below are just representative examples from the available range; we have dozens more configurations that we can tailor based on budget, density, or model roadmap.
| Vendor | Model | Summary |
|---|---|---|
| Asus | ESC N8‑E11 | 7U chassis, HGX H100/H200 backplane, 8 GPUs, NVSwitch, Blackwell‑ready |
| Gigabyte | G593‑SD0 | 5U chassis, 8 GPUs, DDR5‑8800, OCP 3.0 |
| ASRock Rack | 4U8G‑ICX2/2T | 4U chassis, 8 GPUs, dual Ice Lake, redundant 3+1 PSU |
| Supermicro | GPU/AI Liquid‑Cooled Series | Up to 12 kW per node, certified DLC, over 100,000 GPUs delivered in 2024 |
Liquid‑cooled racks
Ibertrónica’s DLC racks integrate DLC manifolds, redundant coolant distribution units (CDUs), and quick‑connect fittings. Each rack supports up to 96 GPUs and dissipates up to 100 kW of heat while monitoring temperature, flow and pressure through a centralized BMC.
200/400 Gb backbone network
The backbone network is essentially the data highway connecting all parts of the AI Cloud. Its mission is to ensure GPUs, storage and data science teams can exchange data without delays.
- Speed: We currently work with 200 and 400 Gb links, with pilots at 800 Gb in coming years.
- Common vendors: Broadcom, NVIDIA (formerly Mellanox), and Intel dominate this market, guaranteeing interoperability and long-term support.
- Why it matters: if the network is slow, GPUs spend more time waiting for data than computing. A fast network enables faster model training, fewer machine hours and earlier project go‑live.
In summary: investing in a high-speed backbone is key to maximizing each euro spent on GPU servers and reducing the overall cluster cost.
Conclusion
For serious development projects involving intensive workloads over several years, a private AI Cloud proves significantly more cost-effective than third-party public cloud (AWS, Microsoft, Google, etc.), even with a relatively small initial investment. CAPEX amortizes quickly, while public cloud bills scale linearly with GPU hours consumed.
The larger the investment, the faster the payback: the effective cost per hour can be two to three times lower than AWS, and the private cluster’s lower latency speeds up training cycles, accelerating model deployment. Ultimately, the more ambitious and long-term the roadmap, the more pronounced the cost and performance advantage of a private Cloud.
If your organization plans a mid-term AI roadmap, evaluating an on‑premise cluster from the start is the strategy that maximizes return and data sovereignty.
Ready to take your AI to the next level? Fill in our contact form and one of our architects will reach out to analyze your case and design the best solution for your project, with no obligation.
Request a customised proposal
![]() |
![]() |
![]() |
![]() |
La tecnología detrás del Cloud para IA |
Cómo desplegar tu Cloud para IA con Racks VibeRack de Ibertrónica |
Nuevos armarios OCP V3 |
Nvidia HGX: Plataforma abierta que impulsa la IA y HPC a gran escala |



