On-premise AI scales flexibly and step by step: you can scale horizontally (add more servers) or scale vertically (more powerful GPUs). The onprem.ai software manages all systems centrally and distributes the load automatically.
Scaling strategies
Scale horizontally (more servers)
Advantage:
- Capacity grows in 2-GPU steps
- No need to replace existing hardware
- Flexible expansion from S1 via S2 up to the data center configurations DC4 and DC8
All configurations at a glance: server page
Scale vertically (more powerful GPUs)
Advantage:
- Higher performance and more memory per server
- Less management overhead
The software supports the full NVIDIA spectrum from the RTX PRO 6000 Blackwell up to the data center GPUs H100, H200, and GH200. Details on supported GPUs: software page
Hybrid approach
Combination:
- Base servers for standard workloads
- High-performance servers for critical applications
- Central management via the onprem.ai software
Scaling without data loss
Advantages of the modular architecture:
- Servers can be added without changing existing configurations
- Models remain available on all servers
- No data migration required
- Rolling updates without downtime, self-healing on failures
Costs when scaling
On-premise:
- Additional hardware only when needed
- No usage dependency, predictable costs
Cloud:
- Every additional user means more token costs
- Unpredictable bills
The more users, the faster your own hardware pays for itself. What that means for your team in concrete terms is shown by the cost calculator with break-even and residual value over three years.
Next steps
- Open the cost calculator and plan your scaling
- Contact us for individual scaling options
Sources and further information:
- Start small and scale (scaling strategies)
- On-premise AI for SMEs (scaling and clusters)