Member of Inception Program

How does the on-premise AI solution scale?

On-premise AI scales flexibly and step by step: you can scale horizontally (add more servers) or scale vertically (more powerful GPUs). The onprem.ai software manages all systems centrally and distributes the load automatically.

Scaling strategies

Scale horizontally (more servers)

Advantage:

  • Capacity grows in 2-GPU steps
  • No need to replace existing hardware
  • Flexible expansion from S1 via S2 up to the data center configurations DC4 and DC8

All configurations at a glance: server page

Scale vertically (more powerful GPUs)

Advantage:

  • Higher performance and more memory per server
  • Less management overhead

The software supports the full NVIDIA spectrum from the RTX PRO 6000 Blackwell up to the data center GPUs H100, H200, and GH200. Details on supported GPUs: software page

Hybrid approach

Combination:

  • Base servers for standard workloads
  • High-performance servers for critical applications
  • Central management via the onprem.ai software

Scaling without data loss

Advantages of the modular architecture:

  • Servers can be added without changing existing configurations
  • Models remain available on all servers
  • No data migration required
  • Rolling updates without downtime, self-healing on failures

Costs when scaling

On-premise:

  • Additional hardware only when needed
  • No usage dependency, predictable costs

Cloud:

  • Every additional user means more token costs
  • Unpredictable bills

The more users, the faster your own hardware pays for itself. What that means for your team in concrete terms is shown by the cost calculator with break-even and residual value over three years.

Next steps


Sources and further information: