Member of Inception Program

Which hardware does onprem.ai use?

We build on NVIDIA GPUs, from the enterprise server to the AI data center. The onprem.ai software supports the full spectrum:

  • NVIDIA RTX PRO 6000 Blackwell: Blackwell GPU with an excellent price-performance ratio for inference, the basis of the configurations S1, S2, DC4, and DC8
  • NVIDIA H100: proven Hopper data center GPU for demanding inference and fine-tuning workloads
  • NVIDIA H200: Hopper GPU with extended memory and higher bandwidth for more context and larger models per GPU
  • NVIDIA GH200: Grace Hopper superchip for the highest demands

Why this selection?

Professional LLM inference needs above all memory capacity and memory bandwidth. The GPUs we deploy offer both, across a spectrum from cost-efficient team servers to maximum-performance scenarios, with mature drivers and broad software support.

Our configurations

Entry starts with the Server S1 and grows in 2-GPU steps: S2 (2x RTX PRO 6000 Blackwell, 192 GB VRAM) for whole teams, DC4 and DC8 for data centers. All details on the server page.

Note: GPU prices are currently volatile. What your configuration costs is shown by the cost calculator, binding figures come with your quote.

Next steps