Member of Inception Program

Schweizer FlaggeSwiss Engineering Server S2: enterprise AI for your whole team

Our most popular server: 2x NVIDIA RTX PRO 6000 Blackwell with 192 GB VRAM, preinstalled onprem.ai software, and the latest LLM models. Runs standalone, scales as a cluster.

One server that serves an entire team

Designed for teams in regulated industries such as legal, finance, and healthcare. All data is processed exclusively on premises, fully air-gapped if required. Thanks to OpenAI-compatible APIs, the S2 replaces cloud AI services without code changes.

onprem.ai Server S2
GPUs
2x NVIDIA RTX PRO 6000 Blackwell
VRAM
192 GB @ 1.8 TB/s
AI performance
8000 AI TOPS
Power draw
~1.8 kW max
Form factor
4U rack
Users
5 - 20 concurrent

Built for speed

A single S2 runs DeepSeek v4 flash with the full 1M context and serves an entire team at once. The key measurements:

1,300 tokens/s

Text generation under parallel requests, 200 tokens/s single-stream

10,000 tokens/s

Input processing under parallel requests

1M tokens

Context window, with a 1.4M token KV cache

44 users

Concurrent in office use, plus 10 parallel coding agents

60,000 pages/h

Document classification in continuous operation

3.4 bn tokens/month

Generated tokens, plus 26 bn processed input tokens

Preliminary measurements with DeepSeek v4 flash on the Server S2, as of August 2026. Verification in progress.