Skip to Content
DE
NVIDIA H200 NVL · Hopper · 141 GB HBM3e
5-year NVIDIA AI Enterprise included In stock — ships from stock / typ. 1–2 weeks from €26,500 excl. VAT

Buy NVIDIA H200 — 141 GB HBM3e, ready to deploy

The data-center GPU for LLM inference in standard servers: 141 GB HBM3e at 4.8 TB/s — Llama 70B runs in FP8 on a single card. As a PCIe card, the H200 NVL fits into existing Gen5 servers, air-cooled, 1 to 8 cards. Nelpx delivers cards, NVLink bridges and, on request, the matching server — quote within 48 h.

IT systems house for HPC & AI infrastructure · DACH region · Daily updated pricing & availability

NVIDIA H200 NVL — PCIe data-center GPU with 141 GB HBM3e, dual-slot, passively cooled
NVIDIA H200 NVL — PCIe Gen5, dual-slot, passive. Image: NVIDIA/PNY.
141 GBHBM3e per card
4.8 TB/sMemory bandwidth
3.3 PFLOPSFP8 with sparsity (dense ~1.7)
900 GB/sNVLink via 2-/4-way bridge
Quick answer

What does an NVIDIA H200 cost — and what is it?

The NVIDIA H200 is the Hopper data-center GPU with 141 GB of HBM3e and 4.8 TB/s of memory bandwidth — the individually purchasable PCIe card is called the H200 NVL. At Nelpx it starts at €26,500 excl. VAT per card and is available at short notice (from stock up to typically 1–2 weeks). A 5-year NVIDIA AI Enterprise subscription is included in the price. The card runs air-cooled in standard PCIe Gen5 servers with 1 to 8 GPUs and holds a 70B model in FP8 entirely on one card — making it the price-performance buy of 2026 for on-premises LLM inference, fine-tuning and RAG.

Price
From €26,500 net per card — GPU prices are currently volatile; Nelpx quotes at daily rates, usually within 48 h.
In stock
Available through German distribution — from stock up to typically 1–2 weeks; Nelpx checks inventory daily.
Software included
5 years of NVIDIA AI Enterprise (incl. NIM) per card — activated via the GPU serial number.

As of: 27 July 2026 · Sources: NVIDIA H200 product page & datasheet, German distribution pricing — manufacturer claims are labelled. Details: H200 datasheet.

The product

H200 NVL: data-center power for standard servers

No baseboard, no liquid cooling, no facility project: the NVL card brings the H200 into any qualified PCIe Gen5 server — and scales via NVLink bridges.

01 · The cardH200 NVL (PCIe)

141 GB HBM3e, 4.8 TB/s, PCIe Gen5 x16, dual-slot, passively cooled, up to 600 W (configurable). MIG-capable (7× 16.5 GB), Confidential Computing, 7 NVDEC + 7 JPEG decoders for video and multimodal pipelines.

02 · The poolNVLink bridges

2- or 4-way bridges connect cards at 900 GB/s per GPU — a 4-way pool acts like one accelerator with 564 GB of HBM3e. Up to 8 cards per server (2× 4-way domains) per NVIDIA reference architecture.

03 · The software5 years of AI Enterprise included

Every H200 NVL includes a 5-year NVIDIA AI Enterprise subscription incl. NIM microservices — production stack, validated containers and software support from day one, activated via the GPU serial number.

GPU memory141 GB HBM3e · 4.8 TB/s
Tensor performanceFP8/INT8: 3,341 TFLOPS/TOPS sparse (~1,671 dense) · FP16/BF16: 1,671 sparse · TF32: 835 sparse
HPC performanceFP64: 30 TFLOPS · FP64 Tensor: 60 TFLOPS · FP32: 60 TFLOPS
NVLink2-/4-way bridge, 900 GB/s per GPU · PCIe Gen5: 128 GB/s
MIGup to 7 instances of 16.5 GB each
Form factor / coolingPCIe Gen5 x16 · dual-slot · passive (server airflow) · TDP up to 600 W configurable
Server integrationNVIDIA-Certified/MGX systems with 1–8 cards — incl. Supermicro, GIGABYTE G293/G493, Dell, HPE, Lenovo
SXM distinctionThe H200 SXM (up to 700 W, NVLink across 8 GPUs) only ships in HGX/DGX systems — Nelpx supplies those too

Values: current NVIDIA H200 product page (NVL column) and H200 NVL datasheet, as of 27 Jul 2026. Beware of third-party sources: comparison portals sometimes list SXM values (3,958 TFLOPS / 700 W) for the NVL.

Reality check

Does my model fit into 141 GB?

The most important question before buying — here are the common models with memory footprints (guide values: weights + reserve for KV cache; long contexts need more):

Llama 3.3 70B · FP8The standard case: comfortable on 1 card
~75 GB · 1 card ✓
Qwen2.5-72B · FP8incl. vision variant
~72 GB · 1 card ✓
Mistral Large 123B · FP8tight, but it fits
~123 GB · 1 card ✓
Llama 70B · FP16full precision — little KV-cache headroom
~141 GB · 1 card (tight)
Mistral Large 123B · FP16needs the 2-way NVLink pool (282 GB)
~246 GB · 2 cards
DeepSeek-R1 671B (MoE) · FP8a case for the 8-GPU system → HGX class
>700 GB · 8-GPU system

Guide values from common VRAM calculators (apxml.com, gigagpu.com), as of 27 Jul 2026 — rule of thumb: FP16 ≈ 2 bytes, FP8 ≈ 1 byte, INT4 ≈ 0.5 bytes per parameter, plus 10–20 % overhead. Alternatively, up to 7 smaller models (7B/14B) run in parallel on one card via MIG.

H200 configurator

Your H200 configuration.
Requested in 2 minutes.

Choose quantity, NVLink, server and services — download as a PDF or send directly to Nelpx. Quote at the current daily price, usually within 48 hours.

01 · GPU fixed — H200 NVL, individually purchasable
NVIDIA H200 NVL · 141 GB HBM3e · 4.8 TB/s · FP8 3,341 TFLOPS (sparse) · PCIe Gen5, dual-slot, passive, up to 600 W · incl. 5 years of NVIDIA AI Enterprise
02 · Number of cards

= 141 GB HBM3e · up to ~0.6 kW GPU load

03 · NVLink pool bridges as accessories — from 2 cards
04 · Deployment environment
Existing serverNelpx checks compatibility free of charge
With a new servere.g. GIGABYTE G293/G493
OpenNelpx recommends to match your workload
05 · Integration & services
Delivery onlycard(s) + bridges + licence
Install + setupincl. driver & CUDA stack
Turnkeyserver, stack, licence activation
06 · Procurement model
Purchasestandard — pays for itself under sustained load vs. cloud rental
Leasingpredictable monthly rates
OpenNelpx compares both paths for you
07 · Your details for the PDF & quote
Technical data

H200 NVL compared: H100 NVL & RTX PRO 6000

The H200 advantage is memory — exactly where LLM inference is limited. The compute engines are practically identical to the H100.

GPU memory per card

GB — HBM3e vs. HBM3 vs. GDDR7

H200 NVL141 GB HBM3e
RTX PRO 6000 SE96 GB GDDR7
H100 NVL94 GB HBM3

Memory bandwidth

TB/s — decisive for token throughput

H200 NVL4.8 TB/s
H100 NVL3.9 TB/s
RTX PRO 6000 SE~1.6 TB/s

GPU-to-GPU interconnect

GB/s — NVLink bridge vs. PCIe only

H200 NVL (bridge)900 GB/s
H100 NVL (bridge)600 GB/s
RTX PRO 6000 SE (PCIe)~128 GB/s

One 70B model (FP8, ~75 GB)

How many cards does it take? — fewer is better

H200 NVL1 card (53 % used)
H100 NVL / RTX PRO 60001 card (packed) up to 2 cards

Values: NVIDIA product pages H200/H100 (NVL columns) and RTX PRO 6000 Blackwell Server Edition data, as of 27 Jul 2026. NVIDIA claim “up to 1.7× LLM inference vs. H100 NVL” is a manufacturer figure. The RTX PRO 6000 SE is the far less expensive PCIe alternative without HBM/NVLink — product page.

Show all technical data (NVIDIA H200 NVL)
SpecificationNVIDIA H200 NVL
ArchitectureNVIDIA Hopper · 4th-gen Tensor Cores · Transformer Engine
GPU memory141 GB HBM3e · 4.8 TB/s bandwidth
Tensor performanceFP8/INT8: 3,341 TFLOPS/TOPS sparse (~1,671 dense) · FP16/BF16: 1,671 sparse · TF32: 835 sparse
HPC performanceFP64: 30 TFLOPS · FP64 Tensor: 60 TFLOPS · FP32: 60 TFLOPS
NVLink2-/4-way bridge: 900 GB/s per GPU (4-way pool: 564 GB shared) · PCIe Gen5: 128 GB/s
MIGup to 7 instances of 16.5 GB each
Decoders7× NVDEC · 7× JPEG
Confidential Computingsupported
Form factorPCIe Gen5 x16 · dual-slot · passively cooled (server airflow required)
TDPup to 600 W (configurable) · power via 16-pin connector
SoftwareNVIDIA AI Enterprise: 5-year subscription included (incl. NIM) · CUDA · vGPU support (AI Enterprise Infra 6.1+)
ServersNVIDIA-Certified/MGX: 1–8 cards (2× 4-way NVLink domains) — Supermicro, GIGABYTE G293/G493, Dell, HPE, Lenovo
SXM distinctionH200 SXM: up to 700 W, MIG at 18 GB, FP8 3,958 TFLOPS sparse — HGX/DGX systems only
Launch / statusNVL announced 11/2024 (SC24), shipping since 12/2024 · in the regular 2026 portfolio and available in the EU
WarrantyManufacturer warranty via the distribution channel (typ. 3 years) — Nelpx states warranty terms in the quote
Scaling

From one card to an 8-GPU server

The H200 NVL grows with your project — from a PoC on one card to a full PCIe GPU server per NVIDIA reference architecture.

Entry1 card141 GB — Llama 70B (FP8) entirely on one GPU, MIG for multi-tenant
Duo2 cards282 GB via 2-way NVLink bridge — e.g. 123B models in FP16
Pool4 cards564 GB via 4-way bridge, 900 GB/s per GPU — NVIDIA’s reference sweet spot
Full server8 cards2× 4-way domains, ~1.1 TB HBM3e — e.g. in a GIGABYTE G493 or Supermicro
Card, server or complete system — Nelpx delivers all three paths

Individual cards with a compatibility check for your existing server, fully configured PCIe GPU servers, or the full HGX system: Nelpx prices the options transparently against each other — with daily updated pricing and a quote usually within 48 hours.

Configure now
Use cases

Where does the H200 pay off?

01LLM inference, 70B class

The core use case: 70B models in FP8 on one card, with 4.8 TB/s of bandwidth for high token throughput and room for long contexts — internal RAG, chatbots, copilots.

02Upgrade existing servers

The underrated path to AI capacity: extend existing PCIe Gen5 servers with 1–8 cards instead of buying a new system — air-cooled, no facility rebuild.

03Fine-tuning & LoRA

141 GB holds model, gradients and batches of the 7B–70B class — fine-tuning runs that fail on 80 GB cards complete here.

04Multi-tenant via MIG

Up to 7 isolated instances of 16.5 GB each: several teams, models or environments on one card — cleanly orchestrated with the included AI Enterprise software.

05GDPR-compliant on-prem AI

Prompts, documents and models stay in-house — the alternative to US cloud APIs, with break-even vs. typical cloud rental as a guide value after 10–12 months of sustained load.

06HPC & simulation

60 TFLOPS of FP64 Tensor performance and the fastest memory subsystem in its class — for CFD, molecular dynamics and scientific computing alongside AI workloads.

Positioning

Buy an H200 — or something else?

The honest decision guide — because not every project needs the same GPU:

Why the H200 NVL

  • 141 GB on one card: the entire 70B class in FP8 without sharding — at 4.8 TB/s bandwidth.
  • Available now: from stock up to typ. 1–2 weeks — while Blackwell systems sit in allocation.
  • Standard server instead of a facility project: PCIe Gen5, air-cooled, 1–8 cards — also as a retrofit.
  • Software included: 5 years of NVIDIA AI Enterprise per card — an expensive add-on with other SKUs.
  • Most mature stack: Hopper has the largest installed base in 2026 — every framework runs from day one.

The alternatives, honestly

  • RTX PRO 6000 Blackwell SE: 96 GB GDDR7 at a much lower price — good for smaller models and mixed graphics/AI use, but without HBM and NVLink (product page).
  • HGX B200: from cluster class upwards, with NVLink across 8 GPUs and 1.4 TB HBM3e per system (B200 configurator).
  • HGX B300 / DGX B300: for models beyond the NVL pools, FP4 mass serving and maximum density (HGX B300 · DGX B300).
  • Prototype first: for development below the data-center class there is the DGX Spark.
  • Complete GPU server: card plus matching platform in one project — via the GPU server configurator or directly the 8-GPU server GIGABYTE G493.
FAQ

Frequently asked questions about the NVIDIA H200

What does an NVIDIA H200 cost?

At Nelpx, the NVIDIA H200 NVL (141 GB, PCIe) starts at €26,500 excl. VAT per card — including the 5-year NVIDIA AI Enterprise subscription. In mid-2026 the German market price is roughly €26,000–27,500 net depending on dealer and daily availability; H200 SXM systems (HGX/DGX) cost considerably more. As GPU prices are volatile, Nelpx quotes at daily rates — usually within 48 hours via the configurator on this page.

Is the NVIDIA H200 in stock and how long is delivery?

Yes — the H200 NVL is available through German distribution, typically from stock up to about 1–2 weeks. The H200 remains in NVIDIA’s regular 2026 portfolio and supply to the European market continues — the Blackwell backlog additionally supports demand for immediately available Hopper. For larger quantities Nelpx checks allocation daily — the earlier the enquiry, the better the date.

H200 NVL or H200 SXM — which variant do I need?

The H200 NVL is the individually purchasable PCIe card: dual-slot, passively cooled, up to 600 W, for standard PCIe Gen5 servers with 1 to 8 cards — ideal for retrofits and flexible quantities. The H200 SXM is soldered onto HGX baseboards (4 or 8 GPUs) or ships in the DGX H200, offers up to 700 W and NVLink across all 8 GPUs. Rule of thumb: individual cards and existing servers → NVL; a new 8-GPU system with maximum GPU-to-GPU bandwidth → SXM system. Nelpx supplies both.

What is the difference between H100 and H200?

The H200 is the memory-rich Hopper GPU: 141 GB HBM3e instead of 80/94 GB and 4.8 TB/s instead of 3.35/3.9 TB/s of bandwidth — the compute engines are practically identical. For memory-bound workloads such as LLM inference that is exactly what delivers the boost: NVIDIA quotes up to 1.7× more LLM throughput for the H200 NVL versus the H100 NVL (manufacturer figure). Add faster NVLink (900 instead of 600 GB/s via bridge). New H100 cards are barely available in German retail any more.

Does Llama 70B run on a single H200?

Yes — in FP8 quantisation a 70B model like Llama 3.3 occupies around 70–75 GB and runs comfortably on one H200 with plenty of reserve for KV cache and long contexts. Even FP16 (~141 GB) just about fits on the card. Qwen2.5-72B (FP8 ~72 GB) and Mistral Large 123B (FP8 ~123 GB) also run on a single card — only models beyond 141 GB need the NVLink pool or an 8-GPU system. All figures are guide values for weights plus KV-cache reserve.

How much power does the NVIDIA H200 draw?

The H200 NVL is configurable up to 600 W (the SXM variant up to 700 W). The card is passively cooled — the server must provide the airflow, which is why it belongs only in qualified chassis. As a planning figure: 4 cards ≈ 2.4 kW, 8 cards ≈ 4.8 kW for the GPUs alone, plus the server. That remains classic air cooling in an existing rack — one of the main reasons to choose the H200 in 2026 over liquid-cooled systems.

Does the H200 NVL fit into my existing server?

If the server offers PCIe Gen5 x16, a dual-slot bay, sufficient front-to-back airflow and PSU headroom for up to 600 W per card — very likely yes. NVIDIA lists qualified systems (NVIDIA-Certified/MGX) from Supermicro, GIGABYTE (G293/G493), Dell, HPE and Lenovo with up to 8 cards. Nelpx checks your server configuration free of charge before you order — and can supply the matching GPU server as well.

How does NVLink work on the H200 NVL?

Plug-on NVLink bridges connect 2 or 4 cards at 900 GB/s per GPU — a 4-way pool then acts like one accelerator with 564 GB of HBM3e. Cards without a bridge communicate via PCIe Gen5 at 128 GB/s. For models above 141 GB or tensor parallelism, the bridge is therefore the most important accessory decision — selectable in the configurator; Nelpx supplies the matching bridges.

Does the H200 support MIG (Multi-Instance GPU)?

Yes — the H200 NVL can be split into up to 7 fully isolated MIG instances of 16.5 GB each. One card then serves several teams or services in parallel: for example multiple 7B/14B models at once, or separate dev/test/prod instances. Together with the included NVIDIA AI Enterprise software, this is the building block for multi-tenant inference on-premises.

Is NVIDIA AI Enterprise included with the H200?

Yes — every H200 NVL includes a 5-year NVIDIA AI Enterprise subscription incl. NIM microservices, activated via the GPU serial number. That is a real cost advantage over SXM systems and other cards where the software is licensed separately. Production stack, enterprise software support and validated containers are on board from day one.

Is the H200 still worth it in 2026 — or go straight to Blackwell?

For models up to 141 GB — the entire 7B-to-70B class in FP8 — the H200 NVL is the price-performance buy of 2026: available now, air-cooled in standard servers, the most mature software stack, and roughly half the cost of a comparable Blackwell system slot. Blackwell (HGX B200/B300) wins for larger models, FP4 mass serving and cluster training. Honest rule of thumb: if the workload fits on 1–8 NVL cards, buy the H200; if it grows beyond that, plan an HGX system straight away — Nelpx prices both paths in the quote.

Why buy an H200 instead of renting GPU cloud?

Two reasons: economics and data sovereignty. H200 cloud instances from established providers typically cost 3 to 4.50 US dollars per GPU-hour (discount providers partly below) — under sustained load at that price level, a purchased card pays for itself in roughly 10 to 12 months as a guide value, then keeps earning for years. And on-premises, prompts, training data and models stay fully in-house — which considerably simplifies GDPR assessment and EU AI Act classification. Nelpx also supports financing and leasing.

Start your project

Your H200 — requested in 2 minutes

From €26,500 excl. VAT per card, including 5 years of NVIDIA AI Enterprise. GPU prices are volatile and stock changes daily: enquire now to lock in price and card. Nelpx delivers individually, as an NVLink pool or with a matching server.

Configure now & request a quote

Daily updated pricing · Quote usually within 48 h · Delivery & integration across DACH