Buy NVIDIA H200 — 141 GB HBM3e, ready to deploy
The data-center GPU for LLM inference in standard servers: 141 GB HBM3e at 4.8 TB/s — Llama 70B runs in FP8 on a single card. As a PCIe card, the H200 NVL fits into existing Gen5 servers, air-cooled, 1 to 8 cards. Nelpx delivers cards, NVLink bridges and, on request, the matching server — quote within 48 h.
IT systems house for HPC & AI infrastructure · DACH region · Daily updated pricing & availability
What does an NVIDIA H200 cost — and what is it?
The NVIDIA H200 is the Hopper data-center GPU with 141 GB of HBM3e and 4.8 TB/s of memory bandwidth — the individually purchasable PCIe card is called the H200 NVL. At Nelpx it starts at €26,500 excl. VAT per card and is available at short notice (from stock up to typically 1–2 weeks). A 5-year NVIDIA AI Enterprise subscription is included in the price. The card runs air-cooled in standard PCIe Gen5 servers with 1 to 8 GPUs and holds a 70B model in FP8 entirely on one card — making it the price-performance buy of 2026 for on-premises LLM inference, fine-tuning and RAG.
- Price
- From €26,500 net per card — GPU prices are currently volatile; Nelpx quotes at daily rates, usually within 48 h.
- In stock
- Available through German distribution — from stock up to typically 1–2 weeks; Nelpx checks inventory daily.
- Software included
- 5 years of NVIDIA AI Enterprise (incl. NIM) per card — activated via the GPU serial number.
As of: 27 July 2026 · Sources: NVIDIA H200 product page & datasheet, German distribution pricing — manufacturer claims are labelled. Details: H200 datasheet.
H200 NVL: data-center power for standard servers
No baseboard, no liquid cooling, no facility project: the NVL card brings the H200 into any qualified PCIe Gen5 server — and scales via NVLink bridges.
141 GB HBM3e, 4.8 TB/s, PCIe Gen5 x16, dual-slot, passively cooled, up to 600 W (configurable). MIG-capable (7× 16.5 GB), Confidential Computing, 7 NVDEC + 7 JPEG decoders for video and multimodal pipelines.
2- or 4-way bridges connect cards at 900 GB/s per GPU — a 4-way pool acts like one accelerator with 564 GB of HBM3e. Up to 8 cards per server (2× 4-way domains) per NVIDIA reference architecture.
Every H200 NVL includes a 5-year NVIDIA AI Enterprise subscription incl. NIM microservices — production stack, validated containers and software support from day one, activated via the GPU serial number.
Values: current NVIDIA H200 product page (NVL column) and H200 NVL datasheet, as of 27 Jul 2026. Beware of third-party sources: comparison portals sometimes list SXM values (3,958 TFLOPS / 700 W) for the NVL.
Does my model fit into 141 GB?
The most important question before buying — here are the common models with memory footprints (guide values: weights + reserve for KV cache; long contexts need more):
Guide values from common VRAM calculators (apxml.com, gigagpu.com), as of 27 Jul 2026 — rule of thumb: FP16 ≈ 2 bytes, FP8 ≈ 1 byte, INT4 ≈ 0.5 bytes per parameter, plus 10–20 % overhead. Alternatively, up to 7 smaller models (7B/14B) run in parallel on one card via MIG.
Your H200 configuration.
Requested in 2 minutes.
Choose quantity, NVLink, server and services — download as a PDF or send directly to Nelpx. Quote at the current daily price, usually within 48 hours.
= 141 GB HBM3e · up to ~0.6 kW GPU load
H200 NVL compared: H100 NVL & RTX PRO 6000
The H200 advantage is memory — exactly where LLM inference is limited. The compute engines are practically identical to the H100.
GPU memory per card
GB — HBM3e vs. HBM3 vs. GDDR7
Memory bandwidth
TB/s — decisive for token throughput
GPU-to-GPU interconnect
GB/s — NVLink bridge vs. PCIe only
One 70B model (FP8, ~75 GB)
How many cards does it take? — fewer is better
Values: NVIDIA product pages H200/H100 (NVL columns) and RTX PRO 6000 Blackwell Server Edition data, as of 27 Jul 2026. NVIDIA claim “up to 1.7× LLM inference vs. H100 NVL” is a manufacturer figure. The RTX PRO 6000 SE is the far less expensive PCIe alternative without HBM/NVLink — product page.
Show all technical data (NVIDIA H200 NVL)
| Specification | NVIDIA H200 NVL |
|---|---|
| Architecture | NVIDIA Hopper · 4th-gen Tensor Cores · Transformer Engine |
| GPU memory | 141 GB HBM3e · 4.8 TB/s bandwidth |
| Tensor performance | FP8/INT8: 3,341 TFLOPS/TOPS sparse (~1,671 dense) · FP16/BF16: 1,671 sparse · TF32: 835 sparse |
| HPC performance | FP64: 30 TFLOPS · FP64 Tensor: 60 TFLOPS · FP32: 60 TFLOPS |
| NVLink | 2-/4-way bridge: 900 GB/s per GPU (4-way pool: 564 GB shared) · PCIe Gen5: 128 GB/s |
| MIG | up to 7 instances of 16.5 GB each |
| Decoders | 7× NVDEC · 7× JPEG |
| Confidential Computing | supported |
| Form factor | PCIe Gen5 x16 · dual-slot · passively cooled (server airflow required) |
| TDP | up to 600 W (configurable) · power via 16-pin connector |
| Software | NVIDIA AI Enterprise: 5-year subscription included (incl. NIM) · CUDA · vGPU support (AI Enterprise Infra 6.1+) |
| Servers | NVIDIA-Certified/MGX: 1–8 cards (2× 4-way NVLink domains) — Supermicro, GIGABYTE G293/G493, Dell, HPE, Lenovo |
| SXM distinction | H200 SXM: up to 700 W, MIG at 18 GB, FP8 3,958 TFLOPS sparse — HGX/DGX systems only |
| Launch / status | NVL announced 11/2024 (SC24), shipping since 12/2024 · in the regular 2026 portfolio and available in the EU |
| Warranty | Manufacturer warranty via the distribution channel (typ. 3 years) — Nelpx states warranty terms in the quote |
From one card to an 8-GPU server
The H200 NVL grows with your project — from a PoC on one card to a full PCIe GPU server per NVIDIA reference architecture.
Individual cards with a compatibility check for your existing server, fully configured PCIe GPU servers, or the full HGX system: Nelpx prices the options transparently against each other — with daily updated pricing and a quote usually within 48 hours.
Where does the H200 pay off?
The core use case: 70B models in FP8 on one card, with 4.8 TB/s of bandwidth for high token throughput and room for long contexts — internal RAG, chatbots, copilots.
The underrated path to AI capacity: extend existing PCIe Gen5 servers with 1–8 cards instead of buying a new system — air-cooled, no facility rebuild.
141 GB holds model, gradients and batches of the 7B–70B class — fine-tuning runs that fail on 80 GB cards complete here.
Up to 7 isolated instances of 16.5 GB each: several teams, models or environments on one card — cleanly orchestrated with the included AI Enterprise software.
Prompts, documents and models stay in-house — the alternative to US cloud APIs, with break-even vs. typical cloud rental as a guide value after 10–12 months of sustained load.
60 TFLOPS of FP64 Tensor performance and the fastest memory subsystem in its class — for CFD, molecular dynamics and scientific computing alongside AI workloads.
Buy an H200 — or something else?
The honest decision guide — because not every project needs the same GPU:
Why the H200 NVL
- 141 GB on one card: the entire 70B class in FP8 without sharding — at 4.8 TB/s bandwidth.
- Available now: from stock up to typ. 1–2 weeks — while Blackwell systems sit in allocation.
- Standard server instead of a facility project: PCIe Gen5, air-cooled, 1–8 cards — also as a retrofit.
- Software included: 5 years of NVIDIA AI Enterprise per card — an expensive add-on with other SKUs.
- Most mature stack: Hopper has the largest installed base in 2026 — every framework runs from day one.
The alternatives, honestly
- RTX PRO 6000 Blackwell SE: 96 GB GDDR7 at a much lower price — good for smaller models and mixed graphics/AI use, but without HBM and NVLink (product page).
- HGX B200: from cluster class upwards, with NVLink across 8 GPUs and 1.4 TB HBM3e per system (B200 configurator).
- HGX B300 / DGX B300: for models beyond the NVL pools, FP4 mass serving and maximum density (HGX B300 · DGX B300).
- Prototype first: for development below the data-center class there is the DGX Spark.
- Complete GPU server: card plus matching platform in one project — via the GPU server configurator or directly the 8-GPU server GIGABYTE G493.
Frequently asked questions about the NVIDIA H200
What does an NVIDIA H200 cost?
At Nelpx, the NVIDIA H200 NVL (141 GB, PCIe) starts at €26,500 excl. VAT per card — including the 5-year NVIDIA AI Enterprise subscription. In mid-2026 the German market price is roughly €26,000–27,500 net depending on dealer and daily availability; H200 SXM systems (HGX/DGX) cost considerably more. As GPU prices are volatile, Nelpx quotes at daily rates — usually within 48 hours via the configurator on this page.
Is the NVIDIA H200 in stock and how long is delivery?
Yes — the H200 NVL is available through German distribution, typically from stock up to about 1–2 weeks. The H200 remains in NVIDIA’s regular 2026 portfolio and supply to the European market continues — the Blackwell backlog additionally supports demand for immediately available Hopper. For larger quantities Nelpx checks allocation daily — the earlier the enquiry, the better the date.
H200 NVL or H200 SXM — which variant do I need?
The H200 NVL is the individually purchasable PCIe card: dual-slot, passively cooled, up to 600 W, for standard PCIe Gen5 servers with 1 to 8 cards — ideal for retrofits and flexible quantities. The H200 SXM is soldered onto HGX baseboards (4 or 8 GPUs) or ships in the DGX H200, offers up to 700 W and NVLink across all 8 GPUs. Rule of thumb: individual cards and existing servers → NVL; a new 8-GPU system with maximum GPU-to-GPU bandwidth → SXM system. Nelpx supplies both.
What is the difference between H100 and H200?
The H200 is the memory-rich Hopper GPU: 141 GB HBM3e instead of 80/94 GB and 4.8 TB/s instead of 3.35/3.9 TB/s of bandwidth — the compute engines are practically identical. For memory-bound workloads such as LLM inference that is exactly what delivers the boost: NVIDIA quotes up to 1.7× more LLM throughput for the H200 NVL versus the H100 NVL (manufacturer figure). Add faster NVLink (900 instead of 600 GB/s via bridge). New H100 cards are barely available in German retail any more.
Does Llama 70B run on a single H200?
Yes — in FP8 quantisation a 70B model like Llama 3.3 occupies around 70–75 GB and runs comfortably on one H200 with plenty of reserve for KV cache and long contexts. Even FP16 (~141 GB) just about fits on the card. Qwen2.5-72B (FP8 ~72 GB) and Mistral Large 123B (FP8 ~123 GB) also run on a single card — only models beyond 141 GB need the NVLink pool or an 8-GPU system. All figures are guide values for weights plus KV-cache reserve.
How much power does the NVIDIA H200 draw?
The H200 NVL is configurable up to 600 W (the SXM variant up to 700 W). The card is passively cooled — the server must provide the airflow, which is why it belongs only in qualified chassis. As a planning figure: 4 cards ≈ 2.4 kW, 8 cards ≈ 4.8 kW for the GPUs alone, plus the server. That remains classic air cooling in an existing rack — one of the main reasons to choose the H200 in 2026 over liquid-cooled systems.
Does the H200 NVL fit into my existing server?
If the server offers PCIe Gen5 x16, a dual-slot bay, sufficient front-to-back airflow and PSU headroom for up to 600 W per card — very likely yes. NVIDIA lists qualified systems (NVIDIA-Certified/MGX) from Supermicro, GIGABYTE (G293/G493), Dell, HPE and Lenovo with up to 8 cards. Nelpx checks your server configuration free of charge before you order — and can supply the matching GPU server as well.
How does NVLink work on the H200 NVL?
Plug-on NVLink bridges connect 2 or 4 cards at 900 GB/s per GPU — a 4-way pool then acts like one accelerator with 564 GB of HBM3e. Cards without a bridge communicate via PCIe Gen5 at 128 GB/s. For models above 141 GB or tensor parallelism, the bridge is therefore the most important accessory decision — selectable in the configurator; Nelpx supplies the matching bridges.
Does the H200 support MIG (Multi-Instance GPU)?
Yes — the H200 NVL can be split into up to 7 fully isolated MIG instances of 16.5 GB each. One card then serves several teams or services in parallel: for example multiple 7B/14B models at once, or separate dev/test/prod instances. Together with the included NVIDIA AI Enterprise software, this is the building block for multi-tenant inference on-premises.
Is NVIDIA AI Enterprise included with the H200?
Yes — every H200 NVL includes a 5-year NVIDIA AI Enterprise subscription incl. NIM microservices, activated via the GPU serial number. That is a real cost advantage over SXM systems and other cards where the software is licensed separately. Production stack, enterprise software support and validated containers are on board from day one.
Is the H200 still worth it in 2026 — or go straight to Blackwell?
For models up to 141 GB — the entire 7B-to-70B class in FP8 — the H200 NVL is the price-performance buy of 2026: available now, air-cooled in standard servers, the most mature software stack, and roughly half the cost of a comparable Blackwell system slot. Blackwell (HGX B200/B300) wins for larger models, FP4 mass serving and cluster training. Honest rule of thumb: if the workload fits on 1–8 NVL cards, buy the H200; if it grows beyond that, plan an HGX system straight away — Nelpx prices both paths in the quote.
Why buy an H200 instead of renting GPU cloud?
Two reasons: economics and data sovereignty. H200 cloud instances from established providers typically cost 3 to 4.50 US dollars per GPU-hour (discount providers partly below) — under sustained load at that price level, a purchased card pays for itself in roughly 10 to 12 months as a guide value, then keeps earning for years. And on-premises, prompts, training data and models stay fully in-house — which considerably simplifies GDPR assessment and EU AI Act classification. Nelpx also supports financing and leasing.
Your H200 — requested in 2 minutes
From €26,500 excl. VAT per card, including 5 years of NVIDIA AI Enterprise. GPU prices are volatile and stock changes daily: enquire now to lock in price and card. Nelpx delivers individually, as an NVLink pool or with a matching server.
Configure now & request a quoteDaily updated pricing · Quote usually within 48 h · Delivery & integration across DACH
Buy NVIDIA H200 — 141 GB HBM3e, ready to deploy
The data-center GPU for LLM inference in standard servers: 141 GB HBM3e at 4.8 TB/s — Llama 70B runs in FP8 on a single card. As a PCIe card, the H200 NVL fits into existing Gen5 servers, air-cooled, 1 to 8 cards. Nelpx delivers cards, NVLink bridges and, on request, the matching server — quote within 48 h.
IT systems house for HPC & AI infrastructure · DACH region · Daily updated pricing & availability
What does an NVIDIA H200 cost — and what is it?
The NVIDIA H200 is the Hopper data-center GPU with 141 GB of HBM3e and 4.8 TB/s of memory bandwidth — the individually purchasable PCIe card is called the H200 NVL. At Nelpx it starts at €26,500 excl. VAT per card and is available at short notice (from stock up to typically 1–2 weeks). A 5-year NVIDIA AI Enterprise subscription is included in the price. The card runs air-cooled in standard PCIe Gen5 servers with 1 to 8 GPUs and holds a 70B model in FP8 entirely on one card — making it the price-performance buy of 2026 for on-premises LLM inference, fine-tuning and RAG.
- Price
- From €26,500 net per card — GPU prices are currently volatile; Nelpx quotes at daily rates, usually within 48 h.
- In stock
- Available through German distribution — from stock up to typically 1–2 weeks; Nelpx checks inventory daily.
- Software included
- 5 years of NVIDIA AI Enterprise (incl. NIM) per card — activated via the GPU serial number.
As of: 27 July 2026 · Sources: NVIDIA H200 product page & datasheet, German distribution pricing — manufacturer claims are labelled. Details: H200 datasheet.
H200 NVL: data-center power for standard servers
No baseboard, no liquid cooling, no facility project: the NVL card brings the H200 into any qualified PCIe Gen5 server — and scales via NVLink bridges.
141 GB HBM3e, 4.8 TB/s, PCIe Gen5 x16, dual-slot, passively cooled, up to 600 W (configurable). MIG-capable (7× 16.5 GB), Confidential Computing, 7 NVDEC + 7 JPEG decoders for video and multimodal pipelines.
2- or 4-way bridges connect cards at 900 GB/s per GPU — a 4-way pool acts like one accelerator with 564 GB of HBM3e. Up to 8 cards per server (2× 4-way domains) per NVIDIA reference architecture.
Every H200 NVL includes a 5-year NVIDIA AI Enterprise subscription incl. NIM microservices — production stack, validated containers and software support from day one, activated via the GPU serial number.
Values: current NVIDIA H200 product page (NVL column) and H200 NVL datasheet, as of 27 Jul 2026. Beware of third-party sources: comparison portals sometimes list SXM values (3,958 TFLOPS / 700 W) for the NVL.
Does my model fit into 141 GB?
The most important question before buying — here are the common models with memory footprints (guide values: weights + reserve for KV cache; long contexts need more):
Guide values from common VRAM calculators (apxml.com, gigagpu.com), as of 27 Jul 2026 — rule of thumb: FP16 ≈ 2 bytes, FP8 ≈ 1 byte, INT4 ≈ 0.5 bytes per parameter, plus 10–20 % overhead. Alternatively, up to 7 smaller models (7B/14B) run in parallel on one card via MIG.
Your H200 configuration.
Requested in 2 minutes.
Choose quantity, NVLink, server and services — download as a PDF or send directly to Nelpx. Quote at the current daily price, usually within 48 hours.
= 141 GB HBM3e · up to ~0.6 kW GPU load
H200 NVL compared: H100 NVL & RTX PRO 6000
The H200 advantage is memory — exactly where LLM inference is limited. The compute engines are practically identical to the H100.
GPU memory per card
GB — HBM3e vs. HBM3 vs. GDDR7
Memory bandwidth
TB/s — decisive for token throughput
GPU-to-GPU interconnect
GB/s — NVLink bridge vs. PCIe only
One 70B model (FP8, ~75 GB)
How many cards does it take? — fewer is better
Values: NVIDIA product pages H200/H100 (NVL columns) and RTX PRO 6000 Blackwell Server Edition data, as of 27 Jul 2026. NVIDIA claim “up to 1.7× LLM inference vs. H100 NVL” is a manufacturer figure. The RTX PRO 6000 SE is the far less expensive PCIe alternative without HBM/NVLink — product page.
Show all technical data (NVIDIA H200 NVL)
| Specification | NVIDIA H200 NVL |
|---|---|
| Architecture | NVIDIA Hopper · 4th-gen Tensor Cores · Transformer Engine |
| GPU memory | 141 GB HBM3e · 4.8 TB/s bandwidth |
| Tensor performance | FP8/INT8: 3,341 TFLOPS/TOPS sparse (~1,671 dense) · FP16/BF16: 1,671 sparse · TF32: 835 sparse |
| HPC performance | FP64: 30 TFLOPS · FP64 Tensor: 60 TFLOPS · FP32: 60 TFLOPS |
| NVLink | 2-/4-way bridge: 900 GB/s per GPU (4-way pool: 564 GB shared) · PCIe Gen5: 128 GB/s |
| MIG | up to 7 instances of 16.5 GB each |
| Decoders | 7× NVDEC · 7× JPEG |
| Confidential Computing | supported |
| Form factor | PCIe Gen5 x16 · dual-slot · passively cooled (server airflow required) |
| TDP | up to 600 W (configurable) · power via 16-pin connector |
| Software | NVIDIA AI Enterprise: 5-year subscription included (incl. NIM) · CUDA · vGPU support (AI Enterprise Infra 6.1+) |
| Servers | NVIDIA-Certified/MGX: 1–8 cards (2× 4-way NVLink domains) — Supermicro, GIGABYTE G293/G493, Dell, HPE, Lenovo |
| SXM distinction | H200 SXM: up to 700 W, MIG at 18 GB, FP8 3,958 TFLOPS sparse — HGX/DGX systems only |
| Launch / status | NVL announced 11/2024 (SC24), shipping since 12/2024 · in the regular 2026 portfolio and available in the EU |
| Warranty | Manufacturer warranty via the distribution channel (typ. 3 years) — Nelpx states warranty terms in the quote |
From one card to an 8-GPU server
The H200 NVL grows with your project — from a PoC on one card to a full PCIe GPU server per NVIDIA reference architecture.
Individual cards with a compatibility check for your existing server, fully configured PCIe GPU servers, or the full HGX system: Nelpx prices the options transparently against each other — with daily updated pricing and a quote usually within 48 hours.
Where does the H200 pay off?
The core use case: 70B models in FP8 on one card, with 4.8 TB/s of bandwidth for high token throughput and room for long contexts — internal RAG, chatbots, copilots.
The underrated path to AI capacity: extend existing PCIe Gen5 servers with 1–8 cards instead of buying a new system — air-cooled, no facility rebuild.
141 GB holds model, gradients and batches of the 7B–70B class — fine-tuning runs that fail on 80 GB cards complete here.
Up to 7 isolated instances of 16.5 GB each: several teams, models or environments on one card — cleanly orchestrated with the included AI Enterprise software.
Prompts, documents and models stay in-house — the alternative to US cloud APIs, with break-even vs. typical cloud rental as a guide value after 10–12 months of sustained load.
60 TFLOPS of FP64 Tensor performance and the fastest memory subsystem in its class — for CFD, molecular dynamics and scientific computing alongside AI workloads.
Buy an H200 — or something else?
The honest decision guide — because not every project needs the same GPU:
Why the H200 NVL
- 141 GB on one card: the entire 70B class in FP8 without sharding — at 4.8 TB/s bandwidth.
- Available now: from stock up to typ. 1–2 weeks — while Blackwell systems sit in allocation.
- Standard server instead of a facility project: PCIe Gen5, air-cooled, 1–8 cards — also as a retrofit.
- Software included: 5 years of NVIDIA AI Enterprise per card — an expensive add-on with other SKUs.
- Most mature stack: Hopper has the largest installed base in 2026 — every framework runs from day one.
The alternatives, honestly
- RTX PRO 6000 Blackwell SE: 96 GB GDDR7 at a much lower price — good for smaller models and mixed graphics/AI use, but without HBM and NVLink (product page).
- HGX B200: from cluster class upwards, with NVLink across 8 GPUs and 1.4 TB HBM3e per system (B200 configurator).
- HGX B300 / DGX B300: for models beyond the NVL pools, FP4 mass serving and maximum density (HGX B300 · DGX B300).
- Prototype first: for development below the data-center class there is the DGX Spark.
- Complete GPU server: card plus matching platform in one project — via the GPU server configurator or directly the 8-GPU server GIGABYTE G493.
Frequently asked questions about the NVIDIA H200
What does an NVIDIA H200 cost?
At Nelpx, the NVIDIA H200 NVL (141 GB, PCIe) starts at €26,500 excl. VAT per card — including the 5-year NVIDIA AI Enterprise subscription. In mid-2026 the German market price is roughly €26,000–27,500 net depending on dealer and daily availability; H200 SXM systems (HGX/DGX) cost considerably more. As GPU prices are volatile, Nelpx quotes at daily rates — usually within 48 hours via the configurator on this page.
Is the NVIDIA H200 in stock and how long is delivery?
Yes — the H200 NVL is available through German distribution, typically from stock up to about 1–2 weeks. The H200 remains in NVIDIA’s regular 2026 portfolio and supply to the European market continues — the Blackwell backlog additionally supports demand for immediately available Hopper. For larger quantities Nelpx checks allocation daily — the earlier the enquiry, the better the date.
H200 NVL or H200 SXM — which variant do I need?
The H200 NVL is the individually purchasable PCIe card: dual-slot, passively cooled, up to 600 W, for standard PCIe Gen5 servers with 1 to 8 cards — ideal for retrofits and flexible quantities. The H200 SXM is soldered onto HGX baseboards (4 or 8 GPUs) or ships in the DGX H200, offers up to 700 W and NVLink across all 8 GPUs. Rule of thumb: individual cards and existing servers → NVL; a new 8-GPU system with maximum GPU-to-GPU bandwidth → SXM system. Nelpx supplies both.
What is the difference between H100 and H200?
The H200 is the memory-rich Hopper GPU: 141 GB HBM3e instead of 80/94 GB and 4.8 TB/s instead of 3.35/3.9 TB/s of bandwidth — the compute engines are practically identical. For memory-bound workloads such as LLM inference that is exactly what delivers the boost: NVIDIA quotes up to 1.7× more LLM throughput for the H200 NVL versus the H100 NVL (manufacturer figure). Add faster NVLink (900 instead of 600 GB/s via bridge). New H100 cards are barely available in German retail any more.
Does Llama 70B run on a single H200?
Yes — in FP8 quantisation a 70B model like Llama 3.3 occupies around 70–75 GB and runs comfortably on one H200 with plenty of reserve for KV cache and long contexts. Even FP16 (~141 GB) just about fits on the card. Qwen2.5-72B (FP8 ~72 GB) and Mistral Large 123B (FP8 ~123 GB) also run on a single card — only models beyond 141 GB need the NVLink pool or an 8-GPU system. All figures are guide values for weights plus KV-cache reserve.
How much power does the NVIDIA H200 draw?
The H200 NVL is configurable up to 600 W (the SXM variant up to 700 W). The card is passively cooled — the server must provide the airflow, which is why it belongs only in qualified chassis. As a planning figure: 4 cards ≈ 2.4 kW, 8 cards ≈ 4.8 kW for the GPUs alone, plus the server. That remains classic air cooling in an existing rack — one of the main reasons to choose the H200 in 2026 over liquid-cooled systems.
Does the H200 NVL fit into my existing server?
If the server offers PCIe Gen5 x16, a dual-slot bay, sufficient front-to-back airflow and PSU headroom for up to 600 W per card — very likely yes. NVIDIA lists qualified systems (NVIDIA-Certified/MGX) from Supermicro, GIGABYTE (G293/G493), Dell, HPE and Lenovo with up to 8 cards. Nelpx checks your server configuration free of charge before you order — and can supply the matching GPU server as well.
How does NVLink work on the H200 NVL?
Plug-on NVLink bridges connect 2 or 4 cards at 900 GB/s per GPU — a 4-way pool then acts like one accelerator with 564 GB of HBM3e. Cards without a bridge communicate via PCIe Gen5 at 128 GB/s. For models above 141 GB or tensor parallelism, the bridge is therefore the most important accessory decision — selectable in the configurator; Nelpx supplies the matching bridges.
Does the H200 support MIG (Multi-Instance GPU)?
Yes — the H200 NVL can be split into up to 7 fully isolated MIG instances of 16.5 GB each. One card then serves several teams or services in parallel: for example multiple 7B/14B models at once, or separate dev/test/prod instances. Together with the included NVIDIA AI Enterprise software, this is the building block for multi-tenant inference on-premises.
Is NVIDIA AI Enterprise included with the H200?
Yes — every H200 NVL includes a 5-year NVIDIA AI Enterprise subscription incl. NIM microservices, activated via the GPU serial number. That is a real cost advantage over SXM systems and other cards where the software is licensed separately. Production stack, enterprise software support and validated containers are on board from day one.
Is the H200 still worth it in 2026 — or go straight to Blackwell?
For models up to 141 GB — the entire 7B-to-70B class in FP8 — the H200 NVL is the price-performance buy of 2026: available now, air-cooled in standard servers, the most mature software stack, and roughly half the cost of a comparable Blackwell system slot. Blackwell (HGX B200/B300) wins for larger models, FP4 mass serving and cluster training. Honest rule of thumb: if the workload fits on 1–8 NVL cards, buy the H200; if it grows beyond that, plan an HGX system straight away — Nelpx prices both paths in the quote.
Why buy an H200 instead of renting GPU cloud?
Two reasons: economics and data sovereignty. H200 cloud instances from established providers typically cost 3 to 4.50 US dollars per GPU-hour (discount providers partly below) — under sustained load at that price level, a purchased card pays for itself in roughly 10 to 12 months as a guide value, then keeps earning for years. And on-premises, prompts, training data and models stay fully in-house — which considerably simplifies GDPR assessment and EU AI Act classification. Nelpx also supports financing and leasing.
Your H200 — requested in 2 minutes
From €26,500 excl. VAT per card, including 5 years of NVIDIA AI Enterprise. GPU prices are volatile and stock changes daily: enquire now to lock in price and card. Nelpx delivers individually, as an NVLink pool or with a matching server.
Configure now & request a quoteDaily updated pricing · Quote usually within 48 h · Delivery & integration across DACH