KPT-B200-SXM6-192G
B200 HGX
- Memory
- 192 GB HBM3e
- Node
- 8 GPUs / 8U
- In-node fabric
- NVLink 1.8 TB/s
- Draw
- ~14.3 kW
- Minimum
- 8 GPUs
Enterprise AI compute · three U.S. regions
We supply enterprise GPU nodes, accelerators and whole racks, two ways to take them: buy them and we manage them for you in a U.S. facility, or pay by the GPU-hour and stop whenever you like. Same hardware, same facility, same serial numbers in the same system.
Across three U.S. regions, every one serial-tracked.
Median for in-stock SKUs, including 72 hours of burn-in.
Dual-fed PDUs, N+1 cooling, written into the contract.
Grace period after a rental ends; beyond it, the day rate applies.
Own or rent
In this industry those are usually two companies doing two things: one sells hardware, the other rents compute. We do both, out of one inventory, one facility and one serial-number system. So you do not have to guess between “what if it sits idle” and “what if renting gets expensive” — both cost curves are plotted below.
Kepter Supply
You put up the capital and keep the asset. We put up the facility, the power, the hands and the spares.
Kepter Flow
Operating expense only. Pay for the hours you use; stopped machines do not bill.
This is not a sales line: at any point in a rental you can convert the machines you are using to a purchase, with 60% of the rent you have paid credited against the price. Same serial numbers — no migration, no reinstall, no downtime. Three months of real utilisation figures beat any procurement model.
Deployment
That figure includes 72 hours of burn-in — the machine has already run at full load for three days before it reaches you, and the report ships with it. We publish the duration of every leg, because a date is the only promise worth anything in this business.
01 · Commercial
Tell us the GPU model, the count, the region and roughly how long for; a plan comes back within 48 hours. Two hours after you confirm, stock is locked and serials are assigned on the spot.
02 · Logistics
Asset-tagged, ESD-packed, re-checked before racking. The waybill is trackable throughout, and cross-region transfers move on our own fleet.
03 · Facility
Assigned a cabinet, wired to both power feeds and to 400G. You will be told which cabinet, which U positions and which switch this batch hangs off.
04 · Acceptance
Seventy-two hours at full load with temperature, draw, memory ECC and NCCL bandwidth logged throughout. Passing is what gets it delivered; failing gets it swapped and re-run.
Inside the rack
Most compute suppliers give you a GPU model and a number. We have drawn a whole cabinet apart: which U positions your machines sit in, which switch they hang off, which power feed they draw, where the tag goes. All of it lands in your topology document on delivery day.
Hardware
Four main product lines, outright price and hourly rate written side by side. The price is not hidden behind “contact sales” — what needs discussing is terms and volume, not the unit price.
KPT-B200-SXM6-192G
B200 HGX
KPT-H200-SXM5-141G
H200 HGX
KPT-H100-SXM5-80G
H100 HGX
KPT-L40S-PCIE-48G
L40S PCIe
Rates are on-demand and exclude commitment discounts. Rent continuously for 3 months and the rate drops 20% automatically; at 12 months, 35%. No contract to sign, no application to file — it shows up on the bill. The outright price includes the first year of managed ops; from year two the managed fee is billed per rack and per kW. All prices are illustrative; the quote governs.
U.S. footprint
All hardware is inside the United States. Between regions we run dark fibre we lease ourselves, so cross-region traffic never touches the public internet and is never billed as egress — which is a real line item when you are training data-parallel across regions, not a decoration on a spec sheet.
All three regions are inside the United States; hardware, backups and logs do not leave it. We will supply the facility locations and a device list on request.
Dual-fed PDUs, N+1 cooling, and 99.9% power and cooling availability written into the SLA. Miss it and we credit a share of that month’s managed fee.
Outright customers can book a quarterly site visit and see their own cabinet and their own tags. It is the plainest way to verify that the asset is in your name.
Serial tracking
The serial is assigned the moment a unit arrives and does not change until it is retired. Whose hands it passed through, what it ran, how burn-in went, how many repairs, which U of which cabinet it is in right now — scan the QR code on the tag and it is all there.
Pricing
There is only one honest answer: it depends how long you will run it. The chart below runs the real price list for one 8-GPU H200 node out to 36 months; the curves cross in month 16. Renting wins below 16 months, owning wins above it — we do both, so we have no side to argue for.
| SKU | On demand | ≥ 3 months | ≥ 12 months | Minimum |
|---|---|---|---|---|
| B200-SXM6-192G | $8.90 | $7.12 | $5.79 | 8 GPUs |
| H200-SXM5-141G | $3.42 | $2.74 | $2.22 | 8 GPUs |
| H100-SXM5-80G | $2.28 | $1.82 | $1.48 | 8 GPUs |
| L40S-PCIE-48G | $1.15 | $0.92 | $0.75 | 1 GPU |
Figures are USD per GPU-hour. Discounts apply automatically, accrued by continuous usage, with nothing to sign or request. Stopped machines do not bill, but storage kept for a retained instance is charged at $0.04 / GB-month.
| Item | Us | Common in the industry |
|---|---|---|
| Inter-region traffic | $0 | $0.02–0.09 / GB |
| Egress (first 50 TB/month) | $0 | Metered |
| Starting and stopping instances | $0 | Minimum billing increment |
| Early exit | $0 | 30–100% of the remaining term |
| Downtime during a swap | Not billed | Billed as normal |
| Technical support | Included | Charged by tier |
The “common in the industry” column reflects general practice in published price lists and is not aimed at any one vendor. We have folded these into the base rate — the unit price does not look like the cheapest, but there is no second line on the bill.
Console
Whether a machine is your asset or a rental, it looks identical in the console: utilisation, temperature, draw, billing, serial number, rack position. The only difference is whether the tag in the corner reads Supply or Flow.
Utilisation, memory in use, temperature, draw and ECC counts, sampled every 15 seconds and kept for 13 months.
Every line on the invoice opens up to show which serial, which window and which rate produced it.
Anything the console does, the REST API does: provision, scale, query stock, pull telemetry, export invoices.
Three reminders before a term ends — 14 days, 7 days, 48 hours — with automatic renewal or release available.
FAQ
If the answer is not here, just ask. Our sales team does not work from a script: what they can answer they answer on the spot, and what they cannot, they tell you when they can.
Ask something specificYes. Same inventory pool, same 72-hour burn-in standard, same facility. The only difference is the commercial relationship: bought units are transferred into your name and taken out of the shared pool, rented units remain ours. Every unit you rent has a serial number, and the console shows its full history — including who used it before you, and for how long.
Yes, whenever you like. The hardware is your asset. When the managed contract ends we de-rack, ESD-pack and ship it out, charged as a one-off relocation service. We do not hold hardware back in any form, and there is no “minimum managed term” clause locking your asset inside our facility.
You can convert to a purchase at any point in the rental, with 60% of the rent already paid credited against the price, capped at 40% of that price. What converts is the machines you are already using — same serial numbers, no migration, no reinstall, no downtime, effective the same day.
HGX models start at eight GPUs (one whole node, because an NVLink domain cannot be split); L40S starts at one. There is no minimum term — billing is hourly and stops when the machine stops. There is no minimum monthly spend and no committed volume either.
Rental customers: we swap the unit, the swap window is not billed, and the time it took counts against our SLA. P1 faults get a response within 4 hours and a replacement within 24.
Outright customers: Kepter Care covers parts and labour on the same response times. Parts beyond warranty are billed at cost, with no markup.
No. We do not train models, we do not serve inference, we do not run any business that competes with a customer, and we do not use idle capacity for mining or for internal jobs. Idle means idle — which is what lets us publish per-cabinet utilisation.
All three facilities are inside the United States; hardware, backups and logs do not leave the country. Rented instances are wiped on return and a re-inspection report is issued. We do not look inside your instance: telemetry collects hardware-level metrics only — temperature, draw, memory in use, ECC — never memory contents or disk data.
Get started
Within 48 hours you get a plan naming specific serial numbers and arrival dates, with the arithmetic for both buying and renting. If it is not a fit we will say so — we do not chase quotes.