Skip to content

Enterprise AI compute · three U.S. regions

Buy it outright,
or rent it by the hour.
Racked in three days.

We supply enterprise GPU nodes, accelerators and whole racks, two ways to take them: buy them and we manage them for you in a U.S. facility, or pay by the GPU-hour and stop whenever you like. Same hardware, same facility, same serial numbers in the same system.

112 nodes in stock, schedulable now8-GPU minimumNo commitment threshold
Live inventoryUpdated 2 minutes ago
KPT-B200-SXM6-192G
SJC-02 · Santa Clara, CA
Ready
24nodes
KPT-H200-SXM5-141G
IAD-01 · Ashburn, VA
Ready
64nodes
KPT-H200-SXM5-141G
DFW-01 · Dallas, TX
Racking
14nodes
KPT-L40S-PCIE-48G
IAD-01 · Ashburn, VA
Ready
46cards
Stock and lead times are public. See the full catalogue →
148nodes
Under management

Across three U.S. regions, every one serial-tracked.

72h
Order to racked

Median for in-stock SKUs, including 72 hours of burn-in.

99.9%
Power & cooling

Dual-fed PDUs, N+1 cooling, written into the contract.

14days
Return window

Grace period after a rental ends; beyond it, the day rate applies.

Own or rent

Same machines,
two ways to take them.

In this industry those are usually two companies doing two things: one sells hardware, the other rents compute. We do both, out of one inventory, one facility and one serial-number system. So you do not have to guess between “what if it sits idle” and “what if renting gets expensive” — both cost curves are plotted below.

Figure 01How the two paths runFig. 01 · Ownership vs. hourly
Your workloadFor how long?OWN · KEPTER SUPPLYSpec & quotePlan back within 48 hPurchase & handoverTitle transferred to youWe manage itMonthly · parts and labourThe asset is yoursDepreciation and residual tooMovable at any timeRENT · KEPTER FLOWSpec & provisionAvailable the same dayBilled by GPU-hour8-GPU minimumScale any timeNo term, no exit feeStop when you are done14-day return windowNo asset, no balance owedWhat both paths share · SHAREDSame hardware · same facility · same serial tracking · same 72-hour burn-in report · same console
How to read: the upper and lower paths differ only in the commercial relationship — everything inside the dashed band in the middle is identical. That is why we can sell and rent at the same time: there is one ops team, one tracking system, one facility and one burn-in standard, so “the rental units are the worse batch” cannot be true.

Kepter Supply

Buy outright + managed

You put up the capital and keep the asset. We put up the facility, the power, the hands and the spares.

  • The asset is on your booksInvoices, serial numbers, depreciation and residual value are yours — deductible, financeable, resellable.
  • Unit cost thins out over timeRun for 36 months, total cost lands at about 59% of renting the same configuration by the hour (the arithmetic is in figure 06 below).
  • Capacity is certainThe machines are yours. They are not scheduled into a shared pool, so they never queue behind someone else at peak.
  • You can move them outWhen the managed contract ends we de-rack, pack and ship them out. We do not hold hardware back.
  • The managed fee is fixedBilled per rack and per kW, written into the contract, not settled against a floating power price.
Fits: steady workloads, an 18-month-plus horizon, teams that need the asset on the balance sheet.Ask for a purchase plan

Kepter Flow

Rent by the GPU-hour

Operating expense only. Pay for the hours you use; stopped machines do not bill.

  • Available todayFor in-stock SKUs, connection details land within hours of provisioning.
  • No long contractHourly settlement, no minimum term, no early-exit penalty.
  • Scale without paperworkEight GPUs to start; adding more is confirmed live against remaining stock.
  • Hardware risk is not yoursWe swap failed units. The replacement time counts against our SLA, not against your bill.
  • Convertible to ownershipAt any point in a rental you can convert the machines you are on to a purchase, with 60% of rent already paid credited against the price.
Fits: training peaks, pilots, workloads that have not settled into a shape yet.See hourly rates

Not sure which? Rent first, decide after three months.

This is not a sales line: at any point in a rental you can convert the machines you are using to a purchase, with 60% of the rent you have paid credited against the price. Same serial numbers — no migration, no reinstall, no downtime. Three months of real utilisation figures beat any procurement model.

Deployment

Order to training job:
a 72-hour median.

That figure includes 72 hours of burn-in — the machine has already run at full load for three days before it reaches you, and the report ships with it. We publish the duration of every leg, because a date is the only promise worth anything in this business.

Figure 0272-hour racking timelineFig. 02 · T+0 → T+72h
Commercial8 hLogistics28 hFacility36 hT+0Order confirmedContract / POT+2hStock lockedSerials assignedT+8hPicked & packedTagged · ESD packedT+36hRacked on sitePower · network · bayT+60hBurn-in passed72 h full load · filedT+72hDeliveredCredentials · in consoleWhat you get at each step · WHAT YOU GET① Quote & spec sheet → ② Serial number list → ③ Waybill & live location④ Rack position & topology → ⑤ Burn-in report PDF → ⑥ SSH / API credentials
How to read: the three bands along the top are who owns which stretch — commercial 8 h, logistics 28 h, facility 36 h. Whichever leg overruns is the leg at fault; SLA penalties are assessed per band. The row underneath is what you should have received at each milestone. If it has not arrived, that step is not finished — do not accept “in progress”.

01 · Commercial

Select and lock

Tell us the GPU model, the count, the region and roughly how long for; a plan comes back within 48 hours. Two hours after you confirm, stock is locked and serials are assigned on the spot.

You get: quote · serial list

02 · Logistics

Pick and ship

Asset-tagged, ESD-packed, re-checked before racking. The waybill is trackable throughout, and cross-region transfers move on our own fleet.

You get: waybill · live location

03 · Facility

Rack and cable

Assigned a cabinet, wired to both power feeds and to 400G. You will be told which cabinet, which U positions and which switch this batch hangs off.

You get: rack position · topology

04 · Acceptance

Burn in and deliver

Seventy-two hours at full load with temperature, draw, memory ECC and NCCL bandwidth logged throughout. Passing is what gets it delivered; failing gets it swapped and re-run.

You get: burn-in report · access credentials

Inside the rack

What you buy — or rent —
looks exactly like this.

Most compute suppliers give you a GPU model and a number. We have drawn a whole cabinet apart: which U positions your machines sit in, which switch they hang off, which power feed they draw, where the tag goes. All of it lands in your topology document on delivery day.

Figure 0342U standard cabinet, dissectedFig. 03 · Rack anatomy · 10.4 kW
24 °C intake≤40 °C return42U · W600 × D1200ABToR-A · 400GToR-B · 400GHGX NODE 01 · 8USN KC 7F20-B4418HGX NODE 02 · 8USN KC 7F20-B4419HGX NODE 03 · 8USN KC 7F20-B4420HGX NODE 04 · 8USN KC 7F20-B44216U reserved · expansionNetwork · 2 × 400G ToR switchesEvery node dual-homes to switches A and B. Either can fail without dropping the link.Tracking · asset tag position (orange block)Lower right of every faceplate: serial, SKU, intake date, QR code. It opens the full history.Compute · 4 × HGX nodes = 32 GPUsEight GPUs per node, 8U high. A full cabinet is 32 GPUs, our minimum whole-cabinet unit.Fabric · NVLink plus 400G, two tiersNVLink meshes 8 GPUs at 900 GB/s inside a node; 400G between nodes; 3.2 Tb/s out of the cabinet.Cooling · contained cold aisle, front to backIntake at 24 °C, return no higher than 40 °C, N+1 redundant. Curves are readable per cabinet.Power · dual PDUs (the vertical strips)10.4 kW per cabinet, A and B fed independently, so losing one feed does not drop the load.
How to read: on the left is the cabinet in front elevation. The two vertical strips are the A and B power feeds, the pair along the top are the dual-homed switches, the four blocks in the middle are 8U HGX nodes, and the small orange square at the lower right of each one is where the asset tag goes. The dashed box at the bottom is 6U left deliberately empty — we do not fill a cabinet, so expanding it never means moving machines that are already in it. 6 callouts in all.

Hardware

Models on sale and for rent.

Four main product lines, outright price and hourly rate written side by side. The price is not hidden behind “contact sales” — what needs discussing is terms and volume, not the unit price.

KPT-B200-SXM6-192G

B200 HGX

Newly rackedSJC-02 in stock
Memory
192 GB HBM3e
Node
8 GPUs / 8U
In-node fabric
NVLink 1.8 TB/s
Draw
~14.3 kW
Minimum
8 GPUs
$8.90 / GPU-hr
Outright $392,000 / 8-GPU node

KPT-H200-SXM5-141G

H200 HGX

In stock, all regions
Memory
141 GB HBM3e
Node
8 GPUs / 8U
In-node fabric
NVLink 900 GB/s
Draw
~10.4 kW
Minimum
8 GPUs
$3.42 / GPU-hr
Outright $248,000 / 8-GPU node

KPT-H100-SXM5-80G

H100 HGX

IAD-01 in stock
Memory
80 GB HBM3
Node
8 GPUs / 8U
In-node fabric
NVLink 900 GB/s
Draw
~10.2 kW
Minimum
8 GPUs
$2.28 / GPU-hr
Outright $164,000 / 8-GPU node

KPT-L40S-PCIE-48G

L40S PCIe

IAD-01 in stockSingle GPU
Memory
48 GB GDDR6
Node
1–8 GPUs / 4U
In-node fabric
PCIe Gen5
Draw
~3.1 kW
Minimum
1 GPU
$1.15 / GPU-hr
Outright $11,400 / per GPU

Rates are on-demand and exclude commitment discounts. Rent continuously for 3 months and the rate drops 20% automatically; at 12 months, 35%. No contract to sign, no application to file — it shows up on the bill. The outright price includes the first year of managed ops; from year two the managed fee is billed per rack and per kW. All prices are illustrative; the quote governs.

U.S. footprint

Three regions,
one private backbone.

All hardware is inside the United States. Between regions we run dark fibre we lease ourselves, so cross-region traffic never touches the public internet and is never billed as egress — which is a real line item when you are training data-parallel across regions, not a decoration on a spec sheet.

Figure 04Region layout and inter-region latencyFig. 04 · SJC · DFW · IAD
SJC ↔ IAD · 58 ms33 ms27 msSJC-02Santa Clara, CA · WestRacks 18 / 24 used576 GPU · 250 kWDFW-01Dallas, TX · CentralRacks 9 / 20 used288 GPU · 208 kWIAD-01Ashburn, VA · EastRacks 26 / 32 used832 GPU · 333 kWWestCentralEastOwn dark fibreRelayed via Central
How to read: the bar on each of the three cards is current rack occupancy. A green dot means there is headroom, amber means close to full (DFW-01 is amber because it is being expanded, not because it is out of stock). Solid lines are our own point-to-point dark fibre; the dashed line marks SJC↔IAD relaying through the central site — which is why 58 ms is a little more than the two legs added together. That gap is real and we do not round it away.

Data stays in country

All three regions are inside the United States; hardware, backups and logs do not leave it. We will supply the facility locations and a device list on request.

Power is in the contract

Dual-fed PDUs, N+1 cooling, and 99.9% power and cooling availability written into the SLA. Miss it and we credit a share of that month’s managed fee.

Come and look

Outright customers can book a quarterly site visit and see their own cabinet and their own tags. It is the plainest way to verify that the asset is in your name.

Serial tracking

Every machine carries
a complete history.

The serial is assigned the moment a unit arrives and does not change until it is retired. Whose hands it passed through, what it ran, how burn-in went, how many repairs, which U of which cabinet it is in right now — scan the QR code on the tag and it is all there.

Figure 05Serial-number lifecycle loopFig. 05 · One SN, cradle to grave
Records attached to this machine · ATTACHED TO SNSN KC 7F20-B4418Purchase invoiceBurn-in reportRacking recordTemperature curvesBilling detailRepairs & swaps01Intake checkUnboxed · photographed0272 h burn-inFull load · logged03Asset taggedSerial bound to QR04Racked & stockedCabinet and U logged05In customer useTelemetry · billing06Return / RMARe-inspected · wipedRe-stocked once it passes inspection · same SN throughout, history accrues
How to read: the strip along the top is the six kinds of record hanging off one serial number, and they do not reset because the machine changed tenants. The purple dashed line at the bottom is the recovery loop: a machine that comes back from a customer is re-inspected, wiped and re-stocked, and its history keeps accumulating. So “this unit has 412 days of service and one fan replacement”, as the console puts it, is true — not counted from the day you rented it. 6 stages around the loop.

What outright customers get

  • A per-device invoice cross-referenced to serial numbers
  • The 72-hour burn-in report as a PDF, with NCCL bandwidth and ECC logs
  • Rack position and topology, refreshed quarterly
  • The right to book a site visit each quarter
  • De-racking, packing and outbound shipping when the managed contract ends

What rental customers get

  • The serial numbers of the units you are on (not “an H200 somewhere”)
  • The same burn-in report — rentals are not a different batch
  • Hour-by-hour billing detail, reconcilable to a single device
  • A record of every swap and how long it took, counted against the SLA
  • A return inspection report confirming the data was wiped

Pricing

When to buy,
and when to rent.

There is only one honest answer: it depends how long you will run it. The chart below runs the real price list for one 8-GPU H200 node out to 36 months; the curves cross in month 16. Renting wins below 16 months, owning wins above it — we do both, so we have no side to argue for.

Figure 0636-month cumulative cost: outright vs. hourlyFig. 06 · 1× KPT-H200 8-GPU node
$0$100k$200k$300k$400k$500k$600k061218243036Months · MONTHSBreaks even in month 16Owning saves from here onSaves $209.7k over 36 monthsOwning is 59% of renting$248k one-offOutright + managed (first year included)Rented by the GPU-hour (automatic tiers included)What owning saves
How to read: the purple line steps vertically at month 0 — that is the $248k one-off capital outlay — then runs flat for 12 months (the first year of managed ops is inside the purchase price), and from month 13 adds $2.4k/month of managed fee. The amber line is renting by the hour; its slope eases twice as the commitment tiers land. The two lines meet in month 16 — that month moves with the GPU model and the utilisation, and the quote recomputes it against your actual numbers.

Hourly rates

SKUOn demand≥ 3 months≥ 12 monthsMinimum
B200-SXM6-192G$8.90$7.12$5.798 GPUs
H200-SXM5-141G$3.42$2.74$2.228 GPUs
H100-SXM5-80G$2.28$1.82$1.488 GPUs
L40S-PCIE-48G$1.15$0.92$0.751 GPU

Figures are USD per GPU-hour. Discounts apply automatically, accrued by continuous usage, with nothing to sign or request. Stopped machines do not bill, but storage kept for a retained instance is charged at $0.04 / GB-month.

What we do not charge for

ItemUsCommon in the industry
Inter-region traffic$0$0.02–0.09 / GB
Egress (first 50 TB/month)$0Metered
Starting and stopping instances$0Minimum billing increment
Early exit$030–100% of the remaining term
Downtime during a swapNot billedBilled as normal
Technical supportIncludedCharged by tier

The “common in the industry” column reflects general practice in published price lists and is not aimed at any one vendor. We have folded these into the base rate — the unit price does not look like the cheapest, but there is no second line on the bill.

Console

Bought and rented,
in one console.

Whether a machine is your asset or a rental, it looks identical in the console: utilisation, temperature, draw, billing, serial number, rack position. The only difference is whether the tag in the corner reads Supply or Flow.

KEPTERCOMPUTE/ fleetSupply + Flow
78.4%Queue utilisation · all regions
112Ready
14Racking
19Term ending
3Fault
214,880Billed GPU-hours this month
$734kThis month’s invoice
3Regions
IAD-01832 GPU · 81% usedSJC-02576 GPU · 74% usedDFW-01288 GPU · 62% usedSLA99.94% MTD

Per-device telemetry

Utilisation, memory in use, temperature, draw and ECC counts, sampled every 15 seconds and kept for 13 months.

Bills that reconcile to hardware

Every line on the invoice opens up to show which serial, which window and which rate produced it.

Complete API

Anything the console does, the REST API does: provision, scale, query stock, pull telemetry, export invoices.

Term warnings

Three reminders before a term ends — 14 days, 7 days, 48 hours — with automatic renewal or release available.

FAQ

Frequently asked

If the answer is not here, just ask. Our sales team does not work from a script: what they can answer they answer on the spot, and what they cannot, they tell you when they can.

Ask something specific
Are rented machines the same batch as the ones you sell?

Yes. Same inventory pool, same 72-hour burn-in standard, same facility. The only difference is the commercial relationship: bought units are transferred into your name and taken out of the shared pool, rented units remain ours. Every unit you rent has a serial number, and the console shows its full history — including who used it before you, and for how long.

Once I have bought them, can I take the machines away?

Yes, whenever you like. The hardware is your asset. When the managed contract ends we de-rack, ESD-pack and ship it out, charged as a one-off relocation service. We do not hold hardware back in any form, and there is no “minimum managed term” clause locking your asset inside our facility.

If I rent and then want to buy, what happens to the rent I paid?

You can convert to a purchase at any point in the rental, with 60% of the rent already paid credited against the price, capped at 40% of that price. What converts is the machines you are already using — same serial numbers, no migration, no reinstall, no downtime, effective the same day.

What is the minimum, and is there a minimum term?

HGX models start at eight GPUs (one whole node, because an NVLink domain cannot be split); L40S starts at one. There is no minimum term — billing is hourly and stops when the machine stops. There is no minimum monthly spend and no committed volume either.

What if a machine fails? Whose time is the downtime?

Rental customers: we swap the unit, the swap window is not billed, and the time it took counts against our SLA. P1 faults get a response within 4 hours and a replacement within 24.

Outright customers: Kepter Care covers parts and labour on the same response times. Parts beyond warranty are billed at cost, with no markup.

Do you run your own workloads on these GPUs?

No. We do not train models, we do not serve inference, we do not run any business that competes with a customer, and we do not use idle capacity for mining or for internal jobs. Idle means idle — which is what lets us publish per-cabinet utilisation.

What about data security and compliance?

All three facilities are inside the United States; hardware, backups and logs do not leave the country. Rented instances are wiped on return and a re-inspection report is issued. We do not look inside your instance: telemetry collects hardware-level metrics only — temperature, draw, memory in use, ECC — never memory contents or disk data.

Get started

Tell us the GPU model, the count, and roughly how long for.

Within 48 hours you get a plan naming specific serial numbers and arrival dates, with the arithmetic for both buying and renting. If it is not a fit we will say so — we do not chase quotes.