AIForge Works · Market Research
The Sticker Is Not the Product: Five Realities of GPU Catalog Pricing
Pricing evidence through the September 27, 2026 Harvest

On September 27, two eight-GPU H100 SXM machines sat in public catalogs at $3.90 and $12.29 per GPU-hour. Same chip, same form factor, same eight GPUs per machine, same day. The first is Crusoe in New York. The second is Azure in West US 3. The Azure instance costs more than three times the Crusoe one.
Neither price is a mistake and neither provider is doing anything wrong. They are selling different products that happen to share an accelerator.
That is the reality underneath every GPU budget. Somewhere in yours sits a number for what an H100 costs per hour. Before you defend it in a planning meeting, ask three questions. Which H100? Sold by whom? On what terms?
One note on timing, which makes this article’s argument for it. Nebius’s published pricing lists $4.50 per H100 GPU-hour in Finland effective October 1, 2026, up from $3.85, an increase of about 17 percent. The table shows $3.85 because that is what our September 27 Harvest observed. The next Harvest will check the updated catalog.
So if you are pricing an H100 at Nebius now, this is the moment to check the provider rather than any index, including ours. That is not a disclaimer. It is the whole point.
Why this is worth ten minutes
GPU pricing punishes rounding errors. One dollar per GPU-hour on a single eight-GPU machine running around the clock is about $70,000 a year, and most production estates are larger than one machine.
Run that arithmetic on the two machines above. The difference between the Crusoe and Azure examples is $8.39 per GPU-hour, so at continuous use one eight-GPU machine costs about $588,000 more a year on Azure than on Crusoe, before accounting for the platform and service differences that may justify part of it. That is the scale of decision hiding inside a number most budgets carry to two decimal places.
Every figure below comes from the AIForge Works September 27 catalog across nine providers, and every one is dated. None of it is a live quote, proof of capacity, or a guarantee that you can order anything. Verify current pricing, terms, fees, availability, and discounts with the provider before acting. The methodology, the exclusions, and the known qualifications are at the end, where you can check them.
Reality 1: The sticker is not the product
Two listings that both say "H100" can sell very different systems. The GPU may be an SXM accelerator attached to a high-speed interconnect or a PCIe card in a different topology. A listing may price a whole machine, a GPU-backed virtual machine, or a component. Memory capacity, GPU count, CPUs, host memory, storage, networking, operating system, and support all vary. Region and purchase model (e.g. on-demand versus a one-year term) move the price again.
This is why a per-GPU-hour figure is a normalization rather than proof that two things are the same. It makes prices comparable enough to reason about. It does not make the products identical. The table below holds every row at the same pack size, eight GPUs sold as one unit, so what remains different between the rows is the product and the provider selling it. The recorded region for each row, including the one Lambda does not publish, is in the Methodology note.
Reality 2: Who sells it sets the starting point
Each provider's lowest qualifying eight-GPU H100 catalog price in the September 27 snapshot, split into the two tiers that sell it: the hyperscalers, and the neoclouds, meaning the GPU-focused providers built around accelerators rather than a broad platform. This is a provider-floor comparison, not a claim that every row offers the same platform, service level, network, or available capacity.
| Provider | Type | Lowest qualifying 8-GPU option | $/GPU-hr |
|---|---|---|---|
| Nebius (Finland) | Neocloud | gpu-h100-sxm/8gpu-128vcpu-1600gb |
$3.85 |
| Crusoe | Neocloud | h100-80gb-sxm-ib.8x |
$3.90 |
| Lambda | Neocloud | gpu_8x_h100_sxm5 |
$3.99 |
| CoreWeave | Neocloud | hgx-h100 |
$6.16 |
| AWS * | Hyperscaler | p5.48xlarge |
$6.88 |
| OCI | Hyperscaler | B98415-us-ashburn-1 |
$10.00 |
| GCP | Hyperscaler | gcp-machine-a3-highgpu-8g |
$11.06 |
| Azure | Hyperscaler | Standard_ND96isr_H100_v5 |
$12.29 |
| Vultr | Neocloud | No confirmed comparable price |
Nebius: $3.85 is the September 27 observed price in eu-north1 (Finland). Its published pricing lists $4.50 effective October 1. * AWS: the saved source establishes H100 80 GB for the qualifying P5 offer but does not state the SXM form factor.
Every priced entry above is an eight-GPU machine. Across the nine providers we track, neocloud floors in this snapshot ran from $3.85 to $6.16, and hyperscaler floors ran from $6.88 to $12.29. The cheapest hyperscaler floor across those nine was about 1.8 times the cheapest neocloud floor, and the full spread across the eight qualifying provider minimums was about 3.2 times.
The simplest way to read that: the two bands do not overlap. None of the eight provider minimums shown falls between the most expensive neocloud floor at $6.16 and the cheapest hyperscaler floor at $6.88.
That is not a value ranking. A hyperscaler rate may carry identity integration, managed services, enterprise agreements, compliance controls, and a global footprint you actually need. The point is narrower and more useful: provider tier sets your starting price before any discount conversation begins. If you negotiate hard against a number that started three times higher, you can win the negotiation and still lose the decision.
Marketplace and community capacity is excluded because hardware bundles, reliability, and listing persistence vary enough that blending them in would turn this provider-catalog comparison into a different analysis.
Reality 3: A commitment is a bet on your own utilization
A committed discount is not free money. It trades flexibility for a lower rate, and you carry the risk of using less than you planned.
For a like-for-like machine billed across the full term, the break-even test is one line:
Break-even utilization = effective committed hourly rate ÷ on-demand hourly rate
Effective means the committed rate with any upfront payment amortized across the term, so the numerator is what the hour actually costs you.
If on-demand is $10 and the committed rate is $6, the commitment only wins once you use more than about 60 percent of the committed hours. Below that, paying $10 for the hours you actually use beats paying $6 for every hour in the term. The ratio ignores reservation scope, resale rights, burst needs, and operational cost, but it exposes the utilization assumption that a discount headline can obscure.
Two cautions. AIForge Works exposes comparable term-specific public rates for AWS, Azure, and GCP, where those providers publish them in structured form. Other providers may offer commitment discounts privately without publishing machine-by-machine terms, so a missing row is not evidence that no discount exists. And a commitment optimizes within one provider and one product. It does not erase a cross-tier gap. Start with the workload, then compare purchase models inside the set that can actually run it.
Reality 4: The hourly rate is not the invoice
Four categories routinely change the answer after the GPU rate is settled.
Data transfer. Moving training data, checkpoints, and inference output can carry separate charges, and the policy differs by provider and by direction.
Persistent storage. Volumes and snapshots keep billing after the GPU stops.
Interruption and recovery. A discounted machine that vanishes mid-run costs repeated compute, lost time, and engineering hours.
Software and platform charges. Operating-system licenses and other software charges can add to the GPU rate. Compare the software included in each offer as well as the hardware.
The right comparison is not the smallest hourly number. It is the cost of completing the required work at the required service level.
Reality 5: Match the buying model to the workload, then pick the tier
Interruption-tolerant batch work, checkpointed training, and experiments can often use spot, preemptible, or marketplace capacity. Production services carrying latency or availability commitments usually need on-demand or reserved capacity and stronger operational guarantees. A long-running but elastic workload may need a mix rather than a single buying model.
Five checks before your next GPU purchase or renewal:
- Define the product. GPU model, memory, form factor, interconnect, pack size, host bundle, software, region, terms.
- Normalize the comparison. Whole-machine price, GPU count, per-GPU price, and a dated evidence basis.
- Test the commitment. Committed rate divided by on-demand rate, against honest utilization and your real need for flexibility.
- Price the result, not the hour. Transfer, storage, interruption, software, labor, time to completion.
- Verify the boundary. Current quote, orderability, capacity, minimums, fees, taxes, and service commitments, with the provider.
Those five are for the decision in front of you. If the decision is already made and what is coming is the renewal, the companion piece is 7 Questions to Answer Before Making or Renewing a GPU Decision, which runs the same discipline backward: what you decided, what evidence supported it, and how old that evidence is now.
Why the gaps persist
The products stay different. Hyperscalers price a broad platform and regional footprint. Neoclouds run narrower catalogs closer to the hardware and compete directly on accelerator economics. Supply, networking, support, contract structure, and regional demand all differ, and large buyers negotiate away from list while published catalogs move on their own schedule.
There is a subtler point in the September 27 evidence. Exact matched on-demand prices were stable, yet the market still changed underneath them: Lambda held all 22 of its GPU configurations at the same prices while its availability feed reported fewer locations. Stable price does not mean stable availability, and a newly visible price line does not prove new capacity. Catalogs answer what something costs, not whether you can buy it.
Where AIForge Works fits
We turn provider-published catalogs and specifications into dated, comparable decision evidence across AWS, Azure, GCP, OCI, CoreWeave, Lambda, Vultr, Nebius, and Crusoe. The Multi-Cloud Pricing Analyzer is where you inspect public options by GPU, region, and purchase model. GPU Audit pressure-tests a current or proposed rate and its operating assumptions against dated public-market context.
For teams with one active decision, the GPU Decision Review is a complimentary, private, analyst-led path offered to a limited number of accepted applicants. It includes a short workload confirmation, a 30-minute Review, and a written Decision Brief covering comparable public-market context, evidence-supported alternatives, assumptions, limitations, the questions worth putting to your provider, and a next action. No invoice upload and no cloud credentials.
Methodology. This article uses the September 27, 2026 catalog evidence published in the Weekly Pricing Pulse prepared September 28. All nine configured provider collections completed successfully. The snapshot contains 6,318,998 public pricing rows and 13,999 normalized public on-demand GPU catalog observations. Of the 13,991 observations matched exactly on provider, region, and GPU configuration against the prior snapshot, none moved by more than 1 percent.
The H100 table shows each provider's lowest qualifying saved-catalog price in the H100 SXM cohort, normalized to dollars per GPU-hour by dividing the listed instance or pack price by GPU count. It excludes software-only charges, component-only prices, unconfirmed GPU specifications, fractional vGPU slices, zero-priced listings, private quotes, negotiated discounts, taxes, provider-specific fees, spot and preemptible prices, and reserved or committed terms. Marketplace and community capacity is excluded as not comparable to governed provider catalogs.
Three qualifications are retained rather than inferred away. The saved AWS source establishes H100 80 GB for the qualifying P5 offer but does not literally state the SXM form factor. Lambda did not publish a region for its floor. No confirmed comparable Vultr price was found in this snapshot.
Recorded region for each table row: Nebius eu-north1, Crusoe US East (New York), Lambda not published, CoreWeave global list price, AWS US East (N. Virginia), OCI us-ashburn-1, GCP us-west1, Azure westus3. Region is a qualification on the comparison rather than a variable the table controls for, which is why it is recorded here rather than beside every price.
The Nebius row was re-verified for this version against gpu-h100-sxm/8gpu-128vcpu-1600gb at $30.80 per machine-hour, which is $3.85 per GPU-hour. Earlier versions carried the one-GPU configuration at the same per-GPU rate. Using the eight-GPU configuration holds pack size constant across every row in the table.
Counts need labels as much as prices do. Five newly visible AWS lines in this Harvest were one B200 instance family in one market, the eight-GPU p6-b200.48xlarge in Paris at $174.32 per instance-hour, or $21.79 per GPU-hour on Linux, plus four operating-system variants from $21.81 to $22.27 per GPU-hour. That is one product with five price lines, not five product launches, and reporting it the other way would overstate what changed.
Catalog presence is not proof of capacity, orderability, or a live quote.
Sources. AIForge Works Weekly Pricing Pulse, September 28, 2026 · AIForge Works products and evidence boundary · Nebius AI Cloud pricing and Nebius compute pricing documentation, accessed September 30 and re-checked October 1, 2026 · GPU Decision Review
