You pay for the whole GPU. Your work uses a slice.
Podstack splits GPUs into shares, bills by the minute or token, and helps data centers sell their spare GPU capacity.
GPU compute is sold in the wrong shape.
Too big
Running a notebook shouldn't require renting an entire H100.
Too long
Billing in large blocks charges you for time you didn't use.
Always on
Inference endpoints keep billing overnight, even with zero traffic.
The GPUs are installed. The cloud layer is missing.
Operators have the hardware, but lack the software to sell it as a cloud and the customers to keep it busy.
- Provisioning and GPU slicing
- Self-service portal under your brand
- Metering and billing
- Customers who need the capacity
One platform that serves both sides.
Podstack owns no data centers. All capacity comes from partners running Podstack OS.
Compute that matches your workload.
Rent a slice, not the whole card.
One-click AI stacks on fractional NVIDIA GPUs.
Pay per unit of work.
No minimum commitment on pods. Rates appear before you launch.
No traffic, no bill.
One OpenAI-compatible API that scales to zero when idle.
One CLI, one SDK, one API.
Start with a GPU slice and scale up to reserved clusters and production endpoints.
One CLI, one SDK, one API
$ podstack gpu launch
? pick GPU › H100 80GB · 50% slice
✓ ssh ready · billed per GPU-minute
$ podstack train create --budget 25
→ LoRA job started · adapter saved when it finishes
$ podstack models list
llama-3-70b · mistral-7b · sdxl · +18Run the same cloud under your own brand.
| GPU | Shape | Price |
|---|---|---|
| H100 SXM | 50% slice | [YOUR PRICE] |
| A100 80GB | Bare metal | [YOUR PRICE] |
| L40S | VM | [YOUR PRICE] |
| Llama 3 70B | API | [YOUR PRICE] |
Need customers? Opt in to Grid.
Receive Podstack demand at a floor price you set. You can also skip Grid and still run the full OS.
Built so a region never goes dark.
- 01Podstack OS runs on every node, so each GPU's health is visible directly
- 02Each region has multiple operators, so workloads fail over
- 03Your uptime terms carry through to every operator
If one site fails, another operator in the region takes the workload.
Questions, answered.
What is Podstack?
Podstack is an AI cloud for developers plus the operating system that powers it: Podstack Cloud for building, training and serving, and Podstack OS for data center operators to run the same platform white-labeled on their own GPUs.
What are the three layers?
Podstack Cloud is what developers buy, Podstack OS is what operators run in their facility, and Podstack Grid is the optional layer that routes Cloud demand to operators who opt in.
What is a fractional GPU?
A fractional GPU is a share of one physical NVIDIA GPU, from 12.5 to 100 percent, isolated by PodVirt so several users can run on one card at once.
How does billing work?
QuickPods bill per GPU-minute, Inference bills per token with nothing charged when idle, TrainPods bill per hour or per reserved term, and Podstack OS is licensed per GPU per year.
Does Podstack own the data centers?
No, Podstack owns no GPUs and no data centers by design; capacity comes from partner operators running Podstack OS, and Podstack's own test racks run the same software.
Can I run Podstack in my own data center?
Yes, operators license Podstack OS, which includes DC Suite, NextGen DC Suite, PodVirt and Token Factory, and run it fully white-labeled inside their facility.
What does an operator keep on Grid?
On pods and clusters routed through Podstack Grid, the operator sets a floor price and keeps 65 to 80 percent of what the capacity earns; joining is optional.
Is Podstack certified?
Yes, the platform is ISO 27001 certified.