Serve tokens from your own racks.
Token Factory is the inference serving stack in Podstack OS: an OpenAI-compatible API, model catalogue, keys, quotas and usage analytics, white-labeled and run in your data center.
Token Factory is the inference serving stack in Podstack OS. It gives your users an OpenAI-compatible API, a model catalogue, keys and quotas, and gives you usage analytics and observability, all under your brand and inside your data center.
Token Factory at a glance
What Token Factory does
OpenAI-compatible serving
Chat, completion and embedding endpoints that existing SDKs talk to after a base-URL change.
Model catalogue
Deploy open-weight models from the catalogue onto your GPUs, with whole-GPU and fractional placement.
Keys, quotas and per-token metering
Issue API keys to your users, cap spend per key and meter every token for billing.
Usage analytics and observability
Requests, tokens, latency and GPU health per model and per user, in one console.
White-label by configuration
Name, logo and accent colour are runtime configuration, so the console and API are yours.
Sovereign-ready
Runs entirely inside your facility, which is why in-country programs start with a Token Factory.
You run the facility. Podstack ships the software.
Podstack OS is installed inside your data center and operated by your team. Podstack maintains the software and, if you opt in, brings demand.
You run
- The GPUs, servers, network and power, in your facility.
- Your brand on the storefront, the console and the invoices.
- Your prices, your customers and your customer relationships.
- Day-to-day operations of the racks and the tenants on them.
Podstack runs
- The operating system: provisioning, pooling, metering, storefront and console.
- Updates, security patches and support for the stack.
- The same software Podstack runs its own cloud on, so behaviour matches.
- Optional demand through Podstack Grid, at a floor price you set.
How Token Factory works
- 1
Install on your GPU nodes
Token Factory deploys onto a Podstack OS fleet or onto a Kubernetes cluster you already run.
- 2
Deploy models
Pick models from the catalogue and place them on whole or fractional GPUs.
- 3
Hand out keys
Create keys for your users or teams, set quotas, and point their clients at your endpoint.
- 4
Watch and bill
Usage analytics show tokens per user and per model; export them to your billing.
How Token Factory bills
one-time license plus annual maintenance
Token Factory is a one-time license with an annual maintenance contract for updates and support. Ask for a proposal.
Questions about Token Factory
What is Token Factory?
Token Factory is the inference serving stack in Podstack OS, providing an OpenAI-compatible API, model catalogue, keys, quotas and analytics under your brand inside your data center.
Is it the same stack as Podstack Inference?
Yes, Podstack Inference on Podstack Cloud runs on Token Factory, so behaviour is identical in a partner facility.
How is Token Factory licensed?
It is a one-time license with an annual maintenance contract that covers updates and support.
Can I brand it?
Yes, the name, logo and accent colour are runtime configuration, and your users never see Podstack unless you choose to show it.
Does it need DC Suite?
No, Token Factory runs on a Podstack OS fleet or on a Kubernetes cluster you already operate.
What do I get for observability?
Per-model and per-user request counts, token usage, latency and GPU health, in the same console your team uses to deploy models.
Start with Token Factory
Token Factory is the inference serving stack in Podstack OS: an OpenAI-compatible API, model catalogue, keys, quotas and usage analytics, white-labeled and run in your data center.