OS

Serve tokens from your own racks.

Token Factory is the inference serving stack in Podstack OS: an OpenAI-compatible API, model catalogue, keys, quotas and usage analytics, white-labeled and run in your data center.

Token Factory is the inference serving stack in Podstack OS. It gives your users an OpenAI-compatible API, a model catalogue, keys and quotas, and gives you usage analytics and observability, all under your brand and inside your data center.

Open Token Factory
At a glance

Token Factory at a glance

API
OpenAI-compatible
Branding
white-label
Licensing
one-time + AMC
Powers
Podstack Inference
What it does

What Token Factory does

OpenAI-compatible serving

Chat, completion and embedding endpoints that existing SDKs talk to after a base-URL change.

Model catalogue

Deploy open-weight models from the catalogue onto your GPUs, with whole-GPU and fractional placement.

Keys, quotas and per-token metering

Issue API keys to your users, cap spend per key and meter every token for billing.

Usage analytics and observability

Requests, tokens, latency and GPU health per model and per user, in one console.

White-label by configuration

Name, logo and accent colour are runtime configuration, so the console and API are yours.

Sovereign-ready

Runs entirely inside your facility, which is why in-country programs start with a Token Factory.

What you run and what Podstack runs

You run the facility. Podstack ships the software.

Podstack OS is installed inside your data center and operated by your team. Podstack maintains the software and, if you opt in, brings demand.

You run

  • The GPUs, servers, network and power, in your facility.
  • Your brand on the storefront, the console and the invoices.
  • Your prices, your customers and your customer relationships.
  • Day-to-day operations of the racks and the tenants on them.

Podstack runs

  • The operating system: provisioning, pooling, metering, storefront and console.
  • Updates, security patches and support for the stack.
  • The same software Podstack runs its own cloud on, so behaviour matches.
  • Optional demand through Podstack Grid, at a floor price you set.
How it works

How Token Factory works

  1. 1

    Install on your GPU nodes

    Token Factory deploys onto a Podstack OS fleet or onto a Kubernetes cluster you already run.

  2. 2

    Deploy models

    Pick models from the catalogue and place them on whole or fractional GPUs.

  3. 3

    Hand out keys

    Create keys for your users or teams, set quotas, and point their clients at your endpoint.

  4. 4

    Watch and bill

    Usage analytics show tokens per user and per model; export them to your billing.

How it bills

How Token Factory bills

Works with

Token Factory works with

FAQ

Questions about Token Factory

What is Token Factory?

Token Factory is the inference serving stack in Podstack OS, providing an OpenAI-compatible API, model catalogue, keys, quotas and analytics under your brand inside your data center.

Is it the same stack as Podstack Inference?

Yes, Podstack Inference on Podstack Cloud runs on Token Factory, so behaviour is identical in a partner facility.

How is Token Factory licensed?

It is a one-time license with an annual maintenance contract that covers updates and support.

Can I brand it?

Yes, the name, logo and accent colour are runtime configuration, and your users never see Podstack unless you choose to show it.

Does it need DC Suite?

No, Token Factory runs on a Podstack OS fleet or on a Kubernetes cluster you already operate.

What do I get for observability?

Per-model and per-user request counts, token usage, latency and GPU health, in the same console your team uses to deploy models.

Start with Token Factory

Token Factory is the inference serving stack in Podstack OS: an OpenAI-compatible API, model catalogue, keys, quotas and usage analytics, white-labeled and run in your data center.

Open Token Factory