Surplus Compute

Real-time AI is
breaking your budget.

For every job beyond need-it-now.
Choose your timeline. Save on tokens.

POST /v1/batches illustrative
{
  "model": "glm-5.2",
  "completion_window": "6h",
  "messages": [
    {
      "role": "user",
      "content": "Summarize these 40,000 support tickets."
    }
  ]
}

A drop-in API for leading open models. One new field: your deadline.

Time turns into efficiency.

Your timeline gives us room to optimize. Whether it's tagging a dataset, summarizing tickets, or clearing an eval queue, we can help you save money. Surplus hosts models to maximize high-quality throughput and passes the savings on. Same models, same performance. Pay for work, not gaps.

Run now~55% utilized
Immediate execution leaves gaps
Given time~98% utilized
Time lets the same work pack tight

Efficiency means savings.

The more time you give, the less you pay. Just say when. e.g. Escalation triage that used to run in seconds can run in hours — same routing, a fraction of the cost.

today's market surplus
Choose your deadline, earn your savings. More time = more savings. *savings may vary by model price now 6h 12h 18h 24h on-demand batch · 24h surplus

Join the beta.

We're onboarding a small group of teams & developers now.