Pricing Cost Audit · $500 · 2 weeks

Your real cost per million tokens, measured on your stack.

For teams that self-host open models on their own or rented GPUs. In two weeks I set up Throttle on your serving stack, measure what you actually pay per million tokens, test the changes you're considering, and leave a CI gate behind so a costlier deploy gets caught before production.

Price$500 one-time
Length2 weeks
Works withvLLM · SGLang · Ollama · TGI
Book a 15-minute call → or email kushthrottle@gmail.com

What you get

  1. Your true cost per million tokens, measured on your own stack with 95% confidence intervals, not estimated from a spec sheet.
  2. Calibrated verdicts on up to three config changes you choose: a vLLM flag, a quantization, an engine version, a different GPU. Each one comes back CHEAPER, MORE EXPENSIVE, or NO WINNER when the difference is inside your measured noise.
  3. A CI cost gate wired into your pipeline with the Throttle GitHub Action, so a change that makes every token more expensive fails before it ships.
  4. A written report in dollars, with the raw results attached so your team can check every number.

There is no savings guarantee. What you are buying is a measurement you can trust. Sometimes the honest answer is that a change your team was excited about is noise, and that is worth knowing before you roll it out.

Timeline

Week 1

Baseline and noise

Kickoff call, access set up, then repeat checks of your current config to measure its cost per million tokens and how much it swings run to run. That noise bound is what every later verdict is judged against.

Week 2

Changes, gate, report

Each change you picked is measured against the baseline and gets a verdict. The CI gate goes into your pipeline, and you get the report plus a 30-minute walkthrough.

What I need from you

  • Access to an OpenAI-compatible endpoint you serve (vLLM, SGLang, Ollama, TGI or similar), or time on a box where I can run Throttle with your team.
  • What your GPUs cost per hour, so tokens and time can be turned into dollars.
  • A representative prompt sample or a description of your traffic shape, so the measurement reflects your real workload.
  • The changes you want tested, and where your CI runs.

Confidentiality

Throttle runs on your side and only sends requests to your own endpoint. Prompts, responses and keys stay in your environment, and shared summaries leave out URLs, hostnames and anything that looks like a secret. I'm happy to work under your NDA.

What a result looks like

On a rented A100 PCIe 80GB running vLLM 0.16.0 with Qwen2.5-0.5B-Instruct at $1.39/hr, changing max_num_seqs from 1 to 8 took cost from $0.746 to $0.234 per million output tokens (−68.6%), measured with a six-position counterbalanced protocol (B-C-B / C-B-C) so warm-up and drift could not fake the result.

The baseline was deliberately bad: one sequence at a time on a small model. It shows the measurement, not a saving you should expect. The raw artifacts are in the repo, and the full write-up is on the blog.

Questions

How is this different from Throttle Pro at $19/month?

Throttle itself is open source, and Pro is the self-serve product for teams that want scheduled checks and alerts. The audit is the hands-on version: I do the setup, the measurement and the write-up with your team, on your stack.

We already benchmark our serving stack.

Most benchmarks report throughput from one run. The audit reports dollars per million tokens with confidence intervals, measures your run-to-run noise first, and only calls a change a win when it clears that noise. It often confirms what you have; sometimes it shows a "win" was noise.

What if the result is inconclusive?

Then the report says so, with the numbers. An inconclusive verdict means the change is smaller than your noise, which is a useful answer before you spend more GPU hours on it.

Do you need access to production?

No. A staging endpoint or a dedicated box with the same model, engine and GPU type is enough, and it keeps real traffic out of the measurement.