Days 1 and 2. Usage review
Which models you use, for which tasks, with how many tokens and at what cost. The starting point is your invoices, the provider dashboard and your application logs.
LLM cost audit
I look at which models your product calls, for which tasks, and what each one costs. Then I install a router in your code that sends simple requests to cheaper models and keeps the expensive ones for the hard cases.
Who it is for
Before and after
The real split comes from your data, not from this drawing.
Before
After
The week
Five working days from the moment access is in place. Every step leaves something you can review.
Which models you use, for which tasks, with how many tokens and at what cost. The starting point is your invoices, the provider dashboard and your application logs.
Which tasks can move to a cheaper model, what saving your data supports and what risk each change carries. If the saving is not worth it, you find out here, before anything is installed.
The router goes into a branch of your repository. Tests compare the cheap model's answers with the expensive model's answers on your real cases.
How to configure it, how to measure the saving and how to switch it off with a configuration change.
Spend per model and per task, with the source data and the maths in plain view.
Configurable per task and switchable off with a configuration change.
A task only moves to the cheap model if it passes the comparison on your cases.
Cost per request and per model, so the saving can be checked after delivery.
So your team can maintain and tune the router without depending on me.
The decision layer
If choosing the model costs almost as much as the answer, the saving disappears. Several models are now built for this job. The best known is Jev, from TypeSafe AI, which takes 70 to 500 ms to return a typed decision, according to its maker.
If you would rather not add another provider, a small model from the provider you already use, or a set of rules, can make the decision. We choose together and it goes in writing.
Price
The price depends on the number of flows, providers and environments. It is agreed in writing together with scope and timeline, before anything is touched.
Fifteen minutes to see which provider you use, what you spend and whether a router makes sense for you.
Scope, timeline and a fixed price. No open-ended hours and no costs that appear later.
The review in the first days shows it before anything is installed, and we decide in writing whether to continue, reduce the scope or stop.
Questions
Request
With these details I can tell you, before we talk, whether a router makes sense for your product. I reply personally.
Next step
If you received an email from me, a reply is enough. Otherwise, write to this address with the provider you use and what you pay per month, even a rough figure. If you would rather talk it through, ask for a 15-minute call and I will suggest times.