Control AI usage at runtime.
Put Vibe in front of your AI calls. Define the rules in the Console, not in your code, and decide what each customer may do before you pay the provider.
await usage.chat({
identity: 'user_123',
workflow: 'ai-chat',
model: 'gpt-5',
messages,
});
usage 184 tokens · Free plan
decision ROUTE gpt-5 → gpt-4.1-nano
settled 637 tokensThe idea in four steps
What you can do
SDK
Install, create the client once, then make policy-checked calls. Provider keys stay on your server.
npm install @usageflow/vibeconst result = await usage.chat({
identity: 'cust_acme', // who this usage belongs to
workflow: 'support-agent', // which policy (workflow) governs it
provider: 'anthropic',
model: 'claude-sonnet-5',
messages: [{ role: 'user', content: 'Summarize this ticket.' }],
customerMetadata: { plan: 'pro' }, // optional: durable facts about this customer
});
console.log(result.content);Choose your path
Control AI usage at runtime
Vibe checks every AI request before it reaches the provider and applies the policy for that customer. Rules live in the Console, not in your code, so you can change what each customer is allowed to do without a deploy.
1. Without Vibe, you write the rules yourself
- User
- Your app
- AI provider
Nothing stops one customer from spending the whole budget, and every plan rule (“free users get the cheap model after 100 tokens”) ends up as if statements in your app.
2. With Vibe, the rules sit in front of the provider
- Your app
- Vibeidentity + usage + policy
- Decisionallow · route · degrade · block
- AI provideronly if allowed
- Settlementactual usage recorded
A real policy, end to end
Start with the outcome. This is a policy for a free customer, read top to bottom as their usage grows. It is an example; you set your own thresholds and models in the Console.
Free customer · workflow ai-chat
Paid customer · same workflow
The same request, at 184 tokens of usage, for the free customer:
- Identity
- user_123
- Usage
- 184 tokens
- Policy
- Free plan · ai-chat
- Decision
- Route
- Requested → ran
- GPT-5 → GPT-4.1 nano
Same code, different policy
This is the entire integration. The call is identical for the free and the paid customer; only the policy behind ai-chat differs.
// Your app: no plan checks, no model switching, no limit math.
const result = await usage.chat({
identity: 'user_123', // who the usage belongs to
workflow: 'ai-chat', // which policy applies
provider: 'openai',
model: 'gpt-5',
messages,
});- Application code
- Vibe
- Policy from the Console
| Action | Provider called? | Result |
|---|---|---|
| Allow | Yes | Requested model |
| Route | Yes | A different model, possibly another provider |
| Degrade | Yes | Requested model; you apply the lower-cost behavior |
| Block | No | Request rejected |
How Vibe works
Before the call
Vibe knows who is calling (identity), which workflow, how much they have used, their metadata, the requested model and your thresholds. It then decides: allow, route, degrade or block.
After the call
Vibe does not just count requests. It reserves a worst-case amount first, so a request can never push a customer past their allowance, and then settles what the provider actually used.
Settlement never charges more than was reserved. Concepts that make this work, in the order you meet them: Identity (whose usage), Ledger (how much), Policy (what is allowed), then tiers and settlement.
Where your data goes
Your server
Your appVibe SDKProvider keysUsageFlow
Policy + usage decisionAI provider
OpenAI · Anthropic- Provider credentials stay in your environment.
- Vibe asks UsageFlow for the decision and reports usage numbers for metering and settlement.
- Prompts and responses go from your server straight to the provider, not through UsageFlow.
What is UsageFlow Vibe?
UsageFlow Vibe sits between your app and the AI providers. You make your AI calls through one SDK. Before each call UsageFlow checks the customer's usage against your policies, and after the call it records what was really used. You decide what happens at each usage level: carry on, route to a cheaper model, degrade the request, or block it.
What you get
- •One interface for Anthropic and OpenAI: the same call shape for chat, streaming and embeddings.
- •Per-customer usage: every customer (identity) has a running total, so you always know who used what.
- •Limits you enforce before you pay: a blocked call never reaches the provider, so it costs you nothing.
- •Rules you change in the Console: tiers, thresholds and who they apply to are edited without redeploying your app.
- •Alerts and billing hooks: notify Slack, call a webhook, or report to a Stripe meter when a policy fires.
- •A full trail: every call appears in Usage, with what was requested, what ran, and why.
The life of one call:
What happens on every call
- →Your code calls the SDK with an
identity(the customer) and aworkflow(which policy applies). - →The SDK asks UsageFlow whether this customer may make the call. UsageFlow applies your policy and answers allow, redirect to another model, or block.
- →If allowed, the SDK calls the AI provider from your server, using your own provider keys.
- →When the call finishes, the SDK reports the real usage and UsageFlow adds it to the customer's ledger.
🔑 Your provider keys stay with you
The SDK calls Anthropic and OpenAI directly from your server using your own ANTHROPIC_API_KEY and OPENAI_API_KEY. UsageFlow never receives your provider keys or the content of your prompts and responses — it only sees usage numbers and the identifiers you send.
What is an Identity?
The identity is the one field UsageFlow always needs: who a call's usage belongs to. It is the identity string you pass on every SDK call, and you decide how to identify your users: a customer ID, a tenant ID, a user ID (for example cust_acme).
✅ Use the same identity everywhere
Whenever you can, use the same identity you report to your other third-party services, such as Stripe. Then usage lines up with billing and reporting without any mapping work. This matters in practice: the Stripe effect looks up the Stripe customer for the identity you send, and if none matches, that usage is not reported.
Good to know
- •UsageFlow keeps one ledger per identity, per application (see the next section).
- •Use a stable ID. If a customer's identity changes, they start again from zero.
- •Nothing needs to be created in advance. The first call with a new identity creates its ledger automatically.
- •Business facts about the customer, such as their plan, are not part of the identity. Send those as customer metadata (see Your Business Logic).
What is the Ledger?
The ledger is UsageFlow's running account for one identity in one application. It is the single place that answers “how much has this customer used?” — and it is the number your policies compare against.
A ledger keeps
- •Usage this period: the running total for the identity, measured in tokens for AI calls. This is the value a tier's threshold is checked against.
- •The renewal schedule: when the total resets, set by the policy (see Usage Renewal).
- •A history of every call: what you see on the Usage page.
| Call | What is added to the ledger |
|---|---|
| Chat / streaming | Input tokens + output tokens actually used |
| Embeddings | Input tokens |
| Withdraw / credit | The amount you specify (you choose the unit; it is not converted) |
⚠️ One ledger per identity, shared by all its workflows
A customer that uses two workflows has one running total. Each workflow defines its own rules, but they are all measured against that same total. If you need a separate counter per workflow, meter each one under its own identity (for example cust_acme:support and cust_acme:search).
Policy = Workflow
In UsageFlow a policy and a workflow are the same thing. Any UsageFlow user can create a workflow for a kind of AI call in their product, for example support-agent or free-landing-page. Your code names the workflow on an SDK call to say which policy governs it.
What binding a call to a workflow gives you
- •A way to define the usage rules: how much each customer may use, what happens as they use it, and when it resets. Everything is measured against the customer's ledger.
- •Different treatment per customer: the same workflow can give a free user one ladder and a paid user another, without changing your code.
- •Change it without redeploying: you edit the workflow in the Console and it applies to the next call.
💡 Example: a free plan, then a paid plan
A free user can use up to 100 tokens with GPT-5. After that they can only use GPT-4.1 nano. After that they are blocked for the rest of the day. A paid user of the same workflow has a different ladder, with much higher thresholds. That is one workflow, two ladders. The next sections show how it is built.
How the SDK finds the workflow
The workflow ID
- •The workflow ID comes from the policy's name: lowercased, with spaces and symbols turned into dashes. A policy named
Downgrade OpenAI Modelsis addressed asdowngrade-openai-models. - •Workflow IDs are unique in your account. Each workflow's ID is shown where you edit it in the Console.
- •No workflow, no policy. If you leave
workflowout, the call is recorded but no policy is applied: nothing is limited or rerouted. - •A workflow with no matching Active policy (a typo, or a policy that is Draft or Paused) is not an error. The call simply runs without a policy, so double-check the ID.
Anatomy of a Policy
A policy (workflow) tells UsageFlow what to do as a customer's usage grows. It is a short ladder of tiers: “when usage reaches this level, do that”. You build and edit them on the Policies page in the Console (Management → Policies).
A policy is made of
- •Name (and workflow ID): what you call it, and the ID the SDK uses to address it.
- •Status: only Active policies are enforced. Draft and Paused policies are ignored.
- •Scope (optional): apply the workflow only to certain customers, by matching their metadata (for example only
region = eu). The default is Everyone. - •Tiers (1 to 3): each is “when this happens, do that”. The condition is either usage reaching N tokens, or a custom field matching a value. Tiers are checked in order and the first match wins.
- •Branches (optional, up to 3): a separate ladder of tiers for each value of a custom field, for example one per plan, region or contract type (see Metadata). A workflow has either a scope or branches, not both.
- •Effects (optional): what else happens when a tier fires (see below).
- •Renewal: how often usage resets (see Usage Renewal).
Actions
| Action | What it does |
|---|---|
| Route to a different model | Sends the call to the model you name, which can be a cheaper one. The SDK follows this automatically and works out the provider from the model name (claude-* is Anthropic, anything else is OpenAI). |
| Degrade the request | Reduce effort or quality to cut cost. The SDK reports the decision on the result so your code can act on it; it does not change the model by itself. |
| Block the request | The call is rejected outright and the provider is never contacted. Your code receives a rejection error. |
Effects
Effects run whenever a tier matches, including a block — so “tell Slack when someone is blocked” works.
| Effect | What it does |
|---|---|
| Notify Slack | Posts to a Slack integration you configured. The message can use {{policyName}}, {{tier}}, {{action}} and {{model}}. |
| Call a webhook | Sends the event to a URL you provide. |
| Report to a Stripe meter | Reports usage to a Stripe billing meter. It reports the call's amount by default. Slack and Stripe effects need an active integration of that kind. |
💡 Example
A workflow called Free Landing Page: one tier, “usage reaches 5 tokens → Block”, with a Slack effect and a renewal of once a day. Each visitor gets 5 tokens a day. Once they have used them, their next call is blocked and your team is notified.
How a Policy Is Evaluated
Policies are evaluated on every call, before the provider is contacted, in this order:
Step by step
- →The call names its workflow. No workflow means no policy, and the call is recorded only.
- →Customer metadata is saved. Anything you send as
customerMetadatais stored on the identity and used immediately, in this same call. - →The policy is found by its workflow ID. It must be Active.
- →Usage so far is read from the ledger: the identity's total before this call.
- →The first matching tier wins. If the policy has branches (see Your Business Logic), the branch that matches the customer is chosen first, then its tiers are checked in order.
- →The action is applied. Block rejects the call. Route and Degrade are returned to the SDK.
- →Effects fire for the matched tier, including for a block.
- →After the call, the real usage is added to the ledger for the next evaluation.
📏 Thresholds look at usage before the call
Because usage is read before the call, a tier at 5 tokens does not stop the call that takes a customer to 5. It affects the calls that come after usage has reached 5.
↕️ Put the strictest tier first
Only the first matching tier acts. If you list “reaches 100 → route” before “reaches 200 → block”, a customer at 250 matches the first tier and is never blocked. List the highest threshold first.
Metadata (Custom Fields)
The identity says who a customer is. Metadata says what you know about them: their plan, region, company size, a beta flag. Each piece of metadata is a custom field: a named value you configure once in the Console, and then react to in your workflows. This is what lets a single workflow treat a free customer and a paid customer differently.
How a field goes from your code to a rule
Four steps
- →Send it. Add
customerMetadatato your SDK calls, for example{ plan: 'paid', region: 'eu' }. It is saved on the identity, so it applies to that customer's later calls too, and it is used straight away by the same call. - →UsageFlow discovers it. Every new key you send appears under Management → Metadata as a discovered field, with how often it has been seen.
- →Configure it. Promote a discovered field, or create one, to make it a custom field. You give it a display name, the key (exactly the name you send, such as
plan), and a type. You can also ignore a discovered key you don't need. - →React on it. Once configured, the field can be picked in the workflow editor (see below).
What you configure
| Setting | What it is |
|---|---|
| Display name | The friendly label you see in the Console, for example “Plan”. |
| Key | The exact name your code sends, for example plan. Workflows compare against this key, so it must match exactly. |
| Type | Text, number, true/false, or a fixed list of allowed values (for example free, paid). |
| Scope | Identity fields describe the customer and persist across all their calls, like a plan. Request fields apply to a single request only. Metadata you send as customerMetadata is identity metadata. |
The Console also shows how many workflows use each field and the mix of values your customers have, so you can see who is on what before you build a ladder around it.
Three ways a workflow reacts to a field
| Where you use it | What it does | Example |
|---|---|---|
| Branches | A different ladder of tiers for each value of the field. Up to 3 branches. | plan = free gets ladder A, plan = paid gets ladder B |
| Scope | Apply the whole workflow only to customers that match. The default is Everyone. Several conditions must all hold. | Only customers where region = eu |
| A tier's condition | Fire a tier when a field equals, or does not equal, a value. A tier can also compare usage, as in “reaches 100 tokens”. | When plan equals free → Route to a cheaper model |
⚖️ Scope or branches, not both
A workflow uses either a scope or branches. If it has branches, each branch already picks its own customers, so a scope would make the branches unreachable and the editor removes it.
💡 Example
You send customerMetadata: { plan: 'free' }. UsageFlow discovers plan, you promote it to a custom field named “Plan” of type list (free, paid). In the workflow you add one branch per plan. From then on, changing a customer's plan in your own system changes which ladder they follow, with no change to UsageFlow.
Rules for customer metadata
- •Values must be text, numbers or true/false. Nested objects and lists are dropped.
- •Up to 64 keys per identity.
- •Send the latest value whenever it changes, for example when a customer upgrades. The newest value is the one used.
Your Business Logic
Every company has its own business logic. UsageFlow doesn't come with fixed plans or tiers of its own. It gives you three building blocks and you combine them the way your product works:
| Building block | What it is |
|---|---|
| Metadata | Facts about your customers that you send: anything that drives your rules. |
| Workflows | The rules: what happens at each usage level, and for which customers. |
| The ledger | The usage each identity has built up, which the rules are measured against. |
Because the facts and the rules are yours, the flow is fully dynamic. Nothing below is specific to any one kind of business.
Things people build with it
- •Different allowances per customer group: the group is whatever you call it, such as a plan, a contract type or a company size.
- •Different models per region: customers in one region are routed to a model you approve for them.
- •Trial versus production: a generous trial that blocks when its limit is reached.
- •A negotiated allowance: one large customer gets its own limits while everyone else follows the standard ladder.
- •A feature flag: a beta customer is allowed a more expensive model.
How you build it
Four steps
- →Decide which facts drive your rules. Send them as
customerMetadataunder any names you like. UsageFlow stores them on the identity. - →Configure them as custom fields under Management → Metadata (see Metadata above).
- →Build the workflow around them. Use a branch per value, a scope for who it applies to, or a tier condition, and give each group its own ladder of usage tiers.
- →Set the renewal and effects. Choose how often usage resets, and whether to notify Slack, call a webhook or report to Stripe when a tier fires.
Example: two plans
Plans are just one example of a customer group. Say you send customerMetadata: { plan: 'free' } or { plan: 'paid' }. In one workflow, add a branch for each value, with a renewal of once a day:
| Branch | Tier 1 (checked first) | Tier 2 |
|---|---|---|
plan = free | Usage reaches 200 tokens → Block (until the daily reset) | Usage reaches 100 tokens → Route to GPT-4.1 nano |
plan = paid | Usage reaches 2,000,000 tokens → Block (until the daily reset) | Usage reaches 1,000,000 tokens → Route to GPT-4.1 nano |
A free customer's calls run on GPT-5 until their usage reaches 100 tokens, then on GPT-4.1 nano, then they are blocked at 200 until the day resets. A paid customer follows the same shape with far bigger numbers. The higher threshold is listed first so that the strictest tier wins. A customer with no plan matches no branch, so nothing applies to them. Swap plan for any field of your own and the same pattern works.
Bill on the same usage
Add a Stripe meter effect to report each call's usage to your Stripe meter. What you limit and what you invoice then come from the same numbers, as long as the identity matches your Stripe customer.
Usage Renewal
A policy's renewal is how often usage resets to zero for the customers it covers. That gives you “5 calls a day” or “a monthly allowance” without any code.
| Renewal setting | Usage resets |
|---|---|
| None | Never — usage only grows |
| 10 seconds, 30 seconds, 1, 5 or 10 minutes | On that interval (handy for testing) |
| Once a day | Every day |
| Once a month | Every 30 days |
| Once a year | Every 365 days |
After a reset, a customer who was blocked is allowed again. You set the renewal on each policy, on the Policies page.
Quickstart
From nothing to a metered, limited AI call in a few minutes.
1. Create a policy
In the Console, open Management → Policies and create one (a policy is a workflow). Its name sets the workflow ID you will use in code, for example Support Agent becomes support-agent. Set it to Active.
2. Get your API key
Create or copy your application's API key under Management. Keep it on the server and never put it in browser code.
export USAGEFLOW_API_KEY="your-usageflow-api-key"
export ANTHROPIC_API_KEY="your-anthropic-key" # only for the providers you call
export OPENAI_API_KEY="your-openai-key"3. Install the SDK
Node.js / TypeScript
npm install @usageflow/vibePython
pip install usageflow-vibePython 3.9 or newer. Import from usageflow.vibe. The Python client is synchronous; from asyncio, call it with asyncio.to_thread.
Go
go get github.com/usageflow/usageflow-go-middleware/v2In Go you import github.com/usageflow/usageflow-go-middleware/v2/pkg/vibe. It needs Go 1.23.1 or newer.
4. Make a call
Node.js / TypeScript
import { usage } from '@usageflow/vibe';
usage.init({ apiKey: process.env.USAGEFLOW_API_KEY! });
const result = await usage.chat({
identity: 'cust_acme', // who this usage belongs to
workflow: 'support-agent', // which policy (workflow) governs it
provider: 'anthropic',
model: 'claude-sonnet-5',
messages: [{ role: 'user', content: 'Summarize this ticket.' }],
customerMetadata: { plan: 'pro' }, // optional: durable facts about this customer
});
console.log(result.content);Python
from usageflow.vibe import VibeClient, Message
client = VibeClient() # reads USAGEFLOW_API_KEY
result = client.chat(
identity="cust_acme", # who this usage belongs to
workflow="support-agent", # which policy (workflow) governs it
model="claude-sonnet-5", # provider is inferred from the model
messages=[Message("user", "Summarize this ticket.")],
customer_metadata={"plan": "pro"}, # optional
)
print(result.content)Go
res, err := client.Chat(ctx, vibe.ChatRequest{
Identity: "cust_acme", // who this usage belongs to
Workflow: "support-agent", // which policy (workflow) governs it
Provider: vibe.ProviderAnthropic,
Model: "claude-sonnet-5",
Messages: []vibe.Message{{Role: "user", Content: "Summarize this ticket."}},
CustomerMetadata: map[string]any{"plan": "pro"}, // optional
})
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Content)5. See it in the Console
Open Usage: the call appears with the customer, what was requested and what ran. On the Dashboard, Live decisions updates as calls come in. If your policy fired, the reason is shown next to the call.
How the SDK Works
Create the client once when your app starts and reuse it. Each AI call then goes through three steps:
Every call, in three steps
- →Check first. The SDK asks UsageFlow whether the customer may make the call and waits for the answer. UsageFlow sets aside a worst-case amount for the call: the estimated size of your input plus the maximum output. If you don't set
maxTokens, output is capped at 1024 tokens. A denial stops here. - →Call the provider. If allowed, the SDK calls Anthropic or OpenAI using the model in your request, or the model your policy redirected to.
- →Settle. The SDK reports the real usage (never more than was set aside) and the customer's ledger is updated.
Things to rely on
- •A denied call never reaches the provider, so a blocked customer costs you nothing.
- •If UsageFlow can't be reached, the call fails instead of running unmetered. Nothing is sent to the provider.
- •If reporting the final usage fails after a successful chat call, the error is logged but you still get your result, because the provider call already happened.
- •Provider clients are created on first use, so a missing Anthropic key only fails when you actually call Anthropic.
- •OpenAI reasoning models (
o1,o3,o4,gpt-5) are handled for you: the SDK sends the right output-limit parameter and leaves temperature out.
Chat
chat makes one metered, policy-checked call to Anthropic or OpenAI and returns the full response.
| Field | What it is |
|---|---|
identity | Required. The customer this usage belongs to. |
provider | anthropic or openai. |
model | Any model ID your provider accepts. |
messages | The conversation: system, user and assistant turns. |
workflow | Optional. The policy that governs this call. |
customerMetadata | Optional. Durable facts about the customer, such as their plan. |
maxTokens | Optional. Output cap. Defaults to 1024. |
temperature | Optional. Ignored for OpenAI reasoning models. |
tools | Optional. Passed to the provider as you give them. |
Node.js / TypeScript
const result = await usage.chat({
identity: 'cust_acme', // who this usage belongs to
workflow: 'support-agent', // which policy (workflow) governs it
provider: 'anthropic',
model: 'claude-sonnet-5',
messages: [{ role: 'user', content: 'Summarize this ticket.' }],
customerMetadata: { plan: 'pro' }, // optional: durable facts about this customer
});
console.log(result.content);Python
result = client.chat(
identity="cust_acme", # who this usage belongs to
workflow="support-agent", # which policy (workflow) governs it
model="claude-sonnet-5", # provider is inferred from the model
messages=[Message("user", "Summarize this ticket.")],
customer_metadata={"plan": "pro"}, # optional
)
print(result.content)Go
res, err := client.Chat(ctx, vibe.ChatRequest{
Identity: "cust_acme", // who this usage belongs to
Workflow: "support-agent", // which policy (workflow) governs it
Provider: vibe.ProviderAnthropic,
Model: "claude-sonnet-5",
Messages: []vibe.Message{{Role: "user", Content: "Summarize this ticket."}},
CustomerMetadata: map[string]any{"plan": "pro"}, // optional
})
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Content)The result has content, usage (input and output tokens), the model that actually ran, the requestedModel you asked for, any toolCalls, and vibePolicy when a policy tier fired.
Streaming
stream works like chat, with the same check and the same denial behavior. Text arrives as it is generated, and usage is settled when the stream closes. You must read the stream to the end for that to happen.
Node.js / TypeScript
const stream = await usage.stream({ identity: 'cust_acme', workflow: 'support-agent', provider: 'openai', model: 'gpt-4o-mini', messages });
for await (const chunk of stream.textStream) {
process.stdout.write(chunk);
}
const final = await stream.finalResult; // resolves after usage is settledPython
stream = client.stream(identity="cust_acme", model="gpt-4o-mini",
messages=[Message("user", "Tell me a story")])
for chunk in stream:
print(chunk, end="", flush=True)
result = stream.result() # usage is recorded when the stream endsGo
s, err := client.Stream(ctx, req)
if err != nil {
return err
}
for text := range s.Text {
fmt.Print(text)
}
res, err := s.Result() // available after usage is settledEmbeddings
embed creates embeddings with OpenAI. It is checked and metered like chat, but on input tokens only, since embeddings have no output. Embeddings are OpenAI-only.
Node.js / TypeScript
const result = await usage.embed({
identity: 'cust_acme',
model: 'text-embedding-3-small',
input: ['first text', 'second text'],
});
result.embeddings; // one vector per inputPython
result = client.embed(identity="cust_acme", model="text-embedding-3-small", input=["hello"])Go
res, err := client.Embed(ctx, vibe.EmbedRequest{
Identity: "cust_acme",
Model: "text-embedding-3-small",
Input: []string{"first text", "second text"},
})
// res.Embeddings: one vector per inputPolicy Decisions
When a policy tier fires, the SDK acts on it and tells you what happened, so you never have to guess why a call ran the way it did.
What you see on the result
- •Route to a model: the call runs on the model the policy named, even on a different provider.
modelandprovidershow what ran, andrequestedModelshows what you asked for. - •Any fired tier:
vibePolicyholds the policy ID, the tier index and the action. Log it if you want an audit of policy decisions in your own systems. - •Degrade: the original model is used untouched. Read
vibePolicyto apply the degrade yourself. - •Block: there is no result. You get a rejection error instead (next section).
const result = await usage.chat({ identity: 'cust_acme', workflow: 'downgrade-openai-models', provider: 'anthropic', model: 'claude-sonnet-5', messages });
if (result.vibePolicy) {
console.log(`Policy ${result.vibePolicy.policyId}, tier ${result.vibePolicy.tierIndex} fired`);
}
// If the policy rerouted the call, result.provider / result.model already reflect it.Withdraw & Credit
Not all usage is an AI call: a nightly batch job, a correction, a fixed fee. withdraw takes an amount off a customer's allowance directly, and credit reverses one. They go through the same check as AI calls (in Go you can also name a Workflow so a policy applies). Each needs an idempotencyKey that identifies the adjustment.
Node.js / TypeScript
await usage.withdraw({ identity: 'cust_acme', amount: 500, idempotencyKey: 'job-42-attempt-1', reason: 'nightly batch reconciliation' });
await usage.credit({ identity: 'cust_acme', amount: 500, idempotencyKey: 'job-42-reversal-1', reason: 'correct over-count from job-42' });Go
res, err := client.Withdraw(ctx, vibe.WithdrawRequest{
Identity: "cust_acme", Amount: 500, IdempotencyKey: "job-42-attempt-1",
Reason: "nightly batch reconciliation", Workflow: "support-agent", // Workflow is optional
})Reserve now, settle later (Go)
When the final amount isn't known up front, the Go SDK can hold the amount first and settle it when the work is done. WithdrawAsync reserves and returns a CaptureID; Close settles it, by default for the reserved amount or for a smaller final one. A capture can only be closed once. This is available in the Go SDK only.
capture, err := client.WithdrawAsync(ctx, vibe.WithdrawRequest{Identity: "cust_acme", Amount: 100, IdempotencyKey: "job-43"})
// ... do the work ...
final := 60.0
_, err = client.Close(ctx, vibe.CloseRequest{CaptureID: capture.CaptureID, Amount: &final})⚠️ Current limits
idempotencyKey is recorded for your audit trail, but a repeated call with the same key is not yet de-duplicated, so don't retry a withdraw blindly. credit sends the amount as a negative adjustment; confirm it behaves as you expect for your plan before you rely on it in production.
Errors & Blocked Calls
When a policy blocks a call, or the customer is out of allowance, the SDK raises a rejection error. This is the normal way a limit shows up in your code, so handle it explicitly and tell your user.
Node.js / TypeScript
import { usage, VibeRejectionError } from '@usageflow/vibe';
try {
const result = await usage.chat({ /* ... */ });
res.json({ answer: result.content });
} catch (err) {
if (err instanceof VibeRejectionError) {
// UsageFlow denied the call — the provider was never contacted.
res.status(429).json({ error: 'Usage limit reached', reason: err.reason });
} else {
throw err;
}
}Go
res, err := client.Chat(ctx, req)
var rejected *vibe.RejectionError
if errors.As(err, &rejected) {
// UsageFlow denied the call — the provider was never contacted.
http.Error(w, "Usage limit reached: "+rejected.Message, http.StatusTooManyRequests)
return
}
if err != nil {
return err
}| What happened | What you get |
|---|---|
| A policy blocks the call, or usage is over the limit | A rejection error. The provider was not called. |
| UsageFlow can't be reached | An ordinary error. The provider was not called. |
| The provider returns an error | That error, unchanged. |
| A required provider key is missing | An error naming the key to set. |
Reference
Configuration
| Setting | Node.js option | Go option | Environment variable |
|---|---|---|---|
| UsageFlow API key (required) | apiKey | APIKey | USAGEFLOW_API_KEY |
| Anthropic key | anthropicApiKey | AnthropicAPIKey | ANTHROPIC_API_KEY |
| OpenAI key | openaiApiKey | OpenAIAPIKey | OPENAI_API_KEY |
What each SDK supports
| Call | Node.js | Go |
|---|---|---|
| Chat | ✓ | ✓ |
| Streaming | ✓ | ✓ |
| Embeddings (OpenAI) | ✓ | ✓ |
| Withdraw / credit | ✓ | ✓ |
| Reserve now, settle later | — | ✓ |
| Image, speech, transcription, moderation, batch | ✓ (see below) | — |
The two SDKs are not identical. Check this table before you plan around a feature.
Also in the Node.js SDK (early)
- •
image,speakandtranscribeare OpenAI-only. Their usage is counted with a flat placeholder weight, not the provider's real pricing, so don't use them for billing yet. - •
moderateis a passthrough to OpenAI's moderation and is not metered. - •
batchSubmit,batchStatusandbatchSettlerun a provider batch job. Settling a batch must happen in the same running process that submitted it.
Where to See It in the Console
| Page | What it shows |
|---|---|
| Dashboard | Live decisions as calls happen, how requests were changed by your policies (requested versus executed model), and your blocked requests. |
| Usage | Every call: the customer, what was requested and what ran, why (which policy fired), and the impact. You can also see policy decisions, aggregates, top spenders, models and identities. Withdraw and credit adjustments appear as local ledger rows, in credits rather than a model. |
| Management → Policies | Create and edit your policies (workflows): tiers, branches, effects and renewal, and each workflow's ID. |
| Management | Your applications and API keys, your customers (identities), and their metadata. |
🧭 Not seeing a policy fire?
Check the workflow first. Make sure the workflow in your code exactly matches the policy's workflow ID on the Policies page, and that the policy is Active. A call with no matching policy is not an error; it just runs without one.
Framework Packages
Separate from the Vibe SDK, UsageFlow publishes packages that meter the HTTP routes of your own API: Express, Fastify and NestJS for Node.js, FastAPI and Flask for Python, and Gin for Go. Install the one for your runtime below.
Node.js Frameworks
Install the UsageFlow package for your Node.js framework:
npm install @usageflow/express
npm install @usageflow/fastify
npm install @usageflow/nestjsExpress
import express from "express";
import { ExpressUsageFlowAPI } from "@usageflow/express";
const apiKey = process.env.USAGEFLOW_API_KEY;
if (!apiKey) throw new Error("Missing USAGEFLOW_API_KEY");
const app = express();
app.use(express.json());
const usageFlow = new ExpressUsageFlowAPI(apiKey);
app.use(usageFlow.createMiddleware());
app.get("/api/users", (_req, res) => res.json([]));Fastify
import Fastify from "fastify";
import { FastifyUsageFlowAPI } from "@usageflow/fastify";
const apiKey = process.env.USAGEFLOW_API_KEY;
if (!apiKey) throw new Error("Missing USAGEFLOW_API_KEY");
const app = Fastify();
const usageFlow = new FastifyUsageFlowAPI(apiKey);
await app.register(usageFlow.createPlugin());
app.get("/api/users", async () => []);NestJS
import { Module } from "@nestjs/common";
import { UsageFlowModule } from "@usageflow/nestjs";
const apiKey = process.env.USAGEFLOW_API_KEY;
if (!apiKey) throw new Error("Missing USAGEFLOW_API_KEY");
@Module({
imports: [UsageFlowModule.forRoot({ apiKey })],
})
export class AppModule {}Current limitation
Current limitation: module registration alone does not start request processing reliably, so NestJS request traces are not currently supported. Use the Express or Fastify package directly where possible.
Python Frameworks
Install the UsageFlow package for your Python framework:
pip install usageflow-flask
pip install usageflow-fastapi uvicornFastAPI
import os
from fastapi import FastAPI
from usageflow.fastapi import UsageFlowMiddleware
app = FastAPI()
app.add_middleware(
UsageFlowMiddleware,
api_key=os.environ["USAGEFLOW_API_KEY"],
)
@app.get("/api/users")
def users():
return []FastAPI production limitation
Start the app with `uvicorn app:app`, send the test request, then confirm the Trace in Console. Do not enable this middleware on routes that depend on FastAPI or Starlette response background tasks; the current middleware replaces an existing response background task.
Flask
import os
from flask import Flask
from usageflow.flask import UsageFlowMiddleware
app = Flask(__name__)
UsageFlowMiddleware(app, api_key=os.environ["USAGEFLOW_API_KEY"])
@app.get("/api/users")
def users():
return []Go
Install the UsageFlow middleware package for Go (Gin framework):
go get github.com/usageflow/usageflow-go-middleware/v2Gin v2
The Go middleware provides request interception, usage tracking, user identification, and automatic configuration updates with graceful degradation.
package main
import (
"log"
"os"
"github.com/gin-gonic/gin"
"github.com/usageflow/usageflow-go-middleware/v2/pkg/middleware"
)
func main() {
apiKey := os.Getenv("USAGEFLOW_API_KEY")
if apiKey == "" {
log.Fatal("Missing USAGEFLOW_API_KEY")
}
router := gin.New()
usageFlow := middleware.New(apiKey)
router.Use(usageFlow.RequestInterceptor())
router.GET("/api/users", func(c *gin.Context) {
c.JSON(200, []string{})
})
if err := router.Run(":8080"); err != nil {
log.Fatal(err)
}
}Migrating from the legacy Gin package
github.com/usageflow/usageflow-gin is deprecated and does not support Console-managed routes. Replace it with the v2 package above, register RequestInterceptor() without local route lists, then configure routes in Console.
Best Practices
API Key Security
Never commit API keys to version control, and keep your UsageFlow and provider keys on the server, never in browser code. Always use environment variables.
Resources & Support
Find packages, examples, and get help:
Need help? Contact our support team:
- Email: [email protected]