UsageFlow Docs

Control AI usage at runtime.

Put Vibe in front of your AI calls. Define the rules in the Console, not in your code, and decide what each customer may do before you pay the provider.

await usage.chat({
  identity: 'user_123',
  workflow: 'ai-chat',
  model: 'gpt-5',
  messages,
});

usage      184 tokens · Free plan
decision   ROUTE gpt-5 → gpt-4.1-nano
settled    637 tokens

The idea in four steps

What you can do

SDK

Install, create the client once, then make policy-checked calls. Provider keys stay on your server.

Node.js
npm install @usageflow/vibe
TypeScript
const result = await usage.chat({
  identity: 'cust_acme',            // who this usage belongs to
  workflow: 'support-agent',        // which policy (workflow) governs it
  provider: 'anthropic',
  model: 'claude-sonnet-5',
  messages: [{ role: 'user', content: 'Summarize this ticket.' }],
  customerMetadata: { plan: 'pro' },  // optional: durable facts about this customer
});

console.log(result.content);

Choose your path

Control AI usage at runtime

Vibe checks every AI request before it reaches the provider and applies the policy for that customer. Rules live in the Console, not in your code, so you can change what each customer is allowed to do without a deploy.

MeterLimitRouteDegradeBlockSettleBill

1. Without Vibe, you write the rules yourself

  1. User
  2. Your app
  3. AI provider

Nothing stops one customer from spending the whole budget, and every plan rule (“free users get the cheap model after 100 tokens”) ends up as if statements in your app.

2. With Vibe, the rules sit in front of the provider

  1. Your app
  2. Vibeidentity + usage + policy
  3. Decisionallow · route · degrade · block
  4. AI provideronly if allowed
  5. Settlementactual usage recorded

A real policy, end to end

Start with the outcome. This is a policy for a free customer, read top to bottom as their usage grows. It is an example; you set your own thresholds and models in the Console.

Free customer · workflow ai-chat

0 – 100 tokensGPT-5
100 – 200 tokensRoute to GPT-4.1 nano
200+ tokensBlock. The provider is never called

Paid customer · same workflow

0 – 1,000,000 tokensGPT-5
1,000,000 – 2,000,000Route to GPT-4.1 nano
2,000,000+ tokensBlock until the daily reset

The same request, at 184 tokens of usage, for the free customer:

Identity
user_123
Usage
184 tokens
Policy
Free plan · ai-chat
Decision
Route
Requested → ran
GPT-5 → GPT-4.1 nano

Same code, different policy

This is the entire integration. The call is identical for the free and the paid customer; only the policy behind ai-chat differs.

TypeScript
// Your app: no plan checks, no model switching, no limit math.
const result = await usage.chat({
  identity: 'user_123',     // who the usage belongs to
  workflow: 'ai-chat',      // which policy applies
  provider: 'openai',
  model: 'gpt-5',
  messages,
});
  1. Application code
  2. Vibe
  3. Policy from the Console
ActionProvider called?Result
AllowYesRequested model
RouteYesA different model, possibly another provider
DegradeYesRequested model; you apply the lower-cost behavior
BlockNoRequest rejected

How Vibe works

Before the call

Vibe knows who is calling (identity), which workflow, how much they have used, their metadata, the requested model and your thresholds. It then decides: allow, route, degrade or block.

After the call

Vibe does not just count requests. It reserves a worst-case amount first, so a request can never push a customer past their allowance, and then settles what the provider actually used.

1. Reserve1,024worst-case estimate
2. Providercallfrom your server
3. Settle637ledger += 637

Settlement never charges more than was reserved. Concepts that make this work, in the order you meet them: Identity (whose usage), Ledger (how much), Policy (what is allowed), then tiers and settlement.

Where your data goes

Your server

Your appVibe SDKProvider keys

UsageFlow

Policy + usage decision

AI provider

OpenAI · Anthropic
  • Provider credentials stay in your environment.
  • Vibe asks UsageFlow for the decision and reports usage numbers for metering and settlement.
  • Prompts and responses go from your server straight to the provider, not through UsageFlow.
Concepts

What is UsageFlow Vibe?

UsageFlow Vibe sits between your app and the AI providers. You make your AI calls through one SDK. Before each call UsageFlow checks the customer's usage against your policies, and after the call it records what was really used. You decide what happens at each usage level: carry on, route to a cheaper model, degrade the request, or block it.

What you get

  • •One interface for Anthropic and OpenAI: the same call shape for chat, streaming and embeddings.
  • •Per-customer usage: every customer (identity) has a running total, so you always know who used what.
  • •Limits you enforce before you pay: a blocked call never reaches the provider, so it costs you nothing.
  • •Rules you change in the Console: tiers, thresholds and who they apply to are edited without redeploying your app.
  • •Alerts and billing hooks: notify Slack, call a webhook, or report to a Stripe meter when a policy fires.
  • •A full trail: every call appears in Usage, with what was requested, what ran, and why.

The life of one call:

What happens on every call

  • →Your code calls the SDK with an identity (the customer) and a workflow (which policy applies).
  • →The SDK asks UsageFlow whether this customer may make the call. UsageFlow applies your policy and answers allow, redirect to another model, or block.
  • →If allowed, the SDK calls the AI provider from your server, using your own provider keys.
  • →When the call finishes, the SDK reports the real usage and UsageFlow adds it to the customer's ledger.

🔑 Your provider keys stay with you

The SDK calls Anthropic and OpenAI directly from your server using your own ANTHROPIC_API_KEY and OPENAI_API_KEY. UsageFlow never receives your provider keys or the content of your prompts and responses — it only sees usage numbers and the identifiers you send.

Concepts

What is an Identity?

The identity is the one field UsageFlow always needs: who a call's usage belongs to. It is the identity string you pass on every SDK call, and you decide how to identify your users: a customer ID, a tenant ID, a user ID (for example cust_acme).

✅ Use the same identity everywhere

Whenever you can, use the same identity you report to your other third-party services, such as Stripe. Then usage lines up with billing and reporting without any mapping work. This matters in practice: the Stripe effect looks up the Stripe customer for the identity you send, and if none matches, that usage is not reported.

Good to know

  • •UsageFlow keeps one ledger per identity, per application (see the next section).
  • •Use a stable ID. If a customer's identity changes, they start again from zero.
  • •Nothing needs to be created in advance. The first call with a new identity creates its ledger automatically.
  • •Business facts about the customer, such as their plan, are not part of the identity. Send those as customer metadata (see Your Business Logic).
Concepts

What is the Ledger?

The ledger is UsageFlow's running account for one identity in one application. It is the single place that answers “how much has this customer used?” — and it is the number your policies compare against.

A ledger keeps

  • •Usage this period: the running total for the identity, measured in tokens for AI calls. This is the value a tier's threshold is checked against.
  • •The renewal schedule: when the total resets, set by the policy (see Usage Renewal).
  • •A history of every call: what you see on the Usage page.
CallWhat is added to the ledger
Chat / streamingInput tokens + output tokens actually used
EmbeddingsInput tokens
Withdraw / creditThe amount you specify (you choose the unit; it is not converted)

⚠️ One ledger per identity, shared by all its workflows

A customer that uses two workflows has one running total. Each workflow defines its own rules, but they are all measured against that same total. If you need a separate counter per workflow, meter each one under its own identity (for example cust_acme:support and cust_acme:search).

Concepts

Policy = Workflow

In UsageFlow a policy and a workflow are the same thing. Any UsageFlow user can create a workflow for a kind of AI call in their product, for example support-agent or free-landing-page. Your code names the workflow on an SDK call to say which policy governs it.

What binding a call to a workflow gives you

  • •A way to define the usage rules: how much each customer may use, what happens as they use it, and when it resets. Everything is measured against the customer's ledger.
  • •Different treatment per customer: the same workflow can give a free user one ladder and a paid user another, without changing your code.
  • •Change it without redeploying: you edit the workflow in the Console and it applies to the next call.

💡 Example: a free plan, then a paid plan

A free user can use up to 100 tokens with GPT-5. After that they can only use GPT-4.1 nano. After that they are blocked for the rest of the day. A paid user of the same workflow has a different ladder, with much higher thresholds. That is one workflow, two ladders. The next sections show how it is built.

How the SDK finds the workflow

The workflow ID

  • •The workflow ID comes from the policy's name: lowercased, with spaces and symbols turned into dashes. A policy named Downgrade OpenAI Models is addressed as downgrade-openai-models.
  • •Workflow IDs are unique in your account. Each workflow's ID is shown where you edit it in the Console.
  • •No workflow, no policy. If you leave workflow out, the call is recorded but no policy is applied: nothing is limited or rerouted.
  • •A workflow with no matching Active policy (a typo, or a policy that is Draft or Paused) is not an error. The call simply runs without a policy, so double-check the ID.
Concepts

Anatomy of a Policy

A policy (workflow) tells UsageFlow what to do as a customer's usage grows. It is a short ladder of tiers: “when usage reaches this level, do that”. You build and edit them on the Policies page in the Console (Management → Policies).

A policy is made of

  • •Name (and workflow ID): what you call it, and the ID the SDK uses to address it.
  • •Status: only Active policies are enforced. Draft and Paused policies are ignored.
  • •Scope (optional): apply the workflow only to certain customers, by matching their metadata (for example only region = eu). The default is Everyone.
  • •Tiers (1 to 3): each is “when this happens, do that”. The condition is either usage reaching N tokens, or a custom field matching a value. Tiers are checked in order and the first match wins.
  • •Branches (optional, up to 3): a separate ladder of tiers for each value of a custom field, for example one per plan, region or contract type (see Metadata). A workflow has either a scope or branches, not both.
  • •Effects (optional): what else happens when a tier fires (see below).
  • •Renewal: how often usage resets (see Usage Renewal).

Actions

ActionWhat it does
Route to a different modelSends the call to the model you name, which can be a cheaper one. The SDK follows this automatically and works out the provider from the model name (claude-* is Anthropic, anything else is OpenAI).
Degrade the requestReduce effort or quality to cut cost. The SDK reports the decision on the result so your code can act on it; it does not change the model by itself.
Block the requestThe call is rejected outright and the provider is never contacted. Your code receives a rejection error.

Effects

Effects run whenever a tier matches, including a block — so “tell Slack when someone is blocked” works.

EffectWhat it does
Notify SlackPosts to a Slack integration you configured. The message can use {{policyName}}, {{tier}}, {{action}} and {{model}}.
Call a webhookSends the event to a URL you provide.
Report to a Stripe meterReports usage to a Stripe billing meter. It reports the call's amount by default. Slack and Stripe effects need an active integration of that kind.

💡 Example

A workflow called Free Landing Page: one tier, “usage reaches 5 tokens → Block”, with a Slack effect and a renewal of once a day. Each visitor gets 5 tokens a day. Once they have used them, their next call is blocked and your team is notified.

Concepts

How a Policy Is Evaluated

Policies are evaluated on every call, before the provider is contacted, in this order:

Step by step

  • →The call names its workflow. No workflow means no policy, and the call is recorded only.
  • →Customer metadata is saved. Anything you send as customerMetadata is stored on the identity and used immediately, in this same call.
  • →The policy is found by its workflow ID. It must be Active.
  • →Usage so far is read from the ledger: the identity's total before this call.
  • →The first matching tier wins. If the policy has branches (see Your Business Logic), the branch that matches the customer is chosen first, then its tiers are checked in order.
  • →The action is applied. Block rejects the call. Route and Degrade are returned to the SDK.
  • →Effects fire for the matched tier, including for a block.
  • →After the call, the real usage is added to the ledger for the next evaluation.

📏 Thresholds look at usage before the call

Because usage is read before the call, a tier at 5 tokens does not stop the call that takes a customer to 5. It affects the calls that come after usage has reached 5.

↕️ Put the strictest tier first

Only the first matching tier acts. If you list “reaches 100 → route” before “reaches 200 → block”, a customer at 250 matches the first tier and is never blocked. List the highest threshold first.

Concepts

Metadata (Custom Fields)

The identity says who a customer is. Metadata says what you know about them: their plan, region, company size, a beta flag. Each piece of metadata is a custom field: a named value you configure once in the Console, and then react to in your workflows. This is what lets a single workflow treat a free customer and a paid customer differently.

How a field goes from your code to a rule

Four steps

  • →Send it. Add customerMetadata to your SDK calls, for example { plan: 'paid', region: 'eu' }. It is saved on the identity, so it applies to that customer's later calls too, and it is used straight away by the same call.
  • →UsageFlow discovers it. Every new key you send appears under Management → Metadata as a discovered field, with how often it has been seen.
  • →Configure it. Promote a discovered field, or create one, to make it a custom field. You give it a display name, the key (exactly the name you send, such as plan), and a type. You can also ignore a discovered key you don't need.
  • →React on it. Once configured, the field can be picked in the workflow editor (see below).

What you configure

SettingWhat it is
Display nameThe friendly label you see in the Console, for example “Plan”.
KeyThe exact name your code sends, for example plan. Workflows compare against this key, so it must match exactly.
TypeText, number, true/false, or a fixed list of allowed values (for example free, paid).
ScopeIdentity fields describe the customer and persist across all their calls, like a plan. Request fields apply to a single request only. Metadata you send as customerMetadata is identity metadata.

The Console also shows how many workflows use each field and the mix of values your customers have, so you can see who is on what before you build a ladder around it.

Three ways a workflow reacts to a field

Where you use itWhat it doesExample
BranchesA different ladder of tiers for each value of the field. Up to 3 branches.plan = free gets ladder A, plan = paid gets ladder B
ScopeApply the whole workflow only to customers that match. The default is Everyone. Several conditions must all hold.Only customers where region = eu
A tier's conditionFire a tier when a field equals, or does not equal, a value. A tier can also compare usage, as in “reaches 100 tokens”.When plan equals free → Route to a cheaper model

⚖️ Scope or branches, not both

A workflow uses either a scope or branches. If it has branches, each branch already picks its own customers, so a scope would make the branches unreachable and the editor removes it.

💡 Example

You send customerMetadata: { plan: 'free' }. UsageFlow discovers plan, you promote it to a custom field named “Plan” of type list (free, paid). In the workflow you add one branch per plan. From then on, changing a customer's plan in your own system changes which ladder they follow, with no change to UsageFlow.

Rules for customer metadata

  • •Values must be text, numbers or true/false. Nested objects and lists are dropped.
  • •Up to 64 keys per identity.
  • •Send the latest value whenever it changes, for example when a customer upgrades. The newest value is the one used.
Concepts

Your Business Logic

Every company has its own business logic. UsageFlow doesn't come with fixed plans or tiers of its own. It gives you three building blocks and you combine them the way your product works:

Building blockWhat it is
MetadataFacts about your customers that you send: anything that drives your rules.
WorkflowsThe rules: what happens at each usage level, and for which customers.
The ledgerThe usage each identity has built up, which the rules are measured against.

Because the facts and the rules are yours, the flow is fully dynamic. Nothing below is specific to any one kind of business.

Things people build with it

  • •Different allowances per customer group: the group is whatever you call it, such as a plan, a contract type or a company size.
  • •Different models per region: customers in one region are routed to a model you approve for them.
  • •Trial versus production: a generous trial that blocks when its limit is reached.
  • •A negotiated allowance: one large customer gets its own limits while everyone else follows the standard ladder.
  • •A feature flag: a beta customer is allowed a more expensive model.

How you build it

Four steps

  • →Decide which facts drive your rules. Send them as customerMetadata under any names you like. UsageFlow stores them on the identity.
  • →Configure them as custom fields under Management → Metadata (see Metadata above).
  • →Build the workflow around them. Use a branch per value, a scope for who it applies to, or a tier condition, and give each group its own ladder of usage tiers.
  • →Set the renewal and effects. Choose how often usage resets, and whether to notify Slack, call a webhook or report to Stripe when a tier fires.

Example: two plans

Plans are just one example of a customer group. Say you send customerMetadata: { plan: 'free' } or { plan: 'paid' }. In one workflow, add a branch for each value, with a renewal of once a day:

BranchTier 1 (checked first)Tier 2
plan = freeUsage reaches 200 tokens → Block (until the daily reset)Usage reaches 100 tokens → Route to GPT-4.1 nano
plan = paidUsage reaches 2,000,000 tokens → Block (until the daily reset)Usage reaches 1,000,000 tokens → Route to GPT-4.1 nano

A free customer's calls run on GPT-5 until their usage reaches 100 tokens, then on GPT-4.1 nano, then they are blocked at 200 until the day resets. A paid customer follows the same shape with far bigger numbers. The higher threshold is listed first so that the strictest tier wins. A customer with no plan matches no branch, so nothing applies to them. Swap plan for any field of your own and the same pattern works.

Bill on the same usage

Add a Stripe meter effect to report each call's usage to your Stripe meter. What you limit and what you invoice then come from the same numbers, as long as the identity matches your Stripe customer.

Concepts

Usage Renewal

A policy's renewal is how often usage resets to zero for the customers it covers. That gives you “5 calls a day” or “a monthly allowance” without any code.

Renewal settingUsage resets
NoneNever — usage only grows
10 seconds, 30 seconds, 1, 5 or 10 minutesOn that interval (handy for testing)
Once a dayEvery day
Once a monthEvery 30 days
Once a yearEvery 365 days

After a reset, a customer who was blocked is allowed again. You set the renewal on each policy, on the Policies page.

Get started

Quickstart

From nothing to a metered, limited AI call in a few minutes.

1. Create a policy

In the Console, open Management → Policies and create one (a policy is a workflow). Its name sets the workflow ID you will use in code, for example Support Agent becomes support-agent. Set it to Active.

2. Get your API key

Create or copy your application's API key under Management. Keep it on the server and never put it in browser code.

Bash
export USAGEFLOW_API_KEY="your-usageflow-api-key"
export ANTHROPIC_API_KEY="your-anthropic-key"   # only for the providers you call
export OPENAI_API_KEY="your-openai-key"

3. Install the SDK

Node.js / TypeScript

Node.js
npm install @usageflow/vibe

Python

Python
pip install usageflow-vibe

Python 3.9 or newer. Import from usageflow.vibe. The Python client is synchronous; from asyncio, call it with asyncio.to_thread.

Go

Go
go get github.com/usageflow/usageflow-go-middleware/v2

In Go you import github.com/usageflow/usageflow-go-middleware/v2/pkg/vibe. It needs Go 1.23.1 or newer.

4. Make a call

Node.js / TypeScript

TypeScript
import { usage } from '@usageflow/vibe';

usage.init({ apiKey: process.env.USAGEFLOW_API_KEY! });

const result = await usage.chat({
  identity: 'cust_acme',            // who this usage belongs to
  workflow: 'support-agent',        // which policy (workflow) governs it
  provider: 'anthropic',
  model: 'claude-sonnet-5',
  messages: [{ role: 'user', content: 'Summarize this ticket.' }],
  customerMetadata: { plan: 'pro' },  // optional: durable facts about this customer
});

console.log(result.content);

Python

Python
from usageflow.vibe import VibeClient, Message

client = VibeClient()  # reads USAGEFLOW_API_KEY

result = client.chat(
    identity="cust_acme",          # who this usage belongs to
    workflow="support-agent",      # which policy (workflow) governs it
    model="claude-sonnet-5",       # provider is inferred from the model
    messages=[Message("user", "Summarize this ticket.")],
    customer_metadata={"plan": "pro"},  # optional
)

print(result.content)

Go

Go
res, err := client.Chat(ctx, vibe.ChatRequest{
	Identity:         "cust_acme",       // who this usage belongs to
	Workflow:         "support-agent",   // which policy (workflow) governs it
	Provider:         vibe.ProviderAnthropic,
	Model:            "claude-sonnet-5",
	Messages:         []vibe.Message{{Role: "user", Content: "Summarize this ticket."}},
	CustomerMetadata: map[string]any{"plan": "pro"}, // optional
})
if err != nil {
	log.Fatal(err)
}
fmt.Println(res.Content)

5. See it in the Console

Open Usage: the call appears with the customer, what was requested and what ran. On the Dashboard, Live decisions updates as calls come in. If your policy fired, the reason is shown next to the call.

Vibe SDK

How the SDK Works

Create the client once when your app starts and reuse it. Each AI call then goes through three steps:

Every call, in three steps

  • →Check first. The SDK asks UsageFlow whether the customer may make the call and waits for the answer. UsageFlow sets aside a worst-case amount for the call: the estimated size of your input plus the maximum output. If you don't set maxTokens, output is capped at 1024 tokens. A denial stops here.
  • →Call the provider. If allowed, the SDK calls Anthropic or OpenAI using the model in your request, or the model your policy redirected to.
  • →Settle. The SDK reports the real usage (never more than was set aside) and the customer's ledger is updated.

Things to rely on

  • •A denied call never reaches the provider, so a blocked customer costs you nothing.
  • •If UsageFlow can't be reached, the call fails instead of running unmetered. Nothing is sent to the provider.
  • •If reporting the final usage fails after a successful chat call, the error is logged but you still get your result, because the provider call already happened.
  • •Provider clients are created on first use, so a missing Anthropic key only fails when you actually call Anthropic.
  • •OpenAI reasoning models (o1, o3, o4, gpt-5) are handled for you: the SDK sends the right output-limit parameter and leaves temperature out.
Vibe SDK

Chat

chat makes one metered, policy-checked call to Anthropic or OpenAI and returns the full response.

FieldWhat it is
identityRequired. The customer this usage belongs to.
provideranthropic or openai.
modelAny model ID your provider accepts.
messagesThe conversation: system, user and assistant turns.
workflowOptional. The policy that governs this call.
customerMetadataOptional. Durable facts about the customer, such as their plan.
maxTokensOptional. Output cap. Defaults to 1024.
temperatureOptional. Ignored for OpenAI reasoning models.
toolsOptional. Passed to the provider as you give them.

Node.js / TypeScript

TypeScript
const result = await usage.chat({
  identity: 'cust_acme',            // who this usage belongs to
  workflow: 'support-agent',        // which policy (workflow) governs it
  provider: 'anthropic',
  model: 'claude-sonnet-5',
  messages: [{ role: 'user', content: 'Summarize this ticket.' }],
  customerMetadata: { plan: 'pro' },  // optional: durable facts about this customer
});

console.log(result.content);

Python

Python
result = client.chat(
    identity="cust_acme",          # who this usage belongs to
    workflow="support-agent",      # which policy (workflow) governs it
    model="claude-sonnet-5",       # provider is inferred from the model
    messages=[Message("user", "Summarize this ticket.")],
    customer_metadata={"plan": "pro"},  # optional
)

print(result.content)

Go

Go
res, err := client.Chat(ctx, vibe.ChatRequest{
	Identity:         "cust_acme",       // who this usage belongs to
	Workflow:         "support-agent",   // which policy (workflow) governs it
	Provider:         vibe.ProviderAnthropic,
	Model:            "claude-sonnet-5",
	Messages:         []vibe.Message{{Role: "user", Content: "Summarize this ticket."}},
	CustomerMetadata: map[string]any{"plan": "pro"}, // optional
})
if err != nil {
	log.Fatal(err)
}
fmt.Println(res.Content)

The result has content, usage (input and output tokens), the model that actually ran, the requestedModel you asked for, any toolCalls, and vibePolicy when a policy tier fired.

Vibe SDK

Streaming

stream works like chat, with the same check and the same denial behavior. Text arrives as it is generated, and usage is settled when the stream closes. You must read the stream to the end for that to happen.

Node.js / TypeScript

TypeScript
const stream = await usage.stream({ identity: 'cust_acme', workflow: 'support-agent', provider: 'openai', model: 'gpt-4o-mini', messages });

for await (const chunk of stream.textStream) {
  process.stdout.write(chunk);
}
const final = await stream.finalResult; // resolves after usage is settled

Python

Python
stream = client.stream(identity="cust_acme", model="gpt-4o-mini",
                       messages=[Message("user", "Tell me a story")])
for chunk in stream:
    print(chunk, end="", flush=True)
result = stream.result()  # usage is recorded when the stream ends

Go

Go
s, err := client.Stream(ctx, req)
if err != nil {
	return err
}
for text := range s.Text {
	fmt.Print(text)
}
res, err := s.Result() // available after usage is settled
Vibe SDK

Embeddings

embed creates embeddings with OpenAI. It is checked and metered like chat, but on input tokens only, since embeddings have no output. Embeddings are OpenAI-only.

Node.js / TypeScript

TypeScript
const result = await usage.embed({
  identity: 'cust_acme',
  model: 'text-embedding-3-small',
  input: ['first text', 'second text'],
});
result.embeddings; // one vector per input

Python

Python
result = client.embed(identity="cust_acme", model="text-embedding-3-small", input=["hello"])

Go

Go
res, err := client.Embed(ctx, vibe.EmbedRequest{
	Identity: "cust_acme",
	Model:    "text-embedding-3-small",
	Input:    []string{"first text", "second text"},
})
// res.Embeddings: one vector per input
Vibe SDK

Policy Decisions

When a policy tier fires, the SDK acts on it and tells you what happened, so you never have to guess why a call ran the way it did.

What you see on the result

  • •Route to a model: the call runs on the model the policy named, even on a different provider. model and provider show what ran, and requestedModel shows what you asked for.
  • •Any fired tier: vibePolicy holds the policy ID, the tier index and the action. Log it if you want an audit of policy decisions in your own systems.
  • •Degrade: the original model is used untouched. Read vibePolicy to apply the degrade yourself.
  • •Block: there is no result. You get a rejection error instead (next section).
TypeScript
const result = await usage.chat({ identity: 'cust_acme', workflow: 'downgrade-openai-models', provider: 'anthropic', model: 'claude-sonnet-5', messages });

if (result.vibePolicy) {
  console.log(`Policy ${result.vibePolicy.policyId}, tier ${result.vibePolicy.tierIndex} fired`);
}
// If the policy rerouted the call, result.provider / result.model already reflect it.
Vibe SDK

Withdraw & Credit

Not all usage is an AI call: a nightly batch job, a correction, a fixed fee. withdraw takes an amount off a customer's allowance directly, and credit reverses one. They go through the same check as AI calls (in Go you can also name a Workflow so a policy applies). Each needs an idempotencyKey that identifies the adjustment.

Node.js / TypeScript

TypeScript
await usage.withdraw({ identity: 'cust_acme', amount: 500, idempotencyKey: 'job-42-attempt-1', reason: 'nightly batch reconciliation' });

await usage.credit({ identity: 'cust_acme', amount: 500, idempotencyKey: 'job-42-reversal-1', reason: 'correct over-count from job-42' });

Go

Go
res, err := client.Withdraw(ctx, vibe.WithdrawRequest{
	Identity: "cust_acme", Amount: 500, IdempotencyKey: "job-42-attempt-1",
	Reason: "nightly batch reconciliation", Workflow: "support-agent", // Workflow is optional
})

Reserve now, settle later (Go)

When the final amount isn't known up front, the Go SDK can hold the amount first and settle it when the work is done. WithdrawAsync reserves and returns a CaptureID; Close settles it, by default for the reserved amount or for a smaller final one. A capture can only be closed once. This is available in the Go SDK only.

Go
capture, err := client.WithdrawAsync(ctx, vibe.WithdrawRequest{Identity: "cust_acme", Amount: 100, IdempotencyKey: "job-43"})
// ... do the work ...
final := 60.0
_, err = client.Close(ctx, vibe.CloseRequest{CaptureID: capture.CaptureID, Amount: &final})

⚠️ Current limits

idempotencyKey is recorded for your audit trail, but a repeated call with the same key is not yet de-duplicated, so don't retry a withdraw blindly. credit sends the amount as a negative adjustment; confirm it behaves as you expect for your plan before you rely on it in production.

Vibe SDK

Errors & Blocked Calls

When a policy blocks a call, or the customer is out of allowance, the SDK raises a rejection error. This is the normal way a limit shows up in your code, so handle it explicitly and tell your user.

Node.js / TypeScript

TypeScript
import { usage, VibeRejectionError } from '@usageflow/vibe';

try {
  const result = await usage.chat({ /* ... */ });
  res.json({ answer: result.content });
} catch (err) {
  if (err instanceof VibeRejectionError) {
    // UsageFlow denied the call — the provider was never contacted.
    res.status(429).json({ error: 'Usage limit reached', reason: err.reason });
  } else {
    throw err;
  }
}

Go

Go
res, err := client.Chat(ctx, req)
var rejected *vibe.RejectionError
if errors.As(err, &rejected) {
	// UsageFlow denied the call — the provider was never contacted.
	http.Error(w, "Usage limit reached: "+rejected.Message, http.StatusTooManyRequests)
	return
}
if err != nil {
	return err
}
What happenedWhat you get
A policy blocks the call, or usage is over the limitA rejection error. The provider was not called.
UsageFlow can't be reachedAn ordinary error. The provider was not called.
The provider returns an errorThat error, unchanged.
A required provider key is missingAn error naming the key to set.
Vibe SDK

Reference

Configuration

SettingNode.js optionGo optionEnvironment variable
UsageFlow API key (required)apiKeyAPIKeyUSAGEFLOW_API_KEY
Anthropic keyanthropicApiKeyAnthropicAPIKeyANTHROPIC_API_KEY
OpenAI keyopenaiApiKeyOpenAIAPIKeyOPENAI_API_KEY

What each SDK supports

CallNode.jsGo
Chat✓✓
Streaming✓✓
Embeddings (OpenAI)✓✓
Withdraw / credit✓✓
Reserve now, settle later—✓
Image, speech, transcription, moderation, batch✓ (see below)—

The two SDKs are not identical. Check this table before you plan around a feature.

Also in the Node.js SDK (early)

  • •image, speak and transcribe are OpenAI-only. Their usage is counted with a flat placeholder weight, not the provider's real pricing, so don't use them for billing yet.
  • •moderate is a passthrough to OpenAI's moderation and is not metered.
  • •batchSubmit, batchStatus and batchSettle run a provider batch job. Settling a batch must happen in the same running process that submitted it.
Console

Where to See It in the Console

PageWhat it shows
DashboardLive decisions as calls happen, how requests were changed by your policies (requested versus executed model), and your blocked requests.
UsageEvery call: the customer, what was requested and what ran, why (which policy fired), and the impact. You can also see policy decisions, aggregates, top spenders, models and identities. Withdraw and credit adjustments appear as local ledger rows, in credits rather than a model.
Management → PoliciesCreate and edit your policies (workflows): tiers, branches, effects and renewal, and each workflow's ID.
ManagementYour applications and API keys, your customers (identities), and their metadata.

🧭 Not seeing a policy fire?

Check the workflow first. Make sure the workflow in your code exactly matches the policy's workflow ID on the Policies page, and that the policy is Active. A call with no matching policy is not an error; it just runs without one.

Framework Packages

Separate from the Vibe SDK, UsageFlow publishes packages that meter the HTTP routes of your own API: Express, Fastify and NestJS for Node.js, FastAPI and Flask for Python, and Gin for Go. Install the one for your runtime below.

Node.js Frameworks

Install the UsageFlow package for your Node.js framework:

Bash
npm install @usageflow/express
npm install @usageflow/fastify
npm install @usageflow/nestjs

Express

JavaScript
import express from "express";
import { ExpressUsageFlowAPI } from "@usageflow/express";

const apiKey = process.env.USAGEFLOW_API_KEY;
if (!apiKey) throw new Error("Missing USAGEFLOW_API_KEY");

const app = express();
app.use(express.json());

const usageFlow = new ExpressUsageFlowAPI(apiKey);
app.use(usageFlow.createMiddleware());

app.get("/api/users", (_req, res) => res.json([]));

Fastify

JavaScript
import Fastify from "fastify";
import { FastifyUsageFlowAPI } from "@usageflow/fastify";

const apiKey = process.env.USAGEFLOW_API_KEY;
if (!apiKey) throw new Error("Missing USAGEFLOW_API_KEY");

const app = Fastify();
const usageFlow = new FastifyUsageFlowAPI(apiKey);
await app.register(usageFlow.createPlugin());

app.get("/api/users", async () => []);

NestJS

TypeScript
import { Module } from "@nestjs/common";
import { UsageFlowModule } from "@usageflow/nestjs";

const apiKey = process.env.USAGEFLOW_API_KEY;
if (!apiKey) throw new Error("Missing USAGEFLOW_API_KEY");

@Module({
  imports: [UsageFlowModule.forRoot({ apiKey })],
})
export class AppModule {}

Current limitation

Current limitation: module registration alone does not start request processing reliably, so NestJS request traces are not currently supported. Use the Express or Fastify package directly where possible.

Python Frameworks

Install the UsageFlow package for your Python framework:

Bash
pip install usageflow-flask
pip install usageflow-fastapi uvicorn

FastAPI

Python
import os
from fastapi import FastAPI
from usageflow.fastapi import UsageFlowMiddleware

app = FastAPI()
app.add_middleware(
    UsageFlowMiddleware,
    api_key=os.environ["USAGEFLOW_API_KEY"],
)

@app.get("/api/users")
def users():
    return []

FastAPI production limitation

Start the app with `uvicorn app:app`, send the test request, then confirm the Trace in Console. Do not enable this middleware on routes that depend on FastAPI or Starlette response background tasks; the current middleware replaces an existing response background task.

Flask

Python
import os
from flask import Flask
from usageflow.flask import UsageFlowMiddleware

app = Flask(__name__)
UsageFlowMiddleware(app, api_key=os.environ["USAGEFLOW_API_KEY"])

@app.get("/api/users")
def users():
    return []

Go

Install the UsageFlow middleware package for Go (Gin framework):

Bash
go get github.com/usageflow/usageflow-go-middleware/v2

Gin v2

The Go middleware provides request interception, usage tracking, user identification, and automatic configuration updates with graceful degradation.

Go
package main

import (
    "log"
    "os"

    "github.com/gin-gonic/gin"
    "github.com/usageflow/usageflow-go-middleware/v2/pkg/middleware"
)

func main() {
    apiKey := os.Getenv("USAGEFLOW_API_KEY")
    if apiKey == "" {
        log.Fatal("Missing USAGEFLOW_API_KEY")
    }

    router := gin.New()
    usageFlow := middleware.New(apiKey)
    router.Use(usageFlow.RequestInterceptor())

    router.GET("/api/users", func(c *gin.Context) {
        c.JSON(200, []string{})
    })

    if err := router.Run(":8080"); err != nil {
        log.Fatal(err)
    }
}

Migrating from the legacy Gin package

github.com/usageflow/usageflow-gin is deprecated and does not support Console-managed routes. Replace it with the v2 package above, register RequestInterceptor() without local route lists, then configure routes in Console.

Best Practices

API Key Security

Never commit API keys to version control, and keep your UsageFlow and provider keys on the server, never in browser code. Always use environment variables.

Resources & Support

Find packages, examples, and get help:

Need help? Contact our support team: