Four prices, in the order they will bite. Every one has a lever next to it, and every one fails open: when a limit is reached, events flow through unjudged and nothing in the request path waits.
| At the default | The lever | |
|---|---|---|
| Money | $0.042 per million input tokens, about $6 per 100k judged events | when, cacheKey, state |
| Rate | 1,200 calls per minute per account | budget.perMinute, cacheKey |
| Latency | up to 2 s before a keep event drains, client unaffected | timeoutMs, a tight when |
| Egress | the whole event, before redaction | state, providerOptions.gateway.zeroDataRetention |
Money
Jev bills input tokens only, at $0.042 per million at the time of writing. The fifteen requests on the overview cost $0.00007. A 1.5k-token event is about $0.00006, so 100k judged events are about $6, and a million are about $60. Prices move; the model page has the current figure.
What decides the bill is how many events reach the model, not how many signals you define. All signals due for one event share one call, so adding a question to a signal that already runs is close to free:
whenis the first lever.faultonstatus >= 400costs nothing on successful requests.cacheKeyis the second. An outage produces thousands of identical 502s; a key onpath + error.namemakes it one call and serves the rest from cache.stateis the third. The default sends the whole event. Picking the six fields the question needs cuts tokens and is the same change that fixes egress below.
Rate
The limit that binds before money is the account rate: 1,200 requests per minute per TypeSafe account, about 20 a second. budget.perMinute (default 600) is a fixed one-minute window per process, so with four instances on one account set it to 300. Past the window, events pass through unjudged and stats().skipped goes up.
A failed call opens a breaker for budget.cooldownMs (default 30 s), so a 429 or an outage at the model costs you a 30 s gap in verdicts, not a queue.
Latency
A keep signal runs before the sampling decision and waits for the verdict, up to timeoutMs (default 2 s). Two things make that acceptable:
- The response is already sent.
keepandenrichrun after the request finishes, so the client never waits on the model. durationMson the event is the request's duration, measured before keep hooks run. The judgment does not inflate your latency charts.
What it does delay is the event reaching your drain, by the model's round trip. Keep signals belong on narrow paths (/api/checkout, not every 200), and enrich-only signals have no such wait.
Egress
The model reads what state gives it. By default that is the whole wide event minus the signals columns, and for a keep signal it is the request context before redaction runs. If your handlers log anything you would not paste into a third-party API, set state to pick fields:
createSignals({
signals,
state: e => ({ status: e.status, path: e.path, durationMs: e.durationMs, error: e.error, payment: e.payment }),
providerOptions: { gateway: { zeroDataRetention: true } },
})
zeroDataRetention is passed through to AI Gateway untouched. Inference runs in your process with your key; evlog holds no key and proxies nothing.
What never happens
- A keep signal never drops an event. It promotes, and head sampling stays deterministic.
- A verdict is never prose. Columns are typed values with a probability, and nothing else is added to the event unless you ask for it (
stampModel). - A limit never surfaces as an error. Budget spent, breaker open, state too large, model down: the event drains as it would have without signals, and the counter in
stats()moves.
What You Can Ask
Three question shapes, each with the column it produces: a yes/no with a probability, one option out of a list, or a position on a rubric.
Reference
Every option of defineSignal and createSignals with its default, the column format, the stats() counters, and the test helpers in @evlog/signals/testing.