Reference
Everything configurable, on one page. The other pages link here instead of repeating it.
pnpm add @evlog/signals ai
ai 7.0.105 or later is a peer dependency. Entry points: @evlog/signals and @evlog/signals/testing.
Options
defineSignal(options) validates and returns a Signal. The shape of the answer follows the input: ask alone is yes/no, ask + choice picks an option, ask + score positions on a rubric.
| Option | Type | Notes |
|---|---|---|
name | string | Column name under event.signals. Letters, digits, dashes, starting with a letter. Unique across registered signals. |
ask | string | The question, in plain English. |
when | (event) => boolean | Predicate over the event. No model call when false. Required with keep. |
choice | Record<option, description> | At least two options. Makes a choice signal. |
score | [level, level, ...] | Ordered lowest first, at least two. Makes a score signal. |
criteria | { true?, false? } | Boolean signals only. What makes each answer true. |
keep | (verdict) => boolean | Promote the event past sampling when true. Never drops. Typed on the verdict the shape produces. |
cacheKey | (event) => string | undefined | Reuse the verdict for events sharing a key. undefined skips caching. Last 1,000 verdicts per process. |
defineSignal throws at definition time on an invalid name, an empty ask, keep without when, fewer than two options or levels.
createSignals(options) returns a SignalsPlugin: an evlog plugin with keep, enrich and stats(). The only required field is signals.
| Option | Default | Notes |
|---|---|---|
signals | The signals to run. Duplicate names throw at startup. | |
model | 'typesafe-ai/jev' | AI SDK evaluation model: a gateway id or a provider instance. |
budget.perMinute | 600 | Model calls per minute, shared by all signals, per process. |
budget.cooldownMs | 30000 | Pause after a failed call. |
timeoutMs | 2000 | Per call. An overrun aborts the call and counts as an error. |
state | whole event minus signals and signalsModel | What the model reads. May return a string. |
maxStateChars | 100000 | Largest state sent, in characters of JSON. Larger events are skipped. |
providerOptions | Forwarded to the model call, e.g. { gateway: { zeroDataRetention: true } }. | |
evaluate | AI SDK experimental_evaluate | Replace the model call. Tests, record and replay. |
stampModel | false | Also write the answering model's id on event.signalsModel. Set it while comparing two models so their verdicts can be told apart in the drain; one more column on every judged event otherwise. |
model takes any AI SDK evaluation model. See Pick the model.
SignalsPlugin is a standard evlog plugin. Where it goes depends on how the framework exposes plugins:
| Framework | Register |
|---|---|
| Nuxt, Nitro | In a server plugin: nitroApp.hooks.hook('evlog:emit:keep', plugin.keep) and nitroApp.hooks.hook('evlog:enrich', plugin.enrich) |
| Next.js, standalone | initLogger({ plugins: [plugin] }) |
| Hono, Express, Fastify, Elysia, oRPC, SvelteKit, React Router | The middleware's plugins option: evlog({ plugins: [plugin] }) |
createMiddlewareLogger (toolkit) | plugins: [plugin] on the call |
Pick the model
The default goes through AI Gateway: one key, every decision model, swap by changing a string. A direct provider is the same option with an instance instead of a string.
The default. Reads AI_GATEWAY_API_KEY, or OIDC on Vercel with no key at all.
createSignals({
model: 'typesafe-ai/jev', // the default; 'liquid/d1' is the other decision model
signals: [fault, silentFailure],
})
Both gateway models return a probability distribution, so every column has a confidence. The id string is the only thing to change to compare them; set stampModel: true while you do, so verdicts from each model can be told apart in the drain.
A direct key to the model behind the default. Install @ai-sdk/typesafe-ai; it reads TYPESAFE_AI_API_KEY.
import { typeSafeAi } from '@ai-sdk/typesafe-ai'
createSignals({
model: typeSafeAi.evaluationModel('jev-latest'),
signals: [fault, silentFailure],
})
Same distributions, same columns as through the gateway.
A general model prompted to answer. Install @ai-sdk/openai; it reads OPENAI_API_KEY.
import { openai } from '@ai-sdk/openai'
createSignals({
model: openai.evaluationModel('gpt-6-luna'),
signals: [fault, silentFailure],
})
A prompted model returns P(true) for yes/no questions and a bare answer for choice and score, so those columns carry value without confidence. Expect a slower call than a decision model; raise timeoutMs if stats().errors climbs.
Any other provider that exposes evaluationModel() plugs in the same way. Options the provider understands go through providerOptions, keyed by provider name.
The column
Verdicts land on event.signals.<name>. With stampModel: true, the answering model's id is on event.signalsModel.
| Shape | value | confidence | Extra |
|---|---|---|---|
| Boolean | boolean | probability of value | |
| Choice | one of the option names | probability of that option, when the model returns a distribution | |
| Score | one of the level names (most likely) | probability of that level, when the model returns a distribution | score: number, probability-weighted position from 0 to levels - 1 |
A signal that promoted the event past sampling adds kept: true to its column.
interface BooleanVerdict { value: boolean; confidence: number }
interface ChoiceVerdict<Option extends string> { value: Option; confidence?: number }
interface ScoreVerdict<Level extends string> { value: Level; score: number; confidence?: number }
type SignalColumn = Verdict & { kept?: true }
Columns are written in enrich, before the console line. They reach stdout, platform logs such as Vercel, and every drain. The response is already sent by then; what waits for the evaluation is the console line, bounded by timeoutMs, so a late line in the dev terminal is the model, not a slow request.
stats()
plugin.stats() returns a snapshot of per-process counters:
| Field | Counts |
|---|---|
calls | Model calls made. |
skipped | Events with due signals not judged: budget spent, breaker open, or state over maxStateChars. |
errors | Calls that failed or timed out. |
cached | Verdicts served from a cacheKey without a call. |
inputTokens | Tokens billed, as reported by the model. |
cached next to calls is what cacheKey saved. skipped climbing while errors stays flat means the budget is below your traffic, not that the model is down.
Testing
@evlog/signals/testing replaces the model with a script. No key, no network, no budget:
import { createSignals } from '@evlog/signals'
import { answers, scriptedEvaluate } from '@evlog/signals/testing'
const mock = scriptedEvaluate((name, question) => {
if (name === 'fault') return answers.choice(question, 'upstream', 0.93)
return answers.first(question)
})
const plugin = createSignals({ signals, evaluate: mock.evaluate })
// after a request has run through evlog
expect(mock.calls).toHaveLength(1)
expect(mock.calls[0].questions).toHaveProperty('fault')
| Export | Does |
|---|---|
scriptedEvaluate(answer, modelId?) | Returns { evaluate, calls }. answer(name, question, state) runs once per question; calls records every request. |
answers.boolean(p) | A yes/no answer with P(true) = p. |
answers.choice(question, pick, p) | Picks pick at p, spreads the rest over the other options. |
answers.score(question, levelIndex, p) | Level levelIndex at p, with the weighted score computed. |
answers.first(question) | A confident default of the right type: first option at 0.9, last level at 0.8, yes at 0.94. |
The same hook replays recorded verdicts: a scriptedEvaluate that looks answers up by requestId runs your signals against yesterday's events without a call.