Signals
Every option of defineSignal and createSignals with its default, the column format, the stats() counters, and the test helpers in @evlog/signals/testing.

Everything configurable, on one page. The other pages link here instead of repeating it.

Terminal
pnpm add @evlog/signals ai

ai 7.0.105 or later is a peer dependency. Entry points: @evlog/signals and @evlog/signals/testing.

Options

defineSignal(options) validates and returns a Signal. The shape of the answer follows the input: ask alone is yes/no, ask + choice picks an option, ask + score positions on a rubric.

OptionTypeNotes
namestringColumn name under event.signals. Letters, digits, dashes, starting with a letter. Unique across registered signals.
askstringThe question, in plain English.
when(event) => booleanPredicate over the event. No model call when false. Required with keep.
choiceRecord<option, description>At least two options. Makes a choice signal.
score[level, level, ...]Ordered lowest first, at least two. Makes a score signal.
criteria{ true?, false? }Boolean signals only. What makes each answer true.
keep(verdict) => booleanPromote the event past sampling when true. Never drops. Typed on the verdict the shape produces.
cacheKey(event) => string | undefinedReuse the verdict for events sharing a key. undefined skips caching. Last 1,000 verdicts per process.

defineSignal throws at definition time on an invalid name, an empty ask, keep without when, fewer than two options or levels.

Pick the model

The default goes through AI Gateway: one key, every decision model, swap by changing a string. A direct provider is the same option with an instance instead of a string.

The default. Reads AI_GATEWAY_API_KEY, or OIDC on Vercel with no key at all.

createSignals({
  model: 'typesafe-ai/jev', // the default; 'liquid/d1' is the other decision model
  signals: [fault, silentFailure],
})

Both gateway models return a probability distribution, so every column has a confidence. The id string is the only thing to change to compare them; set stampModel: true while you do, so verdicts from each model can be told apart in the drain.

Any other provider that exposes evaluationModel() plugs in the same way. Options the provider understands go through providerOptions, keyed by provider name.

The column

Verdicts land on event.signals.<name>. With stampModel: true, the answering model's id is on event.signalsModel.

ShapevalueconfidenceExtra
Booleanbooleanprobability of value
Choiceone of the option namesprobability of that option, when the model returns a distribution
Scoreone of the level names (most likely)probability of that level, when the model returns a distributionscore: number, probability-weighted position from 0 to levels - 1

A signal that promoted the event past sampling adds kept: true to its column.

interface BooleanVerdict { value: boolean; confidence: number }
interface ChoiceVerdict<Option extends string> { value: Option; confidence?: number }
interface ScoreVerdict<Level extends string> { value: Level; score: number; confidence?: number }
type SignalColumn = Verdict & { kept?: true }

Columns are written in enrich, before the console line. They reach stdout, platform logs such as Vercel, and every drain. The response is already sent by then; what waits for the evaluation is the console line, bounded by timeoutMs, so a late line in the dev terminal is the model, not a slow request.

stats()

plugin.stats() returns a snapshot of per-process counters:

FieldCounts
callsModel calls made.
skippedEvents with due signals not judged: budget spent, breaker open, or state over maxStateChars.
errorsCalls that failed or timed out.
cachedVerdicts served from a cacheKey without a call.
inputTokensTokens billed, as reported by the model.

cached next to calls is what cacheKey saved. skipped climbing while errors stays flat means the budget is below your traffic, not that the model is down.

Testing

@evlog/signals/testing replaces the model with a script. No key, no network, no budget:

signals.test.ts
import { createSignals } from '@evlog/signals'
import { answers, scriptedEvaluate } from '@evlog/signals/testing'

const mock = scriptedEvaluate((name, question) => {
  if (name === 'fault') return answers.choice(question, 'upstream', 0.93)
  return answers.first(question)
})

const plugin = createSignals({ signals, evaluate: mock.evaluate })

// after a request has run through evlog
expect(mock.calls).toHaveLength(1)
expect(mock.calls[0].questions).toHaveProperty('fault')
ExportDoes
scriptedEvaluate(answer, modelId?)Returns { evaluate, calls }. answer(name, question, state) runs once per question; calls records every request.
answers.boolean(p)A yes/no answer with P(true) = p.
answers.choice(question, pick, p)Picks pick at p, spreads the rest over the other options.
answers.score(question, levelIndex, p)Level levelIndex at p, with the weighted score computed.
answers.first(question)A confident default of the right type: first option at 0.9, last level at 0.8, yes at 0.94.

The same hook replays recorded verdicts: a scriptedEvaluate that looks answers up by requestId runs your signals against yesterday's events without a call.