Observability
llm-ports exposes two coexisting observability surfaces. Both are opt-in; both fire from the same runtime paths. Pick whichever fits your consumer, or use them together.
The observability contract surface (alpha.28+, primary). Structured, versioned event stream defined by
@llm-ports/observability-contract. Events flow through anObservabilitySink { emit(event) }interface. Every event carries a full envelope (spec_version,event_id, timestamps,source,operation_id,attempt_id, correlation, W3C Trace Context). Lifecycle events at operation and attempt granularity, plusretry_scheduled/fallback.selected/evaluation.recorded. Registry-driven emission lives atRegistryOptions.instrumentation.The alpha.21 fire-and-forget hooks surface (still supported). Six typed callbacks on
RegistryOptions.observability:onComplete,onCost,onTokenUsage,onFallback,onValidationRetryandonCacheHit.onCompletecovers every operation the port routes as of0.1.0-alpha.35.1; inalpha.35it covered two, so upgrade before relying on it. Simple, low-ceremony, aligned with OpenTelemetrygen_ai.*semconv naming where applicable. Suitable when you only need "did the call succeed and what did it cost."
If you're new, start with the contract surface — it's the direction the library is heading, and the contract package can also be used by non-port callers (e.g. your own retry loops, subprocess-driven agents). The alpha.21 hooks stay stable for existing consumers and receive no deprecation timeline in this alpha line.
Contract surface (alpha.28 + alpha.29)
The contract ships as a standalone package with zero runtime dependency on @llm-ports/core. Any caller can construct conformant events using the package's buildEvent / emitLifecycleEvent helpers and forward them to any ObservabilitySink. The Registry uses this same shape when its instrumentation handle is configured.
Wiring the Registry (alpha.29)
import { createRegistryFromEnv } from "@llm-ports/core";
import { createCollectingSink } from "@llm-ports/observability-contract";
const sink = createCollectingSink();
const registry = createRegistryFromEnv({
env: process.env as Record<string, string>,
adapters: { /* ... */ },
instrumentation: {
config: {
sink,
source: { library: "my-app", library_version: "1.0.0" },
},
// Optional: caller-plumbed correlation. When context.operation_id is set,
// withOperation reuses it instead of minting a fresh one — useful when
// an outer scope (e.g. a long-horizon agent run) wants to pin the
// operation_id across many Registry calls.
// context: { operation_id: "op-outer-123" },
//
// Optional: opt into prompt fingerprint compute at attempt.completed.
// Off by default per the contract's CapturePolicy default.
// fingerprint: { algorithm: "sha256", promptId: "triage-classifier" },
},
});With instrumentation supplied, every call on the Registry's LLMPort for generateText, generateStructured, and runAgent emits the full contract lifecycle:
llm.operation.started— before any provider attempt fires. Carriestask_type,method,provider_chain(the intended fallback chain).llm.attempt.started— before each provider attempt. Carriesprovider_alias,model_id,attempt_number(1-indexed),is_retry,is_fallback.llm.attempt.completed— on a successful attempt. Carriesusage,cost,latency_ms,final_model_id, optionalcache_stats,provider_response_id,request_fingerprint.llm.attempt.failed— on a failed attempt. Carrieserror: ErrorInfo(witherror_type,message,cause_category,retryable,fallback_worthy) andlatency_ms.llm.attempt.retry_scheduled— between a failed attempt and the next same-provider retry. Carriesretry_reason,backoff_ms,next_attempt_number. Registry-only.llm.fallback.selected— between a failed attempt and the next chain-provider try. Carriesfrom_provider_alias,to_provider_alias,cause: FallbackCause. Registry-only.llm.operation.completed— on success. Carriesaggregate_usage,aggregate_cost,attempts_made,final_provider_alias,total_duration_ms, optionalresult_summary.llm.operation.failed— when the whole chain failed. Carrieserror,attempts_made,providers_tried,total_duration_ms.llm.operation.cancelled— when anAbortErrorpropagated. Carriescancelled_at_attempt,providers_tried_before_cancel,total_duration_ms.
Every event across one operation shares a single operation_id. Every attempt within an operation gets its own attempt_id. Streaming methods (streamText, streamStructured) emit the same lifecycle since alpha.30, plus AttemptCompletedData.stream_stats (ttft_ms, chunk_count, inter-chunk p50/p99, termination) and optional per-chunk llm.stream.chunk events under CapturePolicy.stream_chunk_capture: "full".
Per-call operation_id precedence (alpha.31)
The operation_id a Registry method uses is picked from three levels of precedence, checked in order:
- Per-call context (alpha.31). Wrap the port with
withObservabilityContext(port, { operation_id })and invoke through the wrapped instance. That id lands on every emitted lifecycle event for that call. This is the level to use when a caller mints its own correlation id (e.g. a capability wrapper that also writes a quality-tracker row keyed on the same id). - Registry-level context (alpha.29). Set
Instrumentation.context.operation_idon the Registry once, and every call reuses it. This is the level to use when an outer scope wants many Registry calls stitched under a single long-horizon id. - Fresh mint (default).
newOperationId()runs at operation start. This is the pre-alpha.29 behavior and remains the default for uninstrumented consumers.
The three levels are independent: any of them can be present or absent, and the Registry falls through cleanly.
import {
createRegistryFromEnv,
withObservabilityContext,
} from "@llm-ports/core";
const registry = createRegistryFromEnv({ /* ... instrumentation.config as above ... */ });
const port = registry.getPort();
// A: fresh mint (default). Registry mints a new operation_id for this call.
await port.generateText({ taskType: "triage", prompt });
// B: Registry-level context. Every call through `port` reuses "op-outer-123".
// (configured in the Registry options via `instrumentation.context: { operation_id: "op-outer-123" }`)
// C: per-call context (alpha.31). This one call is pinned to "op-request-42",
// regardless of any Registry-level context.
const scoped = withObservabilityContext(port, { operation_id: "op-request-42" });
await scoped.generateText({ taskType: "triage", prompt });The precedence applies uniformly to all five port methods: generateText, generateStructured, runAgent, streamText, streamStructured. Adapter-side agent-loop events (agent.step.*, agent.tool.*) continue to correlate via resurrectOperationContext(this) and land under the same chosen operation_id.
The shared instrumentation service
The Registry uses helpers from @llm-ports/core's instrumentation.ts module. You can call the same helpers directly to instrument your own retry loops or subprocess-driven adapters:
import {
withOperation,
withAttempt,
type Instrumentation,
} from "@llm-ports/core";
const instrumentation: Instrumentation = {
config: { sink, source: { library: "my-loop", library_version: "0.1.0" } },
};
const result = await withOperation(
instrumentation,
{ taskType: "triage", method: "runAgent", providerChain: ["openai"] },
async (opCtx) => {
return withAttempt(
opCtx,
{ providerAlias: "openai", modelId: "gpt-4o" },
async () => {
const response = await myProviderCall();
return {
value: response,
usage: response.usage,
cost: response.cost,
modelId: response.model,
};
},
);
},
);withOperation owns the operation.started / .completed / .failed / .cancelled lifecycle (detecting AbortError → cancelled). withAttempt owns attempt.started / .completed / .failed and updates the operation-level counters so operation.completed fills in aggregate_usage, final_provider_alias, and friends correctly. Sink failures never break the primary path (sync throws swallowed; async rejections caught). When instrumentation is undefined, both wrappers no-op and just call the inner work.
For streaming, use the manual escape hatch startAttempt / completeAttempt / failAttempt.
Prompt fingerprint
Setting Instrumentation.fingerprint enables per-attempt request fingerprinting. The Registry computes a RequestFingerprint once per operation (before any attempt runs) and attaches it to every attempt.completed via the optional AttemptCompletedData.request_fingerprint field.
instrumentation: {
config: { sink, source },
fingerprint: {
algorithm: "sha256", // or "hmac-sha256" (requires hmacKey ≥16 UTF-8 bytes)
promptId: "triage-classifier",
promptVersion: "v3.2",
},
}Two calls with identical messages + sampling params produce identical message_hash and request_hash. Different content produces different hashes. Useful for drift detection, template-version tracking, or A/B analysis across attempts.
Persistent evaluation storage (@llm-ports/eval)
Evaluations arrive late — LLM-judge scores after the fact, human annotations hours later, dataset replays days later. The @llm-ports/eval package provides durable storage keyed on the contract's EvaluationRef shape.
Bridge the store to the Registry's sink:
import { createSqliteEvaluationStore, toObservabilitySink } from "@llm-ports/eval";
const store = createSqliteEvaluationStore({ dbPath: "./evaluations.db" });
const sink = toObservabilitySink(store);
const registry = createRegistryFromEnv({
env: process.env as Record<string, string>,
adapters: { /* ... */ },
instrumentation: {
config: { sink, source: { library: "my-app", library_version: "1.0.0" } },
},
});Only evaluation.recorded events land in the store; lifecycle events are silently ignored. For BOTH lifecycle capture AND evaluation storage, compose a fan-out sink yourself.
See docs/concepts/evaluations.md for the full write / query surface.
Shipped in alpha.30
Every alpha.29 carve-out landed:
- Streaming instrumentation.
streamText/streamStructurednow emit the full operation + attempt lifecycle plusAttemptCompletedData.stream_stats(ttft_ms,chunk_count, inter-chunkp50/p99,termination). Per-chunkllm.stream.chunkevents are opt-in viaCapturePolicy.stream_chunk_capture === "full". See the Streaming Observability guide. - Provider cache normalization. OpenAI (
cached_tokens), Anthropic (cache_read/create_input_tokens), and Google (cachedContentTokenCount) fold intoAttemptCompletedData.cache_stats.provider_cachewith a canonical status enum (hit/partial/miss/ omitted). See the Cache Control guide's alpha.30 section. - Agent step + tool events (
agent.step.*,agent.tool.*) — see the "Agent-loop events" section below. - Adapter-side emission — codex + aider adapters route through the shared
withOperation+withAttemptservice instead of hand-rollingsink.emit. Direct-adapter callers see the same lifecycle events the Registry produces. - OpenTelemetry semconv bridge — new companion package
@llm-ports/telemetry-otel.createOtelSink({ tracer, meter? })maps every contract event to OTel gen_ai spans + metrics.
TD-ALPHA29-ADAPTER-EMIT-DEFERRED and TD-ALPHA29-AGENT-STEP-EVENTS-DEFERRED both close on alpha.30.
Agent-loop events (alpha.30+)
The three in-process adapters that own their runAgent tool-use loop (openai, anthropic, google) emit correlated per-step + per-tool events for every runAgent call, threaded through resurrectOperationContext(this) inside the adapter method.
Per LLM turn:
agent.step.started—step_index(1-based),step_type: "llm"(or"tool"/"validation"on adapters that later add those step kinds).agent.step.completed—duration_ms,usage,cost.
Per tool call:
agent.tool.called—tool_name,tool_call_id,arguments_digest: sha256Hex(rawArgs). Content-free by default.agent.tool.returned—result_digest,duration_ms, optionalerror: ErrorInfowhen the tool threw. Content-free by default.
All events share the outer operation_id, so an aggregating sink sees the full step + tool tree stitched under one span (or one OTel span, when wired through the telemetry-otel bridge).
Direct-adapter callers. The same events fire on direct-adapter calls (bypassing the Registry) when the caller has wrapped the port with withObservabilityContext(port, ctx) and set ObservabilityContext.operation_handle to a running operation. When no outer scope has opened one, resurrectOperationContext(this) returns undefined and the four emit helpers no-op — safe by default.
withObservabilityContext binding change (alpha.30)
withObservabilityContext now binds methods to the wrapped proxy (receiver) instead of the underlying target. This is what enables resurrectOperationContext(this) inside adapter runAgent methods — this resolves to the proxy, which is the instance the context is registered against.
Non-breaking: destructured calls (const { generateText } = wrappedPort; generateText(...)) still work because the receiver at bind time is the proxy. Every existing test continues to pass. The change only opens the door for adapters to resurrect the outer op context via this.
Capture policy — responsePreviewMaxChars + stream_chunk_capture
Two CapturePolicy fields exercised by alpha.30 additions:
responsePreviewMaxChars: number(default0in strict,200in permissive). Gates theresponse_previewfield onllm.attempt.completed. Preview lands only whencontent === "full" | "redacted"ANDresponsePreviewMaxChars > 0.response_char_count(the count, not the content) always emits regardless.stream_chunk_capture: "off" | "sampled" | "full"(default"off"). Gates per-chunkllm.stream.chunkevents on streamed methods."full"fires one event per chunk (withchunk_contentfurther gated by the content policy);"sampled"is reserved for future rate-limited emission;"off"yields aggregatestream_statsonly.
Both defaults keep the strict-by-default posture. Opt in via Instrumentation.capturePolicy at Registry setup.
Alpha.21 fire-and-forget hooks
llm-ports exposes five fire-and-forget observability hooks on RegistryOptions.observability (alpha.21+). Event shapes align with the OpenTelemetry gen_ai.* semantic conventions so downstream pipelines (Honeycomb, Datadog, OTel Collector, custom OTLP exporters) can map them onto spans and metrics without re-deriving fields.
The hooks complement the existing per-adapter onRetry hook (alpha.17+). onRetry covers "the adapter decided to retry an in-flight request"; the Registry-level hooks below cover "the Registry decided to move on" and "a successful call is interesting to observe."
Both the alpha.21 hooks and the alpha.28/.29 contract surface can coexist on the same Registry — they observe overlapping runtime paths and are independent choices about how to receive the data.
Quick start
import { createRegistryFromEnv } from "@llm-ports/core";
const registry = createRegistryFromEnv({
env: process.env,
adapters: { /* ... */ },
observability: {
onComplete: (e) => myMetrics.calls.observe({ ok: e.ok, attempts: e.providerAttempts }),
onCost: (e) => myMetrics.cost.observe(e.totalUsd, { model: e.modelId }),
onTokenUsage: (e) => myMetrics.tokens.observe(e.totalTokens, { model: e.modelId }),
onFallback: (e) => myLogger.warn(`fallback ${e.fromAlias} -> ${e.toAlias} (${e.cause})`),
onCacheHit: (e) => myMetrics.cacheHitRatio.observe(e.hitRatio),
onValidationRetry: (e) => myLogger.warn(`validation retry ${e.attempt}/${e.maxAttempts}`),
},
});All six fields are independently optional. Pass only the hooks the downstream pipeline needs.
Hook reference
onComplete
Added in 0.1.0-alpha.35.
Complete since 0.1.0-alpha.35.1. Not complete in alpha.35.
alpha.35 emitted this from generateText and generateChat only, while the operation type named nine, so a consumer adopting it as their single spend event lost every streamed and structured call. If you are on alpha.35, upgrade before relying on it. Fixed in alpha.35.1, which fires for all seven operations the port routes. TD-LLMPORTS-ONCOMPLETE-FIRES-FOR-TWO-OF-NINE-OPERATIONS.
Fires once per call, on success and on failure, which is what makes it different from every other hook here: the rest describe something that happened during a call, and this one describes the call.
observability: {
onComplete: (e) => {
myMetrics.calls.inc({ ok: String(e.ok), task: e.taskType });
myMetrics.latency.observe(e.latencyMs);
if (e.totalUsd !== undefined) myMetrics.spend.observe(e.totalUsd);
if (!e.ok) myLogger.warn({ err: e.error, attempts: e.providerAttempts }, "llm call failed");
},
}| Field | Meaning |
|---|---|
ok | Whether the call returned a result. false means every provider in the chain failed. |
operation | Which port method ran: generateText, generateStructured, generateChat, runAgent, and so on. |
taskType | The task the caller asked for, which is what a route was chosen by. |
providerAttempts | How many providers were tried, not how many answered. A value of 3 on a successful call means two failed over before one worked. |
providerAlias / modelId | The provider and model that answered, or the last one attempted when the call failed. Always present. |
latencyMs | Wall-clock time for the whole call, failover included. |
usage / totalUsd | Absent rather than zero when the call produced neither, so a spend total cannot be inflated by failures or by a model with no known price. |
validationAttempts | Validation rounds, on the structured methods only. 1 means the first response validated. |
taskType / budgetScope / refs | The task a route was chosen by, the scope any spend was counted against, and any artifact references attached to the call. |
error | The error that ended the call, present only when ok is false. |
What "once per call" means for a streamed call. A stream has no single completion instant, so the event fires when the stream is exhausted. On the failure side it fires for an abort or an error, whether that happens before the first chunk or midway through after chunks were already delivered, and usage and cost carry whatever accumulated before the failure. This differs from onStreamComplete on purpose: that one fires only on natural completion and stays that way, while this one promises the failure case.
Why the failure case matters. A completion hook that fired only on success would describe a healthier system than the real one, and reliability is usually the reason to switch it on. It fires for both.
Why providerAttempts is the interesting field. Answering "what did this call cost and how many attempts did it take" previously meant correlating onCost with onFallback and keeping a count yourself. This is that answer in one event.
onCost
Fires after every billable call against generateText, generateStructured, or runAgent. Cost is in USD, broken down by prompt/completion/cache.
interface CostEvent {
promptUsd: number;
completionUsd: number;
totalUsd: number;
cacheReadUsd?: number; // when the provider has a discounted cache-read tier
cacheWriteUsd?: number; // Anthropic-style explicit-cache providers
reasoningUsd?: number; // hidden chain-of-thought billed separately
modelId: string;
providerAlias: string;
operation: "generateText" | "generateStructured" | "streamText" | "streamStructured" | "runAgent" | "embed" | "rerank";
taskType?: string;
budgetScope?: { scope: string; scopeId: string };
}onTokenUsage
Fires alongside onCost with raw token counts (before USD monetization). Useful when downstream metrics care about token volume independent of pricing changes.
interface TokenUsageEvent {
inputTokens: number;
outputTokens: number;
cachedInputTokens?: number;
cacheCreationTokens?: number;
reasoningTokens?: number;
totalTokens: number;
modelId: string;
providerAlias: string;
operation: CostEvent["operation"];
taskType?: string;
budgetScope?: { scope: string; scopeId: string };
}onFallback
Fires when the Registry's provider chain advances. Per-call only; not emitted for the initial selection or for forceProviderAlias calls (which by contract don't fall back).
type FallbackCause =
| "provider-error" // the primary raised an error (budget, 401, 5xx, transient)
| "budget-exhausted" // gate denied the call on the primary
| "validation-exhausted" // structured retries gave up; chain advances
| "empty-response" // primary returned empty after starvation retries
| "circuit-open";
interface FallbackEvent {
fromAlias: string;
toAlias: string;
cause: FallbackCause;
operation: CostEvent["operation"];
taskType?: string;
reason?: unknown; // the error or signal that triggered the advancement
}In alpha.21, only cause: "provider-error" is emitted. Future releases will surface the other causes as the Registry learns to advance on more signals.
onCacheHit
Fires when the response reports cacheReadTokens > 0. Useful for tracking whether prompt-engineering caching is actually firing.
interface CacheHitEvent {
cachedTokens: number;
inputTokensTotal: number;
hitRatio: number; // cachedTokens / inputTokensTotal
savingsUsd?: number; // populated when the provider has a discounted cache-read tier
modelId: string;
providerAlias: string;
operation: CostEvent["operation"];
taskType?: string;
}onValidationRetry
Type-only in alpha.21. Registry-level emission is the alpha.22 follow-up. Consumers wanting validation-retry observability today should use the adapter's existing onRetry hook and filter on reason === "validation-feedback":
const adapter = createOpenAIAdapter({
apiKey: process.env.OPENAI_API_KEY!,
onRetry: (e) => {
if (e.reason === "validation-feedback") {
myMetrics.validationRetries.inc({ model: e.modelId });
}
},
});When emission lands, the event shape will be:
type ValidationRetryCause =
| "schema-mismatch" // valid JSON that failed Zod validation
| "parse-error"; // non-JSON text response
interface ValidationRetryEvent {
attempt: number;
maxAttempts: number;
modelId: string;
providerAlias: string;
cause: ValidationRetryCause;
issues?: unknown;
operation: "generateStructured" | "streamStructured";
}Coverage (as of alpha.25)
| Hook | Emitted? | Where it fires |
|---|---|---|
onCost | ✅ | Every successful generateText / generateStructured / runAgent and at natural stream completion for streamText / streamStructured (alpha.25+) |
onTokenUsage | ✅ | Same as onCost |
onCacheHit | ✅ | When the response reports cacheReadTokens > 0 |
onFallback | ✅ (cause: provider-error) | When the Registry's walkChain advances from one alias to the next |
onValidationRetry | ✅ (alpha.24+) | Registry-level emission via deriveValidationRetryFromAdapterRetry bridging the adapter's onRetry |
Streamed cost surfacing (alpha.25) emits onCost + onTokenUsage once per stream at natural completion; mid-stream errors and consumer-cancelled streams do NOT emit (matches the non-streaming "cost recorded only on success" contract). Enabled by default in adapter-openai via stream_options: { include_usage: true }; opt out per adapter with streamUsage: false if a compat provider rejects the field.
refs field on events (alpha.25+)
Every observability event carries an optional refs?: Record<string, ArtifactRef> field, threaded verbatim from the call options. Use this to attribute events to versioned artifacts (prompts, scaffolds, policies, experiment variants) or attribution tags (tenant, project, session):
const result = await port.generateStructured({
taskType: "extract",
prompt: userInput,
schema: MySchema,
refs: {
prompt: { key: "extractor-v3", version: 3, hash: "sha256:..." },
tenant: { key: "acme-corp" },
session: { key: "sess-abc123" },
},
});
// ...on the observability side:
const registry = createRegistryFromEnv({
observability: {
onCost: (e) => {
metrics.costByPrompt.record(e.totalUsd, {
promptKey: e.refs?.prompt?.key,
promptVersion: e.refs?.prompt?.version,
tenant: e.refs?.tenant?.key,
});
},
},
});Non-goals: refs are not validated, not sent to the model, not persisted, not read by adapters. Pure consumer-owned trace metadata. See the alpha.24-to-alpha.25 migration guide for the full contract.
Aggressive fallback preset (alpha.25+)
RegistryOptions.runtimeFallback: "aggressive" bundles the opinionated classifier three consumers rebuilt by hand (BEPA Plan 29, HomeSignal, SalesCoach Plan 30). Beyond what the unconfigured registry walks on (provider outages including HTTP 5xx, timeouts, empty responses, unsupported content blocks, and a credential that has never worked), it also walks on RateLimitError, ContextWindowExceededError, and a BadRequestError whose message indicates exhausted credit. It does not walk on a credential that worked and then failed, on any other BadRequestError, or on budget exhaustion. The multi-provider guide compares all three policies side by side.
import { createRegistryFromEnv } from "@llm-ports/core";
const registry = createRegistryFromEnv({
adapters: { /* ... */ },
runtimeFallback: "aggressive", // alpha.25+, LP-REQ-01
observability: {
onFallback: (e) => {
// Fires when a rate-limit / credit-exhaustion / empty-response / 5xx
// caused the chain to walk. cause: "provider-error"; reason: the error.
metrics.chainWalks.inc({ cause: e.cause, from: e.fromAlias, to: e.toAlias });
},
},
});For finer control the object form still wins; the classifier is exported so consumers can compose:
import { aggressiveShouldFallback } from "@llm-ports/core";
runtimeFallback: {
shouldFallback: (e) => aggressiveShouldFallback(e) || myCustomCondition(e),
},Per-attempt timeout (alpha.23+)
Independent of the hooks above but related — RegistryOptions.perAttemptTimeoutMs wraps every provider attempt inside walkChain with an AbortController + timer:
const registry = createRegistryFromEnv({
// ...existing options...
perAttemptTimeoutMs: 30000, // 30s cap per provider attempt
});On timeout, the abort propagates to the adapter's HTTP client; the adapter throws ProviderUnavailableError; the Registry's shouldFallback catches it and walks to the next provider with a fresh timer. Per-attempt, not chain-wide — each provider gets its own budget.
When the chain advances due to a timeout, onFallback fires with cause: "provider-error" (the timeout-induced ProviderUnavailableError is the trigger). Caller-supplied signal composes with the timeout — both fire the same wrapped controller; the shorter trigger wins.
This is the ergonomic wrapper for the AbortSignal infrastructure that already existed on the port surface (alpha.6+). Particularly useful for routing around reasoning models that grind on hidden chain-of-thought without erroring.
Error swallowing
Every hook is fire-and-forget. Sync hooks that throw, async hooks that reject — the Registry swallows the failure and continues the inference call. Observability instrumentation can never break inference.
const registry = createRegistryFromEnv({
/* ... */
observability: {
onCost: () => { throw new Error("oops"); }, // swallowed; call proceeds
},
});
const result = await registry.getPort().generateText(/* ... */);
// result is returned normally even though the hook threwThis matches the contract on the existing onRetry hook (emitRetryEvent for the implementation).
Mapping to OpenTelemetry
Each event field maps to a gen_ai.* semantic convention or a vendor-neutral extension. For an OTel Collector pipeline:
| llm-ports field | OTel convention |
|---|---|
modelId | gen_ai.response.model |
providerAlias | gen_ai.system (or vendor extension) |
operation | gen_ai.operation.name |
taskType | gen_ai.request.task_type (vendor extension) |
inputTokens | gen_ai.usage.input_tokens |
outputTokens | gen_ai.usage.output_tokens |
cachedInputTokens | gen_ai.usage.cached_input_tokens (vendor extension) |
totalUsd | gen_ai.usage.cost.total_usd (vendor extension) |
fromAlias / toAlias (onFallback) | span attributes on the fallback event span |
The hooks deliberately stay vendor-neutral on the field names. Map to the conventions your tracing layer expects at the hook callback boundary.
Two related tools that are not hooks
combineSinks: more than one sink on one registry
Added in 0.1.0-alpha.35, exported from @llm-ports/observability-contract. The contract surface takes one sink, and sending events to both a tracing backend and your own logger previously meant writing the fan-out yourself.
import { combineSinks } from "@llm-ports/observability-contract";
import { createOtelSink } from "@llm-ports/telemetry-otel";
const sink = combineSinks(createOtelSink({ tracer }), myAuditSink);Failures are isolated per sink. A sink that throws, or that returns a rejected promise, does not stop the others receiving the event and does not affect the call. That is the same promise every hook here makes, extended to composition: observability never breaks inference.
createRetryRecorder: a bounded view of recent retries
Added in 0.1.0-alpha.35, exported from @llm-ports/core. For a dashboard or a health check that wants to show recent instability without subscribing to every event.
import { createRetryRecorder } from "@llm-ports/core";
const retries = createRetryRecorder(); // retains 50 by default
const adapters = {
openai: createOpenAIAdapter({ apiKey, onRetry: retries.onRetry }),
google: createGoogleAdapter({ apiKey, onRetry: retries.onRetry }),
};
retries.recent(10); // newest first
retries.total; // every retry seen, including evicted onesNote where it attaches, because it is deliberately not a registry method. Retries happen inside adapters: provider backoff, and the retry-with-feedback round when a model returns the wrong shape. What the registry does when a provider fails is walk to the next one, which is a fallback rather than a retry and is already reported by onFallback. A registry.recentRetries() could therefore only have returned fallbacks under a name that says retries, so the recorder is handed to the adapters instead. One recorder is safe across all of them, since every event carries its own provider alias.
It is bounded by construction because the reason to want it is recent instability, and an unbounded log of every retry a long-running worker ever performed is a memory leak wearing a feature's clothes. total survives eviction, so "three retries" stays distinguishable from "three retained out of two hundred".
See also
- Retry observability (
onRetry) — for per-adapter retry signals - Cost vs Request Gating — for the gating mechanics that drive
onFallbackcause discrimination - Cache Control — for the
cacheControlshape that feedsonCacheHit - Validation Strategies — for the
retry-with-feedbackmechanism behindonValidationRetry