Skip to content

@llm-ports/adapter-google ​

Native Google Gemini adapter for llm-ports, built on the unified @google/genai SDK (v2.x). Implements LLMPort with full multimodal support — image content blocks pass through as inlineData (base64) or fileData (URL), not degraded to placeholder text.

Shipped in 0.1.0-alpha.5.

Install ​

bash
pnpm add @llm-ports/core @llm-ports/adapter-google @google/genai zod

Configure ​

ts
import { createRegistryFromEnv } from "@llm-ports/core";
import { createGoogleAdapter } from "@llm-ports/adapter-google";

const registry = createRegistryFromEnv({
  adapters: {
    google: createGoogleAdapter({
      apiKey: process.env.GOOGLE_API_KEY!, // from https://aistudio.google.com/apikey
    }),
  },
});

export const llm = registry.getPort();

.env:

LLM_PROVIDER_FAST=google|gemini-2.5-flash|cost:5/day
LLM_PROVIDER_PREMIUM=google|gemini-2.5-pro|cost:50/day
LLM_TASK_ROUTE_TRIAGE=fast,premium

Adapter options ​

ts
interface GoogleAdapterOptions {
  apiKey: string;
  pricingOverrides?: Record<string, ModelPricing>;
  validationStrategy?: ValidationStrategy;
  imageSizeLimitBytes?: number; // default 20 MB
  onRetry?: OnRetry;            // alpha.17+
}

onRetry observability hook (alpha.17) ​

The adapter fires onRetry whenever it retries a generateStructured call after a Zod validation failure. Sync or async; called fire-and-forget; throwing from the hook does NOT cancel the retry. Pipe events into any tracing or metrics stack.

ts
import { createGoogleAdapter } from "@llm-ports/adapter-google";

const adapter = createGoogleAdapter({
  apiKey: process.env.GOOGLE_API_KEY!,
  onRetry: (event) => {
    // Langfuse / Phoenix / OpenLLMetry / Datadog all accept this shape
    span.addEvent("llm.retry", {
      reason: event.reason,          // "validation-feedback" for Gemini
      attempt: event.attempt,        // 0-indexed retry number
      modelId: event.modelId,
      providerAlias: event.providerAlias,
      delayMs: event.delayMs,
    });
  },
});

Gemini only fires the validation-feedback reason (no transient-auth or capability-fallback retry paths — Gemini doesn't 401 in the burst-protection shape and its parameter compatibility is uniform across models). The event shape matches the OpenAI and Anthropic adapters so consumers can wire one hook across all adapters.

Why this over the OpenAI-compat baseURL ​

Gemini exposes an OpenAI-compatible surface at https://generativelanguage.googleapis.com/v1beta/openai/. It works for most cases. Reasons to prefer this native adapter:

ConcernOpenAI-compat baseURLadapter-google
ImageSource.detailSilently ignored — Gemini has no equivalentIgnored explicitly (consistent with adapter-anthropic)
systemInstructionPrepended to user message, changing Gemini's behaviorNative top-level field
Multimodal richnessimage_url with base64 data URI (lossy)inlineData with explicit mediaType
Bundled pricingNone — bring your ownGemini 2.5 + 2.0 family bundled
Image-block boundary validationInherits from adapter-openaiFirst-class, with imageSizeLimitBytes option
Native responseSchemaNot exposed✓ (alpha.9; falls back to prompted-JSON when the schema contains oneOf/allOf/$ref)

Bundled pricing ​

ModelInput/1MOutput/1MCache read
gemini-2.5-pro$1.25$5.00$0.3125
gemini-2.5-flash$0.075$0.30$0.01875
gemini-2.5-flash-lite$0.0375$0.15$0.009375
gemini-2.0-flash$0.10$0.40$0.025
gemini-2.0-flash-lite$0.075$0.30—

Source: https://ai.google.dev/gemini-api/docs/pricing (verified 2026-05).

Long-context premium: bundled values are the under-200k-token rates. Gemini charges a higher rate above 200k tokens. For long-context workloads, supply pricingOverrides with the over-200k rates.

Supported features (v0.1) ​

FeatureStatus
generateText✓
generateStructured (Zod schemas)✓ (alpha.9: native responseSchema constrained-decoding when the schema converts cleanly; falls back to prompted JSON + alpha.5 repair pass for unsupported schema features)
streamText✓
streamStructured✓ (best-effort partial parse)
runAgent (multi-turn tool use)✓ (alpha.9: full function-calling loop with parallel-call support, aggregated usage, populated toolCalls)
Vision input — base64 images✓ (inlineData)
Vision input — URL images✓ (fileData)
Audio input — base64✓ (inlineData)
Image-block size + URL validation at boundary✓ (alpha.5)
AbortSignal cancellation✓ entry + in-flight (alpha.6)
listModels()✓ (alpha.9; via @google/genai client.models.list())
Embeddings (gemini-embedding-001)✗ — v0.2
Explicit context caching✗ — v0.2
Code execution tool✗ — v0.2

Native responseSchema (alpha.9) ​

generateStructured emits Gemini's native config.responseSchema + config.responseMimeType: "application/json" when the Zod schema converts cleanly. Gemini constrains decoding to the schema before tokens are produced — invalid JSON and missing required fields are impossible (modulo provider bugs). Zod validation + the alpha.5 repair pass + retry-with-feedback remain the safety net.

ts
const result = await llm.generateStructured({
  taskType: "extract",
  prompt: "Extract the person: 'Babak is 42'",
  schema: z.object({ name: z.string(), age: z.number() }),
});
// Adapter sends config.responseSchema; Gemini constrains decoding.
// validationAttempts will typically be 1 (constrained decoding rarely fails Zod).

Fallback path. Gemini's responseSchema accepts the OpenAPI 3.0 subset of JSON Schema. The adapter detects unsupported features and falls back to the prompted-JSON path:

Zod constructJSON Schema outputNative path?
z.object(...){ type: "object", properties: ... }✓
z.array(z.string()){ type: "array", items: ... }✓
z.discriminatedUnion(...){ anyOf: [...] }✓ (Gemini accepts anyOf)
z.enum(...){ enum: [...] }✓
z.intersection(a, b){ allOf: [...] }✗ — falls back to prompted JSON
z.lazy(...) (recursive){ $ref: "#" }✗ — falls back
Hand-rolled oneOf / notas-is✗ — falls back

When a fallback fires, the adapter emits a one-time console.warn per (model, feature) pair naming the unsupported construct. The output is still correct in either case — only the constrained-decoding guarantee differs.

Multi-turn runAgent (alpha.9) ​

runAgent translates options.tools to Gemini's Tool[] shape (function declarations with JSON Schema, OpenAPI 3.0 subset, via zod-to-json-schema), then loops the chat / function-call / function-response cycle until the model returns text only (terminationReason: "completed") or maxSteps is reached (terminationReason: "max_steps").

Gemini emits parallel function calls (multiple functionCall parts in a single response) and expects all functionResponse parts back together — the adapter executes them in order and groups the responses. toolCalls in the result is fully populated; usage aggregates across steps.

ts
const result = await llm.runAgent({
  taskType: "research",
  instructions: "Answer the user's question using tools as needed.",
  messages: [{ role: "user", content: "What's the weather in Paris?" }],
  tools: {
    getWeather: {
      name: "getWeather",
      description: "Fetch current weather for a city.",
      inputSchema: z.object({ city: z.string() }),
      execute: async ({ city }) => ({ tempC: 22, condition: "sunny", city }),
    },
  },
  maxSteps: 4,
});
console.log(result.text);              // "It's 22°C and sunny in Paris."
console.log(result.toolCalls);         // [{ name: "getWeather", input: {city: "Paris"}, output: {...} }]
console.log(result.stepsTaken);        // 2
console.log(result.terminationReason); // "completed"
console.log(result.usage.inputTokens); // aggregated across both steps

listModels() (alpha.9) ​

ts
const models = await port.listModels();
// [{ id: "gemini-2.5-flash", displayName: "Gemini 2.5 Flash", contextWindow: 1048576, ... }, ...]

Pricing is not surfaced (Gemini's /models endpoint exposes catalog metadata but not USD rates). Registry.checkPricingFreshness() uses this to detect added/removed models.

Content blocks supported ​

text, image (base64 → inlineData; URL → fileData), audio (base64 only), tool_use, tool_result. The adapter throws ContentBlockUnsupportedError for unsupported variants (audio URLs).

Cancellation ​

Full AbortSignal support shipped in 0.1.0-alpha.6. Threading the signal cancels the in-flight provider HTTP fetch, not just the JS await:

ts
const controller = new AbortController();
const promise = llm.generateText({
  taskType: "describe_image",
  prompt: [...],
  signal: controller.signal,
});
// User clicks cancel:
controller.abort();
// promise rejects; the HTTP request to generativelanguage.googleapis.com is cancelled.

See the Cancellation guide for the full pattern.

Image cost note ​

Gemini does not have a separate cost-vs-fidelity knob equivalent to OpenAI's image_url.detail. Image cost is determined by the model's automatic tiling — typically ~258 tokens per image for gemini-2.5-flash, ~1,290 for high-resolution inputs to gemini-2.5-pro. If you set ImageSource.detail on a call routed to a Gemini model, the adapter ignores the field (consistent with adapter-anthropic).

Structured output from a JSON Schema ​

Since 0.1.0-alpha.35, generateStructured accepts jsonSchema in place of schema, for a caller that already holds a JSON Schema rather than a Zod one:

ts
await llm.generateStructured<Triage>({ taskType: "classify", messages, jsonSchema });

Exactly one of the two is required. What the JSON Schema path gives up is local validation, because this library carries no JSON Schema validator: the response comes back unvalidated, there is no retry-with-feedback when a model returns the wrong shape, validationAttempts is always 1, and T is yours to assert. Prefer schema unless you genuinely hold a JSON Schema already. streamStructured is unchanged and still requires Zod.

Reading next ​

MIT License