ADR 0016: Gemini on Vertex AI as the agents' daily overflow
- Status: Superseded by ADR 0021 (Anthropic for every tier; no Gemini overflow) on 2026-10-09.
- Date: 2026-10-09
- Decider: Project owner
Context
Section titled “Context”Agents (ADR 0009) used only OpenRouter free models. OpenRouter allows about 50 free requests per
day, and the agent cap was 40 calls per day. One demo Change could not finish in a day. The
owner approved Gemini on Vertex AI, billed to GCP credits, as overflow. ADR 0007 already made
Vertex a capped backup for the Root Planner, with an AI Gateway spend limit of $9 per 30 days
on the google-vertex-ai provider.
Decision
Section titled “Decision”- The agents Worker gets the
VERTEX_SERVICE_ACCOUNTsecret_textbinding (Terraform, dev). - Each tier has two Pi models:
tier-<t>(OpenRouter, $0) andtier-<t>-vertex(Gemini through the gateway compat endpoint,google-vertex-ai/<TIER_VERTEX_MODEL>). Vertex models carry VERTEX_PRICES, so Pi records real cost. Run cost, the Change budget, and the Costs page use it. - The AgentDO picks the provider once, when a run starts. It uses OpenRouter unless today’s
OpenRouter quota is spent. Either of two signals marks it spent: a D1
usage_countersrowopenrouter-exhaustedfor the UTC day (set on an OpenRouter 429 or quota error), or Rhumbatron’s ownopenrouter-callscounter at 45. - An OpenRouter rate limit during a run fails the run as
rate_limited. The same-tier retry (section 42.8) then starts on Vertex. No provider switch happens inside a run. - Caps: 45 OpenRouter calls per day; 400 agent calls per day on all providers. The gateway spend limit is the hard money ceiling.
- A Vertex 429 from the gateway spend limit is
spend_limited. The gateway documents only the status, so a Vertex 429 that is not Google’sRESOURCE_EXHAUSTEDor the gateway request rate limit counts as the spend limit. It is a budget stop. There is no retry, the Task isblocked, and the Change isbudget_blockedwith a reason that names the $9 per 30 days Gemini cap. Avertex-spend-limitedrow for the day stops later runs before they call a model. - Gemini 3 rejects a replayed tool call without its thought signature. pi-ai does not round-trip
extra_content, so the provider keeps signatures per tool call id and sends Google’s documented bypass value when one is lost. - The Decision Receipt records
providerand the reason; the Agents page marks Gemini attempts.
Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens. A sonnet run of 16 calls with about 15,000 input and 1,000 output tokens per call costs about $0.24. The $9 cap allows about 35 such runs per 30 days.
Consequences
Section titled “Consequences”- Agent spend is no longer always $0. Cost budgets per run and per Change now apply to real spend.
- Prompts reach Google on overflow runs. Send only demo fixture code, as for free models.
Update (2026-10-10)
Section titled “Update (2026-10-10)”- The
VERTEX_SERVICE_ACCOUNTbinding is on the agents and integration Workers in dev and prod. MODEL_PROVIDERselects the mode. Dev usesvertex: every call goes to Gemini from the first call. Prod usesopenrouter: free models first, with Gemini as overflow (ADR 0017).- The gateway cap is per environment: dev $40 and prod $10 per 30 days (ADR 0017). Budget-stop
text reads the cap from
VERTEX_SPEND_LIMIT_USD; $9 is only the code fallback. - Tier models (
TIER_VERTEX_MODEL): haiku usesgoogle/gemini-3.5-flash-lite; sonnet and opus usegoogle/gemini-3.8-flash.
Diagram
Section titled “Diagram”Provider choice at run start (selectProvider, packages/model-router/src/provider.ts).
flowchart TD
S["Run starts"] --> V{"MODEL_PROVIDER = vertex<br/>and Vertex configured?"}
V -- yes --> L{"vertex-spend-limited<br/>row today?"}
V -- no --> O{"openrouter-exhausted today,<br/>or 45 OpenRouter calls?"}
O -- no --> OR["OpenRouter free model"]
O -- yes --> C{"Vertex configured?"}
C -- no --> OR
C -- yes --> L
L -- no --> G["Gemini on Vertex<br/>through the gateway"]
L -- yes --> B["Budget stop: Task blocked,<br/>Change budget_blocked"]
Every call also takes one of the 400 agent calls per UTC day on all providers.