Hermes model configuration & upgrade runbook β
How the LLM behind Hermes is selected, upgraded, and rolled back. Scope note: the reader and gateway live in this repo under
supabase/functions/_shared/modelConfig.tsandsupabase/functions/hermes-gateway/index.ts. Edge callers usegetModel()via_shared/aiClient.ts. Production deployshermes-gatewayon projectjrsgosnnyjonxaesqtln; active pins read frompublic.hermes_config(model.chat,model.reasoning). The mobile app is provider-agnostic β it only streams the OpenAI-SSE shape (openai_edge_stream.dart) and never knows which model answered.
Principle β
The model is a swappable component, not a hardcoded constant. Hermes's "house voice" / consistency comes from its retrieval layer (hermes_embeddings, ai_knowledge_base, hermes_search_logs), so the base model can change without changing the product's voice β provided every swap passes a shadow + eval gate (a new model still synthesizes/phrases retrieved knowledge, and can shift Arabic tone, JSON adherence, tool-calling, cost, and latency).
Auto-detect a new model: yes. Auto-deploy it: never.
Where the model is configured: public.hermes_config β
Key/value (jsonb) table. Seeded keys:
| key | meaning |
|---|---|
model.chat | active model for chat / WhatsApp / captions / RAG synthesis |
model.reasoning | active model for heavy reasoning (valuation analysis, agent planning) |
model.transcribe | active foundation model for audio transcription & voice notes (gemini-3.5-transcribe) |
model.chat.target | recommended upgrade for model.chat, staged (not active) |
model.reasoning.target | recommended upgrade for model.reasoning, staged |
The active keys are seeded to the current production models, so reading them is behaviour-neutral until a deliberate promotion. Active: google/gemini-3.7-flash across general agent tiers (chat, reasoning, vision, and video walkthrough comprehension) and google/gemini-3.5-transcribe for dedicated high-throughput audio speech-to-text.
β Live:
hermes-gatewayedge function readshermes_configviamodelConfig.ts. Until a pin is promoted in SQL, fallbacks match the seeded production models below.
Reader snippet (Flutter-owned Hermes service) β
// hermes/lib/modelConfig.ts β TS for supabase-js. Swap Deno.env for
// process.env / Worker env if Hermes isn't a Supabase edge function.
import { createClient, type SupabaseClient } from "@supabase/supabase-js";
type ModelRole = "chat" | "reasoning" | "vision" | "video";
// Fallback = current production models: a DB hiccup degrades to status quo.
const FALLBACK: Record<ModelRole, string> = {
chat: "google/gemini-3.7-flash",
reasoning: "google/gemini-3.7-flash",
vision: "google/gemini-3.7-flash",
video: "google/gemini-3.7-flash",
};
const CONFIG_KEY: Record<ModelRole, string> = {
chat: "model.chat",
reasoning: "model.reasoning",
};
const TTL_MS = 60_000;
let cache: { at: number; values: Record<string, string> } | null = null;
let _client: SupabaseClient | null = null;
function client(): SupabaseClient {
return (_client ??= createClient(
Deno.env.get("SUPABASE_URL")!,
Deno.env.get("SUPABASE_SERVICE_ROLE_KEY")!,
{ auth: { persistSession: false } },
));
}
async function load(): Promise<Record<string, string>> {
if (cache && Date.now() - cache.at < TTL_MS) return cache.values;
try {
const { data, error } = await client()
.from("hermes_config").select("key, value")
.in("key", Object.values(CONFIG_KEY));
if (error) throw error;
const values: Record<string, string> = {};
for (const row of data ?? []) {
if (typeof row.value === "string") values[row.key] = row.value;
}
cache = { at: Date.now(), values };
return values;
} catch (e) {
console.warn("modelConfig: using fallback,", e);
return {};
}
}
export async function getModel(role: ModelRole): Promise<string> {
return (await load())[CONFIG_KEY[role]] ?? FALLBACK[role];
}
export function invalidateModelConfigCache(): void { cache = null; }Wire it into each AI edge function:
// before: model: "google/gemini-3-flash-preview"
import { getModel } from "../lib/modelConfig.ts";
const model = await getModel("chat"); // "reasoning" for valuation analysisPromotion runbook (current β target) β
- Confirm the exact slug for the target against the gateway catalog.
- Freeze an eval set (~20β40 real prompts from
ai_usage_logs/hermes_search_logs): Arabic caption, valuation reasoning, tool-use loop, JSON extraction, WhatsApp auto-reply, Kuwaiti-dialect/RTL case. - Shadow the target (
ai_engine=shadow) β lands inwhatsapp_shadow_repliesbeside the live model. Accumulate a few days. - Score shadow vs live: Arabic/dialect quality, valuation numeric sanity, tool-call correctness, JSON validity, refusals + cost & p50/p95 from
ai_usage_logs(estimated_cost_usd,duration_ms). - Gate: promote only if quality β₯ current and within
ai_usage_budgets. - Promote (one row):sqlLive within the 60s cache TTL (or call
update public.hermes_config set value = '"google/gemini-3.5-flash"'::jsonb, updated_at = now() where key = 'model.chat';invalidateModelConfigCache()). - Canary + rollback: watch
ai_usage_logs.error_message+ latency for the first hours; rollback is the sameUPDATEback to the prior value. - Record date + eval scores in the row
descriptionfor auditability.
Optional: auto-detect (never auto-deploy) β
A weekly cron lists the gateway/provider catalog, diffs vs *.target, and if a newer stable model exists, opens a task "new stable model β run the eval." A human still flips the pin.
Data-residency note β
Every model call ships context (incl. client PII, and anything touching registry_voters) to the provider. Decide deliberately: a DPA + zero-retention setting, region routing, or a self-hosted/regional model for the most sensitive flows. Track which function_names touch sensitive tables.
