Skip to content

Hermes model configuration & upgrade runbook ​

How the LLM behind Hermes is selected, upgraded, and rolled back. Scope note: the reader and gateway live in this repo under supabase/functions/_shared/modelConfig.ts and supabase/functions/hermes-gateway/index.ts. Edge callers use getModel() via _shared/aiClient.ts. Production deploys hermes-gateway on project jrsgosnnyjonxaesqtln; active pins read from public.hermes_config (model.chat, model.reasoning). The mobile app is provider-agnostic β€” it only streams the OpenAI-SSE shape (openai_edge_stream.dart) and never knows which model answered.

Principle ​

The model is a swappable component, not a hardcoded constant. Hermes's "house voice" / consistency comes from its retrieval layer (hermes_embeddings, ai_knowledge_base, hermes_search_logs), so the base model can change without changing the product's voice β€” provided every swap passes a shadow + eval gate (a new model still synthesizes/phrases retrieved knowledge, and can shift Arabic tone, JSON adherence, tool-calling, cost, and latency).

Auto-detect a new model: yes. Auto-deploy it: never.

Where the model is configured: public.hermes_config ​

Key/value (jsonb) table. Seeded keys:

keymeaning
model.chatactive model for chat / WhatsApp / captions / RAG synthesis
model.reasoningactive model for heavy reasoning (valuation analysis, agent planning)
model.transcribeactive foundation model for audio transcription & voice notes (gemini-3.5-transcribe)
model.chat.targetrecommended upgrade for model.chat, staged (not active)
model.reasoning.targetrecommended upgrade for model.reasoning, staged

The active keys are seeded to the current production models, so reading them is behaviour-neutral until a deliberate promotion. Active: google/gemini-3.7-flash across general agent tiers (chat, reasoning, vision, and video walkthrough comprehension) and google/gemini-3.5-transcribe for dedicated high-throughput audio speech-to-text.

βœ… Live: hermes-gateway edge function reads hermes_config via modelConfig.ts. Until a pin is promoted in SQL, fallbacks match the seeded production models below.

Reader snippet (Flutter-owned Hermes service) ​

ts
// hermes/lib/modelConfig.ts β€” TS for supabase-js. Swap Deno.env for
// process.env / Worker env if Hermes isn't a Supabase edge function.
import { createClient, type SupabaseClient } from "@supabase/supabase-js";

type ModelRole = "chat" | "reasoning" | "vision" | "video";

// Fallback = current production models: a DB hiccup degrades to status quo.
const FALLBACK: Record<ModelRole, string> = {
  chat: "google/gemini-3.7-flash",
  reasoning: "google/gemini-3.7-flash",
  vision: "google/gemini-3.7-flash",
  video: "google/gemini-3.7-flash",
};
const CONFIG_KEY: Record<ModelRole, string> = {
  chat: "model.chat",
  reasoning: "model.reasoning",
};

const TTL_MS = 60_000;
let cache: { at: number; values: Record<string, string> } | null = null;
let _client: SupabaseClient | null = null;

function client(): SupabaseClient {
  return (_client ??= createClient(
    Deno.env.get("SUPABASE_URL")!,
    Deno.env.get("SUPABASE_SERVICE_ROLE_KEY")!,
    { auth: { persistSession: false } },
  ));
}

async function load(): Promise<Record<string, string>> {
  if (cache && Date.now() - cache.at < TTL_MS) return cache.values;
  try {
    const { data, error } = await client()
      .from("hermes_config").select("key, value")
      .in("key", Object.values(CONFIG_KEY));
    if (error) throw error;
    const values: Record<string, string> = {};
    for (const row of data ?? []) {
      if (typeof row.value === "string") values[row.key] = row.value;
    }
    cache = { at: Date.now(), values };
    return values;
  } catch (e) {
    console.warn("modelConfig: using fallback,", e);
    return {};
  }
}

export async function getModel(role: ModelRole): Promise<string> {
  return (await load())[CONFIG_KEY[role]] ?? FALLBACK[role];
}
export function invalidateModelConfigCache(): void { cache = null; }

Wire it into each AI edge function:

ts
// before:  model: "google/gemini-3-flash-preview"
import { getModel } from "../lib/modelConfig.ts";
const model = await getModel("chat"); // "reasoning" for valuation analysis

Promotion runbook (current β†’ target) ​

  1. Confirm the exact slug for the target against the gateway catalog.
  2. Freeze an eval set (~20–40 real prompts from ai_usage_logs / hermes_search_logs): Arabic caption, valuation reasoning, tool-use loop, JSON extraction, WhatsApp auto-reply, Kuwaiti-dialect/RTL case.
  3. Shadow the target (ai_engine=shadow) β†’ lands in whatsapp_shadow_replies beside the live model. Accumulate a few days.
  4. Score shadow vs live: Arabic/dialect quality, valuation numeric sanity, tool-call correctness, JSON validity, refusals + cost & p50/p95 from ai_usage_logs (estimated_cost_usd, duration_ms).
  5. Gate: promote only if quality β‰₯ current and within ai_usage_budgets.
  6. Promote (one row):
    sql
    update public.hermes_config
      set value = '"google/gemini-3.5-flash"'::jsonb, updated_at = now()
      where key = 'model.chat';
    Live within the 60s cache TTL (or call invalidateModelConfigCache()).
  7. Canary + rollback: watch ai_usage_logs.error_message + latency for the first hours; rollback is the same UPDATE back to the prior value.
  8. Record date + eval scores in the row description for auditability.

Optional: auto-detect (never auto-deploy) ​

A weekly cron lists the gateway/provider catalog, diffs vs *.target, and if a newer stable model exists, opens a task "new stable model β€” run the eval." A human still flips the pin.

Data-residency note ​

Every model call ships context (incl. client PII, and anything touching registry_voters) to the provider. Decide deliberately: a DPA + zero-retention setting, region routing, or a self-hosted/regional model for the most sensitive flows. Track which function_names touch sensitive tables.

Aldilaijan & Khobara Real Estate Platform