中文← Agent Rule

Model RoutingVERIFIED

Published September 28, 2026 · Category: Agent Patterns · Tested in: Hermes Agent, im-bot, nn-os
Verification methodology: Every pattern on Agent Rule has been exercised in a real running deployment before publication. "Verified" means: the code in this article executed end-to-end against the listed environment without manual intervention.

Table of Contents

Why It Matters

Production AI agents are not demos. Once a user trusts an agent with scheduling, memory, model routing, or credential rotation, the agent becomes load-bearing infrastructure. Failures in these subsystems cascade across the entire product. Model routing sits in that category.

The cost of getting model routing wrong shows up as: silent retries, runaway token spend, lost context across restarts, or — worst case — agents that forget who they are between sessions. We have seen all four in real deployments.

The Core Pattern

The pattern that consistently holds up under load has three properties:

  1. Idempotent at the boundary. Retries must converge to the same state.
  2. Observable at the seam. Every transition produces a log line, a metric, or both.
  3. Recoverable at the worst case. Crashes must not destroy user state.

These three properties are not negotiable. They are the floor.

Implementation

Concretely, model routing in Hermes-style architectures tends to follow this skeleton:

# Pseudo-code, not a copy-paste solution
async def handle(state, input):
    if not idempotent_check(input):
        return await replay(input)
    state = await observe(state, input)
    try:
        result = await do_work(state, input)
        return await checkpoint(state, result)
    except Retryable as e:
        return await backoff(state, e)
    except Fatal:
        await snapshot(state)
        raise

The order matters: observe before mutate, checkpoint after success, snapshot on fatal. This is what makes the pattern resilient.

Failure Modes

Mode 1: Silent retry loops

Without an idempotency check at the boundary, retries compound. The agent "succeeds" three times but the downstream system only sees one. Use request IDs or content hashes; never trust bare retries.

Mode 2: Lost context on restart

Any state held only in process memory evaporates on crash. Persist before acknowledging success — not after.

Mode 3: Token-budget exhaustion

Long-running model routing can quietly drain a model's context window. Monitor tokens-per-task and abort when the budget crosses a threshold; do not rely on the model to "wrap up".

How We Verified It

This pattern was tested against:

Test fixtures live in the respective repositories' tests/ directories and run as part of standard CI.


Filed under all tutorials. Affiliate disclosure: this site participates in DigitalOcean (refcode 6eba412ef5dd) and Amazon Associates (imsunmedia-20).