Model Routing Simulator

Static MOCK presets · deterministic policy · no live model calls · recommendations are advisory
DCP READ-ONLY: CONNECTED DEVELOPMENT / UI PREVIEW DCP BATCH 5/6 SOAK IN PROGRESS TELEMETRY: UNKNOWN COMMAND EXECUTION LOCKED
READ-ONLY SIMULATOR · Deterministic policy evaluation over a STATIC development registry. No live model calls, no routing changes, no execution. Recommendations are advisory; authority stays with the owner.

Task Input (preset scenario)

1 · Git inventory (read-only audit)

Task typeGIT_AUDIT
Agent roleHERMES
Risk levelR1
ComplexityLOW
Context sizeSMALL
VisionNONE
Repo writeno
Long tool loopno
Architecture reasoningno
Code generationno
Security reviewno
Cost sensitivityMEDIUM
Latency sensitivityLOW
Preset expectation: GLM-5.3-Flash(x) eligible/preferred if policy evidence supports.

Recommended Agent

HERMES

Static default policy: planning/reasoning first · prefer verified architecture/reasoning model · fallback only if evidence supports

Primary Model

GLM-5.3-Flash(x) glm-5.3-flashx VERIFIED_FOR_MEDIUM_MULTI_FILE

Fallback Model

DeepSeek V4.1 Flash deepseek-v4.1-flash VERIFIED

DeepSeek V4.1 Flash (deepseek-v4.1-flash)

Status: VERIFIED · last verified: 2026-10-05 · provenance: STATIC_CONFIG

Verified task classes: MEDIUM_FEATURE, LARGE_FEATURE, SMALL_PATCH, CODE_REVIEW, GIT_AUDIT, README_UPDATE · Unverified: UI_VISION_REVIEW

  • heavy coding proven [REAL_OBSERVATION]
  • long agent loops proven [REAL_OBSERVATION]
  • multi-file repo work proven [REAL_OBSERVATION]
  • testing/tool loop proven [REAL_OBSERVATION]

Evidence — primary model

GLM-5.3-Flash(x) (glm-5.3-flashx)

Status: VERIFIED_FOR_MEDIUM_MULTI_FILE · last verified: 2026-10-05 · provenance: STATIC_CONFIG

Verified task classes: GIT_AUDIT, README_UPDATE, SMALL_PATCH, MEDIUM_FEATURE · Unverified: LARGE_FEATURE, SECURITY_REVIEW, UI_VISION_REVIEW

  • read-only Git inventory PASS [REAL_OBSERVATION]
  • untracked-file inspection PASS [REAL_OBSERVATION]
  • small repo hygiene coding patch PASS [REAL_OBSERVATION]
  • 531-test verification PASS [REAL_OBSERVATION]
  • medium multi-file task PASS (MODEL_ROUTING_POLICY_SIMULATOR_V1: ~9 files, 286/286 portal tests, 39 new) [REAL_OBSERVATION]
  • large feature / long agent loop work NOT YET VERIFIED; one medium task is not enough evidence [REAL_OBSERVATION]

SCORE (transparent, advisory)

The weighted score is advisory only. Hard blockers always override the score. UNKNOWN never counts as positive evidence.
DimensionValueStatic weight
Capability Match0.60weight 0.30
Evidence Strength1.00weight 0.25
Task Complexity Fit0.60weight 0.15
Risk Fit0.80weight 0.10
Context Fit1.00weight 0.08
Cost Fit1.00weight 0.07
Latency Fit1.00weight 0.03
Vision Fit0.50weight 0.02
TOTAL0.790

Routing Reasons

WHY THIS MODEL?

  • STATUS — GLM-5.3-Flash(x) status = VERIFIED_FOR_MEDIUM_MULTI_FILE
  • EVIDENCE — verified task classes: GIT_AUDIT, README_UPDATE, SMALL_PATCH, MEDIUM_FEATURE
  • TASK_MATCH — direct verified evidence for task class GIT_AUDIT
  • PARTIAL_SCOPE — VERIFIED_FOR_MEDIUM_MULTI_FILE: verification is scoped to evidenced task classes only; larger task classes remain unverified
  • NOT_PREFERRED_deepseek-v4.1-flash — DeepSeek V4.1 Flash scored 0.604 vs 0.79 (gap 0.186)
  • EXCLUDED_kimi-k3 — Kimi K3 excluded: unavailable (upstream HTTP 503 on trivial prompt; provider retry failed)
  • FALLBACK — fallback DeepSeek V4.1 Flash (deepseek-v4.1-flash): VERIFIED, next-best eligible score 0.604

Hard Blockers

  • none

WHAT IS STILL UNKNOWN?

  • DeepSeek V4.1 Flash: vision support UNKNOWN
  • GLM-5.3-Flash(x): vision support UNKNOWN
  • Kimi K3: vision support UNKNOWN
  • Kimi K3: cost multiplier UNKNOWN
  • GLM-5.3-Flash(x): large feature / long agent loop work not yet verified

WHAT WOULD CHANGE THE DECISION?

  • If GLM-5.3-Flash(x) accumulates 2-3 more successful medium multi-file tasks and comparative token/cost data, the default switch may be reconsidered (currently NOT YET APPROVED)
  • If GLM-5.3-Flash(x) completes a verified LARGE_FEATURE task, it becomes eligible for large-feature routing
  • If Kimi K3 becomes reachable AND accumulates verified evidence in our environment, it re-enters the candidate pool
  • If this task is R3/R4, owner approval is required regardless of model choice

Diversity-of-Failure

DIVERSITY_NOT_APPLICABLE not a developer hand-off scenario

MODEL ROUTING != EXECUTION AUTHORITY

Selecting a model never grants sudo, Docker mutation, production write, DB mutation, deploy, network change or approval power. Authority is governed by risk policy, owner approval and DCP controls.

Model Status (static development registry)

Kimi K3 is UNAVAILABLE (HTTP 503 upstream provider failure) and is never displayed as selectable or recommended.
ModelStatusProven task classesKnown failuresLast verification
DeepSeek V4.1 FlashVERIFIEDMEDIUM_FEATURE, LARGE_FEATURE, SMALL_PATCH, CODE_REVIEW, GIT_AUDIT, README_UPDATE—2026-10-05
GLM-5.3-Flash(x)VERIFIED_FOR_MEDIUM_MULTI_FILEGIT_AUDIT, README_UPDATE, SMALL_PATCH, MEDIUM_FEATURE—2026-10-05
Kimi K3UNAVAILABLEupstream HTTP 503 on trivial prompt; provider retry failed