Task Input (preset scenario)
3 · Medium multi-file feature
| Task type | MEDIUM_FEATURE |
|---|---|
| Agent role | DEVELOPER |
| Risk level | R2 |
| Complexity | MEDIUM |
| Context size | MEDIUM |
| Vision | NONE |
| Repo write | yes |
| Long tool loop | yes |
| Architecture reasoning | no |
| Code generation | yes |
| Security review | no |
| Cost sensitivity | MEDIUM |
| Latency sensitivity | LOW |
Recommended Agent
DEVELOPER
Static default policy: coding evidence dominates · multi-file/long-loop evidence required for larger tasks
Primary Model
DeepSeek V4.1 Flash deepseek-v4.1-flash VERIFIED
Fallback Model
GLM-5.3-Flash(x) glm-5.3-flashx VERIFIED_FOR_MEDIUM_MULTI_FILE
GLM-5.3-Flash(x) (glm-5.3-flashx)
Status: VERIFIED_FOR_MEDIUM_MULTI_FILE · last verified: 2026-10-05 · provenance: STATIC_CONFIG
Verified task classes: GIT_AUDIT, README_UPDATE, SMALL_PATCH, MEDIUM_FEATURE · Unverified: LARGE_FEATURE, SECURITY_REVIEW, UI_VISION_REVIEW
- read-only Git inventory PASS [REAL_OBSERVATION]
- untracked-file inspection PASS [REAL_OBSERVATION]
- small repo hygiene coding patch PASS [REAL_OBSERVATION]
- 531-test verification PASS [REAL_OBSERVATION]
- medium multi-file task PASS (MODEL_ROUTING_POLICY_SIMULATOR_V1: ~9 files, 286/286 portal tests, 39 new) [REAL_OBSERVATION]
- large feature / long agent loop work NOT YET VERIFIED; one medium task is not enough evidence [REAL_OBSERVATION]
Evidence — primary model
DeepSeek V4.1 Flash (deepseek-v4.1-flash)
Status: VERIFIED · last verified: 2026-10-05 · provenance: STATIC_CONFIG
Verified task classes: MEDIUM_FEATURE, LARGE_FEATURE, SMALL_PATCH, CODE_REVIEW, GIT_AUDIT, README_UPDATE · Unverified: UI_VISION_REVIEW
- heavy coding proven [REAL_OBSERVATION]
- long agent loops proven [REAL_OBSERVATION]
- multi-file repo work proven [REAL_OBSERVATION]
- testing/tool loop proven [REAL_OBSERVATION]
SCORE (transparent, advisory)
| Dimension | Value | Static weight |
|---|---|---|
| Capability Match | 0.60 | weight 0.30 |
| Evidence Strength | 1.00 | weight 0.25 |
| Task Complexity Fit | 1.00 | weight 0.15 |
| Risk Fit | 1.00 | weight 0.10 |
| Context Fit | 1.00 | weight 0.08 |
| Cost Fit | 0.72 | weight 0.07 |
| Latency Fit | 0.80 | weight 0.03 |
| Vision Fit | 0.50 | weight 0.02 |
| TOTAL | 0.844 |
Routing Reasons
WHY THIS MODEL?
- STATUS — DeepSeek V4.1 Flash status = VERIFIED
- EVIDENCE — verified task classes: MEDIUM_FEATURE, LARGE_FEATURE, SMALL_PATCH, CODE_REVIEW, GIT_AUDIT, README_UPDATE
- TASK_MATCH — direct verified evidence for task class MEDIUM_FEATURE
- MULTI_FILE — proven multi-file/long-loop evidence for repo implementation
- NOT_PREFERRED_glm-5.3-flashx — GLM-5.3-Flash(x) scored 0.79 vs 0.844 (gap 0.054)
- EXCLUDED_kimi-k3 — Kimi K3 excluded: unavailable (upstream HTTP 503 on trivial prompt; provider retry failed)
- FALLBACK — fallback GLM-5.3-Flash(x) (glm-5.3-flashx): VERIFIED_FOR_MEDIUM_MULTI_FILE, next-best eligible score 0.79
Hard Blockers
- none
WHAT IS STILL UNKNOWN?
- DeepSeek V4.1 Flash: vision support UNKNOWN
- GLM-5.3-Flash(x): vision support UNKNOWN
- Kimi K3: vision support UNKNOWN
- Kimi K3: cost multiplier UNKNOWN
WHAT WOULD CHANGE THE DECISION?
- If a cheaper model accumulates verified evidence for this task class, cost preference may flip the primary
- If Kimi K3 becomes reachable AND accumulates verified evidence in our environment, it re-enters the candidate pool
- If this task is R3/R4, owner approval is required regardless of model choice
Diversity-of-Failure
DIVERSITY_PREFERRED QA reviewer should prefer GLM-5.3-Flash(x) (glm-5.3-flashx) — different failure modes from DeepSeek V4.1 Flash; evidence quality preserved
MODEL ROUTING != EXECUTION AUTHORITY
Selecting a model never grants sudo, Docker mutation, production write, DB mutation, deploy, network change or approval power. Authority is governed by risk policy, owner approval and DCP controls.
Model Status (static development registry)
| Model | Status | Proven task classes | Known failures | Last verification |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | VERIFIED | MEDIUM_FEATURE, LARGE_FEATURE, SMALL_PATCH, CODE_REVIEW, GIT_AUDIT, README_UPDATE | — | 2026-10-05 |
| GLM-5.3-Flash(x) | VERIFIED_FOR_MEDIUM_MULTI_FILE | GIT_AUDIT, README_UPDATE, SMALL_PATCH, MEDIUM_FEATURE | — | 2026-10-05 |
| Kimi K3 | UNAVAILABLE | upstream HTTP 503 on trivial prompt; provider retry failed | ||