Task Input (preset scenario)
2 · Small repo hygiene patch
| Task type | SMALL_PATCH |
|---|---|
| Agent role | DEVELOPER |
| Risk level | R1 |
| Complexity | LOW |
| Context size | SMALL |
| Vision | NONE |
| Repo write | yes |
| Long tool loop | no |
| Architecture reasoning | no |
| Code generation | yes |
| Security review | no |
| Cost sensitivity | HIGH |
| Latency sensitivity | LOW |
Recommended Agent
DEVELOPER
Static default policy: coding evidence dominates · multi-file/long-loop evidence required for larger tasks
Primary Model
GLM-5.3-Flash(x) glm-5.3-flashx VERIFIED_FOR_MEDIUM_MULTI_FILE
Fallback Model
DeepSeek V4.1 Flash deepseek-v4.1-flash VERIFIED
DeepSeek V4.1 Flash (deepseek-v4.1-flash)
Status: VERIFIED · last verified: 2026-10-05 · provenance: STATIC_CONFIG
Verified task classes: MEDIUM_FEATURE, LARGE_FEATURE, SMALL_PATCH, CODE_REVIEW, GIT_AUDIT, README_UPDATE · Unverified: UI_VISION_REVIEW
- heavy coding proven [REAL_OBSERVATION]
- long agent loops proven [REAL_OBSERVATION]
- multi-file repo work proven [REAL_OBSERVATION]
- testing/tool loop proven [REAL_OBSERVATION]
Evidence — primary model
GLM-5.3-Flash(x) (glm-5.3-flashx)
Status: VERIFIED_FOR_MEDIUM_MULTI_FILE · last verified: 2026-10-05 · provenance: STATIC_CONFIG
Verified task classes: GIT_AUDIT, README_UPDATE, SMALL_PATCH, MEDIUM_FEATURE · Unverified: LARGE_FEATURE, SECURITY_REVIEW, UI_VISION_REVIEW
- read-only Git inventory PASS [REAL_OBSERVATION]
- untracked-file inspection PASS [REAL_OBSERVATION]
- small repo hygiene coding patch PASS [REAL_OBSERVATION]
- 531-test verification PASS [REAL_OBSERVATION]
- medium multi-file task PASS (MODEL_ROUTING_POLICY_SIMULATOR_V1: ~9 files, 286/286 portal tests, 39 new) [REAL_OBSERVATION]
- large feature / long agent loop work NOT YET VERIFIED; one medium task is not enough evidence [REAL_OBSERVATION]
SCORE (transparent, advisory)
| Dimension | Value | Static weight |
|---|---|---|
| Capability Match | 0.60 | weight 0.30 |
| Evidence Strength | 1.00 | weight 0.25 |
| Task Complexity Fit | 1.00 | weight 0.15 |
| Risk Fit | 0.80 | weight 0.10 |
| Context Fit | 1.00 | weight 0.08 |
| Cost Fit | 1.00 | weight 0.07 |
| Latency Fit | 1.00 | weight 0.03 |
| Vision Fit | 0.50 | weight 0.02 |
| TOTAL | 0.850 |
Routing Reasons
WHY THIS MODEL?
- STATUS — GLM-5.3-Flash(x) status = VERIFIED_FOR_MEDIUM_MULTI_FILE
- EVIDENCE — verified task classes: GIT_AUDIT, README_UPDATE, SMALL_PATCH, MEDIUM_FEATURE
- TASK_MATCH — direct verified evidence for task class SMALL_PATCH
- COST — low-cost candidate; evidence threshold met
- PARTIAL_SCOPE — VERIFIED_FOR_MEDIUM_MULTI_FILE: verification is scoped to evidenced task classes only; larger task classes remain unverified
- NOT_PREFERRED_deepseek-v4.1-flash — DeepSeek V4.1 Flash scored 0.701 vs 0.85 (gap 0.149)
- EXCLUDED_kimi-k3 — Kimi K3 excluded: unavailable (upstream HTTP 503 on trivial prompt; provider retry failed)
- FALLBACK — fallback DeepSeek V4.1 Flash (deepseek-v4.1-flash): VERIFIED, next-best eligible score 0.701
Hard Blockers
- none
WHAT IS STILL UNKNOWN?
- DeepSeek V4.1 Flash: vision support UNKNOWN
- GLM-5.3-Flash(x): vision support UNKNOWN
- Kimi K3: vision support UNKNOWN
- Kimi K3: cost multiplier UNKNOWN
- GLM-5.3-Flash(x): large feature / long agent loop work not yet verified
WHAT WOULD CHANGE THE DECISION?
- If GLM-5.3-Flash(x) accumulates 2-3 more successful medium multi-file tasks and comparative token/cost data, the default switch may be reconsidered (currently NOT YET APPROVED)
- If GLM-5.3-Flash(x) completes a verified LARGE_FEATURE task, it becomes eligible for large-feature routing
- If Kimi K3 becomes reachable AND accumulates verified evidence in our environment, it re-enters the candidate pool
- If this task is R3/R4, owner approval is required regardless of model choice
Diversity-of-Failure
DIVERSITY_PREFERRED QA reviewer should prefer DeepSeek V4.1 Flash (deepseek-v4.1-flash) — different failure modes from GLM-5.3-Flash(x); evidence quality preserved
MODEL ROUTING != EXECUTION AUTHORITY
Selecting a model never grants sudo, Docker mutation, production write, DB mutation, deploy, network change or approval power. Authority is governed by risk policy, owner approval and DCP controls.
Model Status (static development registry)
| Model | Status | Proven task classes | Known failures | Last verification |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | VERIFIED | MEDIUM_FEATURE, LARGE_FEATURE, SMALL_PATCH, CODE_REVIEW, GIT_AUDIT, README_UPDATE | — | 2026-10-05 |
| GLM-5.3-Flash(x) | VERIFIED_FOR_MEDIUM_MULTI_FILE | GIT_AUDIT, README_UPDATE, SMALL_PATCH, MEDIUM_FEATURE | — | 2026-10-05 |
| Kimi K3 | UNAVAILABLE | upstream HTTP 503 on trivial prompt; provider retry failed | ||