DCP READ-ONLY: CONNECTED DEVELOPMENT / UI PREVIEW DCP BATCH 5/6 SOAK IN PROGRESS TELEMETRY: UNKNOWN COMMAND EXECUTION LOCKED
READ-ONLY SIMULATOR · Deterministic policy evaluation over a STATIC development registry. No live model calls, no routing changes, no execution. Recommendations are advisory; authority stays with the owner.
Task Input (preset scenario)
4 · Large repo feature
| Task type | LARGE_FEATURE |
|---|---|
| Agent role | DEVELOPER |
| Risk level | R2 |
| Complexity | HIGH |
| Context size | LARGE |
| Vision | NONE |
| Repo write | yes |
| Long tool loop | yes |
| Architecture reasoning | no |
| Code generation | yes |
| Security review | no |
| Cost sensitivity | MEDIUM |
| Latency sensitivity | LOW |
Preset expectation: DeepSeek preferred.
Recommended Agent
DEVELOPER
Static default policy: coding evidence dominates · multi-file/long-loop evidence required for larger tasks
Primary Model
DeepSeek V4.1 Flash deepseek-v4.1-flash VERIFIED
Fallback Model
none eligible
Evidence — primary model
DeepSeek V4.1 Flash (deepseek-v4.1-flash)
Status: VERIFIED · last verified: 2026-10-05 · provenance: STATIC_CONFIG
Verified task classes: MEDIUM_FEATURE, LARGE_FEATURE, SMALL_PATCH, CODE_REVIEW, GIT_AUDIT, README_UPDATE · Unverified: UI_VISION_REVIEW
- heavy coding proven [REAL_OBSERVATION]
- long agent loops proven [REAL_OBSERVATION]
- multi-file repo work proven [REAL_OBSERVATION]
- testing/tool loop proven [REAL_OBSERVATION]
SCORE (transparent, advisory)
The weighted score is advisory only. Hard blockers always override the score. UNKNOWN never counts as positive evidence.
| Dimension | Value | Static weight |
|---|---|---|
| Capability Match | 0.60 | weight 0.30 |
| Evidence Strength | 1.00 | weight 0.25 |
| Task Complexity Fit | 1.00 | weight 0.15 |
| Risk Fit | 1.00 | weight 0.10 |
| Context Fit | 1.00 | weight 0.08 |
| Cost Fit | 0.72 | weight 0.07 |
| Latency Fit | 0.80 | weight 0.03 |
| Vision Fit | 0.50 | weight 0.02 |
| TOTAL | 0.844 |
Routing Reasons
WHY THIS MODEL?
- STATUS — DeepSeek V4.1 Flash status = VERIFIED
- EVIDENCE — verified task classes: MEDIUM_FEATURE, LARGE_FEATURE, SMALL_PATCH, CODE_REVIEW, GIT_AUDIT, README_UPDATE
- TASK_MATCH — direct verified evidence for task class LARGE_FEATURE
- MULTI_FILE — proven multi-file/long-loop evidence for repo implementation
- NOT_PREFERRED_glm-5.3-flashx — GLM-5.3-Flash(x) not preferred: hard-blocked (PARTIAL_NO_TASK_EVIDENCE)
- EXCLUDED_kimi-k3 — Kimi K3 excluded: unavailable (upstream HTTP 503 on trivial prompt; provider retry failed)
Hard Blockers
- none
WHAT IS STILL UNKNOWN?
- DeepSeek V4.1 Flash: vision support UNKNOWN
- GLM-5.3-Flash(x): vision support UNKNOWN
- Kimi K3: vision support UNKNOWN
- Kimi K3: cost multiplier UNKNOWN
WHAT WOULD CHANGE THE DECISION?
- If a cheaper model accumulates verified evidence for this task class, cost preference may flip the primary
- If Kimi K3 becomes reachable AND accumulates verified evidence in our environment, it re-enters the candidate pool
- If this task is R3/R4, owner approval is required regardless of model choice
Diversity-of-Failure
DIVERSITY_NOT_AVAILABLE no alternative VERIFIED/PARTIALLY_VERIFIED model qualifies; QA reviewer uses the primary model
MODEL ROUTING != EXECUTION AUTHORITY
Selecting a model never grants sudo, Docker mutation, production write, DB mutation, deploy, network change or approval power. Authority is governed by risk policy, owner approval and DCP controls.
Model Status (static development registry)
Kimi K3 is UNAVAILABLE (HTTP 503 upstream provider failure) and is never displayed as selectable or recommended.
| Model | Status | Proven task classes | Known failures | Last verification |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | VERIFIED | MEDIUM_FEATURE, LARGE_FEATURE, SMALL_PATCH, CODE_REVIEW, GIT_AUDIT, README_UPDATE | — | 2026-10-05 |
| GLM-5.3-Flash(x) | VERIFIED_FOR_MEDIUM_MULTI_FILE | GIT_AUDIT, README_UPDATE, SMALL_PATCH, MEDIUM_FEATURE | — | 2026-10-05 |
| Kimi K3 | UNAVAILABLE | upstream HTTP 503 on trivial prompt; provider retry failed | ||