GLM-5.3-Flash(x) vs DeepSeek V4.1 Flash
INSUFFICIENTComparative claims require comparable task evidence on both sides. DeepSeek V4.1 Flash remains the verified default for heavy coding / long loops / large repo work; GLM-5.3-Flash(x) is verified for medium multi-file work only (one observation).
Successful tasks by model
GLM-5.3-Flash(x) (glm-5.3-flashx) VERIFIED_FOR_MEDIUM_MULTI_FILE
Verified task classes: GIT_AUDIT, README_UPDATE, SMALL_PATCH, MEDIUM_FEATURE · Evidence confidence: REAL_OBSERVATION (committed evidence artifacts / project docs)
| Successful task | Task class | Outcome | Duration | Token usage | Cost | Retry count | Provenance |
|---|---|---|---|---|---|---|---|
| READ-ONLY GIT INVENTORY | GIT_AUDIT | PASS | UNKNOWN | UNKNOWN | UNKNOWN | UNKNOWN | REAL_OBSERVATION |
| DCP UNTRACKED FILE INSPECTION | GIT_AUDIT | PASS | UNKNOWN | UNKNOWN | UNKNOWN | UNKNOWN | REAL_OBSERVATION |
| DCP REPO HYGIENE PATCH | SMALL_PATCH | PASS | UNKNOWN | UNKNOWN | UNKNOWN | UNKNOWN | REAL_OBSERVATION |
| MODEL_ROUTING_POLICY_SIMULATOR_V1 | MEDIUM_MULTI_FILE | PASS | UNKNOWN | UNKNOWN | UNKNOWN | UNKNOWN (one stream retry observed but not reliably recorded) | REAL_OBSERVATION |
DeepSeek V4.1 Flash (deepseek-v4.1-flash) VERIFIED
Verified task classes: MEDIUM_FEATURE, LARGE_FEATURE, SMALL_PATCH, CODE_REVIEW, GIT_AUDIT, README_UPDATE · Evidence confidence: REAL_OBSERVATION (committed evidence artifacts / project docs)
| Successful task | Task class | Outcome | Duration | Token usage | Cost | Retry count | Provenance |
|---|---|---|---|---|---|---|---|
| heavy coding / long agent loops / repo-level work | LARGE_FEATURE | PASS | UNKNOWN | UNKNOWN | UNKNOWN | UNKNOWN | REAL_OBSERVATION |
Observation details
READ-ONLY GIT INVENTORY (GIT_AUDIT) — PASS
| Outcome | PASS PASS |
|---|---|
| Files changed | UNKNOWN |
| Approx lines changed | UNKNOWN |
| Tests added | UNKNOWN |
| Tests total | UNKNOWN |
| Tests result | UNKNOWN |
| Retry count | UNKNOWN |
| Duration | UNKNOWN |
| Input tokens | UNKNOWN |
| Output tokens | UNKNOWN |
| Cached tokens | UNKNOWN |
| Total tokens | UNKNOWN |
| Cost multiplier | UNKNOWN |
| Effective cost | UNKNOWN |
| Instruction following | PASS PASS |
| Evidence discipline | PASS PASS |
| Tool loop stability | UNKNOWN |
| Self correction | UNKNOWN |
| Notes | — |
| Provenance | REAL_OBSERVATION |
| Source | homelab-infra/docs/MODEL-ROUTING-POLICY.md (registry evidence) |
DCP UNTRACKED FILE INSPECTION (GIT_AUDIT) — PASS
| Outcome | PASS PASS |
|---|---|
| Files changed | UNKNOWN |
| Approx lines changed | UNKNOWN |
| Tests added | UNKNOWN |
| Tests total | UNKNOWN |
| Tests result | UNKNOWN |
| Retry count | UNKNOWN |
| Duration | UNKNOWN |
| Input tokens | UNKNOWN |
| Output tokens | UNKNOWN |
| Cached tokens | UNKNOWN |
| Total tokens | UNKNOWN |
| Cost multiplier | UNKNOWN |
| Effective cost | UNKNOWN |
| Instruction following | PASS PASS |
| Evidence discipline | PASS PASS |
| Tool loop stability | UNKNOWN |
| Self correction | UNKNOWN |
| Notes | — |
| Provenance | REAL_OBSERVATION |
| Source | homelab-infra/docs/MODEL-ROUTING-POLICY.md (registry evidence) |
DCP REPO HYGIENE PATCH (SMALL_PATCH) — PASS
| Outcome | PASS PASS |
|---|---|
| Files changed | UNKNOWN |
| Approx lines changed | UNKNOWN |
| Tests added | UNKNOWN |
| Tests total | 531 |
| Tests result | 531 PASS |
| Retry count | UNKNOWN |
| Duration | UNKNOWN |
| Input tokens | UNKNOWN |
| Output tokens | UNKNOWN |
| Cached tokens | UNKNOWN |
| Total tokens | UNKNOWN |
| Cost multiplier | UNKNOWN |
| Effective cost | UNKNOWN |
| Instruction following | PASS PASS |
| Evidence discipline | PASS PASS |
| Tool loop stability | UNKNOWN |
| Self correction | UNKNOWN |
| Notes | — |
| Provenance | REAL_OBSERVATION |
| Source | homelab-infra/docs/MODEL-ROUTING-POLICY.md (registry evidence) |
MODEL_ROUTING_POLICY_SIMULATOR_V1 (MEDIUM_MULTI_FILE) — PASS
| Outcome | PASS PASS |
|---|---|
| Files changed | ~9 |
| Approx lines changed | ~1.9k |
| Tests added | 39 |
| Tests total | 286 |
| Tests result | 286/286 PASS |
| Retry count | UNKNOWN (one stream retry observed but not reliably recorded) |
| Duration | UNKNOWN |
| Input tokens | UNKNOWN |
| Output tokens | UNKNOWN |
| Cached tokens | UNKNOWN |
| Total tokens | UNKNOWN |
| Cost multiplier | LOW (static class, not pricing) |
| Effective cost | UNKNOWN |
| Instruction following | PASS (one routing-rule refinement cycle, self-detected and fixed before commit) |
| Evidence discipline | PASS (evidence artifact committed with task) |
| Tool loop stability | PASS (completed end-to-end) |
| Self correction | OBSERVED (one self-run refinement cycle) |
| Notes | Evidence supports VERIFIED_FOR_MEDIUM_MULTI_FILE only; does NOT support VERIFIED_FOR_LARGE_FEATURE. |
| Provenance | REAL_OBSERVATION |
| Source | homelab-infra/inventory/model-evidence.json |
heavy coding / long agent loops / repo-level work (LARGE_FEATURE) — PASS
| Outcome | PASS PASS |
|---|---|
| Files changed | UNKNOWN |
| Approx lines changed | UNKNOWN |
| Tests added | UNKNOWN |
| Tests total | UNKNOWN |
| Tests result | UNKNOWN |
| Retry count | UNKNOWN |
| Duration | UNKNOWN |
| Input tokens | UNKNOWN |
| Output tokens | UNKNOWN |
| Cached tokens | UNKNOWN |
| Total tokens | UNKNOWN |
| Cost multiplier | UNKNOWN |
| Effective cost | UNKNOWN |
| Instruction following | UNKNOWN |
| Evidence discipline | UNKNOWN |
| Tool loop stability | UNKNOWN |
| Self correction | UNKNOWN |
| Notes | Verified status based on task evidence already present in project history/docs; no comparative run against GLM. |
| Provenance | REAL_OBSERVATION |
| Source | project history/docs (registry evidence; see MODEL-ROUTING-POLICY.md) |
Advisory model decisions (CURRENT_RECOMMENDED_POLICY)
| Task class | Advisory decision |
|---|---|
| Routine Git / docs / read-only | GLM eligible/preferred (evidence supports) |
| Small patches | GLM eligible |
| Medium multi-file | GLM eligible · DeepSeek eligible/proven |
| Large feature / very long agent loop | DeepSeek preferred until GLM gains large-feature evidence |
| Kimi K3 | Excluded while UNAVAILABLE (HTTP 503) |
| Vision | UNKNOWN / MANUAL unless verified evidence exists |
Default model switch: NOT YET APPROVED. Criteria: 2-3 successful medium multi-file GLM tasks · no major instruction-following failures · comparative token/cost data if cost efficiency is part of the decision · one controlled heavier task or clear fallback policy.
Fields with no recorded data are rendered as UNKNOWN and are never estimated. Token counts are never inferred from elapsed time. Request count is not token count. See model-usage-schema.json in homelab-infra for the telemetry schema that will fill these fields with real measurements.