Model Benchmark Scorecard

Recorded observations only · no estimated tokens or costs · no superiority claims without comparative evidence
DCP READ-ONLY: CONNECTED DEVELOPMENT / UI PREVIEW DCP BATCH 5/6 SOAK IN PROGRESS TELEMETRY: UNKNOWN COMMAND EXECUTION LOCKED
READ-ONLY SCORECARD · Recorded observations only — no marketing benchmark data, no estimated tokens, no invented costs. Unknown values stay UNKNOWN. A single model passing a task proves nothing about another model.

GLM-5.3-Flash(x) vs DeepSeek V4.1 Flash

INSUFFICIENT
INSUFFICIENT COMPARATIVE TOKEN DATA — no comparable token/cost measurements exist for both models, so no performance, efficiency or superiority comparison is shown. Neither model is displayed as "better".

Comparative claims require comparable task evidence on both sides. DeepSeek V4.1 Flash remains the verified default for heavy coding / long loops / large repo work; GLM-5.3-Flash(x) is verified for medium multi-file work only (one observation).

Successful tasks by model

GLM-5.3-Flash(x) (glm-5.3-flashx) VERIFIED_FOR_MEDIUM_MULTI_FILE

Verified task classes: GIT_AUDIT, README_UPDATE, SMALL_PATCH, MEDIUM_FEATURE · Evidence confidence: REAL_OBSERVATION (committed evidence artifacts / project docs)

Successful taskTask classOutcomeDurationToken usageCostRetry countProvenance
READ-ONLY GIT INVENTORYGIT_AUDITPASSUNKNOWNUNKNOWNUNKNOWNUNKNOWNREAL_OBSERVATION
DCP UNTRACKED FILE INSPECTIONGIT_AUDITPASSUNKNOWNUNKNOWNUNKNOWNUNKNOWNREAL_OBSERVATION
DCP REPO HYGIENE PATCHSMALL_PATCHPASSUNKNOWNUNKNOWNUNKNOWNUNKNOWNREAL_OBSERVATION
MODEL_ROUTING_POLICY_SIMULATOR_V1MEDIUM_MULTI_FILEPASSUNKNOWNUNKNOWNUNKNOWNUNKNOWN (one stream retry observed but not reliably recorded)REAL_OBSERVATION

DeepSeek V4.1 Flash (deepseek-v4.1-flash) VERIFIED

Verified task classes: MEDIUM_FEATURE, LARGE_FEATURE, SMALL_PATCH, CODE_REVIEW, GIT_AUDIT, README_UPDATE · Evidence confidence: REAL_OBSERVATION (committed evidence artifacts / project docs)

Successful taskTask classOutcomeDurationToken usageCostRetry countProvenance
heavy coding / long agent loops / repo-level workLARGE_FEATUREPASSUNKNOWNUNKNOWNUNKNOWNUNKNOWNREAL_OBSERVATION

Observation details

READ-ONLY GIT INVENTORY (GIT_AUDIT) — PASS
OutcomePASS PASS
Files changedUNKNOWN
Approx lines changedUNKNOWN
Tests addedUNKNOWN
Tests totalUNKNOWN
Tests resultUNKNOWN
Retry countUNKNOWN
DurationUNKNOWN
Input tokensUNKNOWN
Output tokensUNKNOWN
Cached tokensUNKNOWN
Total tokensUNKNOWN
Cost multiplierUNKNOWN
Effective costUNKNOWN
Instruction followingPASS PASS
Evidence disciplinePASS PASS
Tool loop stabilityUNKNOWN
Self correctionUNKNOWN
Notes—
ProvenanceREAL_OBSERVATION
Sourcehomelab-infra/docs/MODEL-ROUTING-POLICY.md (registry evidence)
DCP UNTRACKED FILE INSPECTION (GIT_AUDIT) — PASS
OutcomePASS PASS
Files changedUNKNOWN
Approx lines changedUNKNOWN
Tests addedUNKNOWN
Tests totalUNKNOWN
Tests resultUNKNOWN
Retry countUNKNOWN
DurationUNKNOWN
Input tokensUNKNOWN
Output tokensUNKNOWN
Cached tokensUNKNOWN
Total tokensUNKNOWN
Cost multiplierUNKNOWN
Effective costUNKNOWN
Instruction followingPASS PASS
Evidence disciplinePASS PASS
Tool loop stabilityUNKNOWN
Self correctionUNKNOWN
Notes—
ProvenanceREAL_OBSERVATION
Sourcehomelab-infra/docs/MODEL-ROUTING-POLICY.md (registry evidence)
DCP REPO HYGIENE PATCH (SMALL_PATCH) — PASS
OutcomePASS PASS
Files changedUNKNOWN
Approx lines changedUNKNOWN
Tests addedUNKNOWN
Tests total531
Tests result531 PASS
Retry countUNKNOWN
DurationUNKNOWN
Input tokensUNKNOWN
Output tokensUNKNOWN
Cached tokensUNKNOWN
Total tokensUNKNOWN
Cost multiplierUNKNOWN
Effective costUNKNOWN
Instruction followingPASS PASS
Evidence disciplinePASS PASS
Tool loop stabilityUNKNOWN
Self correctionUNKNOWN
Notes—
ProvenanceREAL_OBSERVATION
Sourcehomelab-infra/docs/MODEL-ROUTING-POLICY.md (registry evidence)
MODEL_ROUTING_POLICY_SIMULATOR_V1 (MEDIUM_MULTI_FILE) — PASS
OutcomePASS PASS
Files changed~9
Approx lines changed~1.9k
Tests added39
Tests total286
Tests result286/286 PASS
Retry countUNKNOWN (one stream retry observed but not reliably recorded)
DurationUNKNOWN
Input tokensUNKNOWN
Output tokensUNKNOWN
Cached tokensUNKNOWN
Total tokensUNKNOWN
Cost multiplierLOW (static class, not pricing)
Effective costUNKNOWN
Instruction followingPASS (one routing-rule refinement cycle, self-detected and fixed before commit)
Evidence disciplinePASS (evidence artifact committed with task)
Tool loop stabilityPASS (completed end-to-end)
Self correctionOBSERVED (one self-run refinement cycle)
NotesEvidence supports VERIFIED_FOR_MEDIUM_MULTI_FILE only; does NOT support VERIFIED_FOR_LARGE_FEATURE.
ProvenanceREAL_OBSERVATION
Sourcehomelab-infra/inventory/model-evidence.json
heavy coding / long agent loops / repo-level work (LARGE_FEATURE) — PASS
OutcomePASS PASS
Files changedUNKNOWN
Approx lines changedUNKNOWN
Tests addedUNKNOWN
Tests totalUNKNOWN
Tests resultUNKNOWN
Retry countUNKNOWN
DurationUNKNOWN
Input tokensUNKNOWN
Output tokensUNKNOWN
Cached tokensUNKNOWN
Total tokensUNKNOWN
Cost multiplierUNKNOWN
Effective costUNKNOWN
Instruction followingUNKNOWN
Evidence disciplineUNKNOWN
Tool loop stabilityUNKNOWN
Self correctionUNKNOWN
NotesVerified status based on task evidence already present in project history/docs; no comparative run against GLM.
ProvenanceREAL_OBSERVATION
Sourceproject history/docs (registry evidence; see MODEL-ROUTING-POLICY.md)

Advisory model decisions (CURRENT_RECOMMENDED_POLICY)

Task classAdvisory decision
Routine Git / docs / read-onlyGLM eligible/preferred (evidence supports)
Small patchesGLM eligible
Medium multi-fileGLM eligible · DeepSeek eligible/proven
Large feature / very long agent loopDeepSeek preferred until GLM gains large-feature evidence
Kimi K3Excluded while UNAVAILABLE (HTTP 503)
VisionUNKNOWN / MANUAL unless verified evidence exists

Default model switch: NOT YET APPROVED. Criteria: 2-3 successful medium multi-file GLM tasks · no major instruction-following failures · comparative token/cost data if cost efficiency is part of the decision · one controlled heavier task or clear fallback policy.

Fields with no recorded data are rendered as UNKNOWN and are never estimated. Token counts are never inferred from elapsed time. Request count is not token count. See model-usage-schema.json in homelab-infra for the telemetry schema that will fill these fields with real measurements.