Download benchmarks/README.md from jcbtc/GLM5.3-Flash-CIRU-STRIX-IU4: direct link, hf CLI and curl.
- Browser
- Download file 8.8 kB
-
https://huggingface.co/jcbtc/GLM5.3-Flash-CIRU-STRIX-IU4/resolve/main/benchmarks/README.md
- Command line
-
hf download hf://jcbtc/GLM5.3-Flash-CIRU-STRIX-IU4/benchmarks/README.md
-
curl -L -o README.md https://huggingface.co/jcbtc/GLM5.3-Flash-CIRU-STRIX-IU4/resolve/main/benchmarks/README.md
Packaged gate and retained quality evidence
256K DFlash depth and cache-safety investigation
The provisional production recipe uses DFlash2 k=5 with prefix caching disabled. A separate local inference workload was later found to have remained resident during the original target-only, k=3, k=5, and HumanEval speed measurements. Those rows are retained as directional, non-isolated evidence. The HumanEval 0–9 gate's 10/10 pass@1 result remains valid.
Two follow-up repetitions of the 65,680-token k=5 stress request cleared the campaign's minimum generation floor without unloading or restarting the GLM pair. This is a long-context stress result, not representative short-prompt chat throughput. A fully isolated depth comparison is deferred to a maintenance window. Full protocol and findings.
The packaged 64K NHI functional gate passed with useful disk-prefix reuse on dev8. Its one-token cache replay is separate from the historical quality results below; it is not a new TG or quality benchmark. The 128K production default passed the short-prompt sanity check below on dev8 and the separate 131,000-token generation/restart/reuse gate on packaged dev9. Both ranks reused 129,024 cached tokens with actual GPU loads and no crash or rank divergence. These are bounded results, not a full quality sweep. The 256K profile is experimental.
Production-chat HumanEval 0–9 sanity, 128K profile
The packaged NHI/DFlash2 k=7 production path ran HumanEval tasks 0–9 through the mirrored chat frontend with the official chat template, temperature 1.0, top-p 0.95, reasoning effort max, and an 8,192-token completion cap. The v3 run was scored once: 10/10 passed, all natural stops, with no caps, exact-loop aborts, retries, or selection among repeated answers. This run used dev8, before the subsequent dev9 cache repairs.
| Measurement | Result |
|---|---|
| Tasks / passed | 10 / 10 |
| Total prompt / new output tokens | 1,537 / 7,892 |
| Sum TTFT / mean TTFT | 13.293937734 s / 1.3294 s |
| Total decode time | 319.626816043 s |
| Weighted decode throughput | 24.6600 tok/s |
| Accepted / proposed draft tokens | 6,352 / 10,745 (59.1159%) |
| Verification steps / cadence | 1,535 / 4.80248 per second |
| Total client wall time | 333.041892 s |
Weighted decode throughput is (7892 - 10) / 319.626816043, excluding each
request's first token and its prefill interval. Cadence is
1535 / 319.626816043; neither is an unweighted average of request rates.
The saved responses and
one-pass scorer output retain the evidence.
This is a production-chat-adapted 10-task sanity check, not full HumanEval 164 or a canonical leaderboard score. These are short prompts in a 128K configured engine, not a near-128K input test. The separate capacity and persistent-prefix write/restart/read gate below passed and carries no quality-score claim.
Near-128K generation and persistent-cache reuse
The 131,000-token cold request completed 128 output tokens: TTFT 377.2719 s,
effective prompt throughput approximately 347.23 tok/s, and client wall
time 390.6911 s. This is a capacity/performance observation, not a quality
score or successful persistent-cache restart/reuse result: one rank's disk
writes reached ENOSPC.
The initial restore fault was fixed in dev9's Mamba cached-state indexing and
paired admission paths. After repairing incomplete cache writes and restarting
both packaged dev9 ranks, the exact 131,000-token request and salt reused
129,024 cached tokens per rank (98.49%), with 7.568049367 s TTFT and
10.713744593 s client wall time. The answer naturally stopped at READY
with 24 output tokens, including 21 reasoning tokens. Both ranks recorded
filesystem reads and actual GPU loads, without crash or rank divergence.
The versioned gate report links the saved response
and rank metrics. This single-prompt functional result is not a quality score
or steady-state TG benchmark.
Free-space changes indicated cache storage on the order of 100 GiB or more per rank; the exact cache footprint was not separately measured. The 8 GiB setting limits RAM staging only. The filesystem tier has no built-in quota or automatic eviction; allow substantial dedicated NVMe headroom and monitor its growth. See the disk-cache instructions.
Historical quality evidence
These public-safe summaries package existing campaign reports. No evaluation was rerun to create them. They characterize the historical production target shape and serving configuration, not a new measurement of the final packaged wheel, persistent prefix cache, new 128K default, or experimental 256K profile.
| Evidence | Retained result | Summary |
|---|---|---|
| ToolEval Standard69 | 120/138 points; 57 pass, 6 partial, 6 fail | standard69.json |
| ToolEval Hard15 | 25/30 points; 11 pass, 3 partial, 1 fail | hard15.json |
| Matched BF16 comparison | 767 positions; PPL 2.148624 vs 2.261154; top-1 agreement 91.917% | bf16-quality.json |
ToolEval scope and configuration
The saved reports identify tool-eval-bench 2.1.0, one trial per scenario,
and completed runs on 2026-09-02 UTC. The benchmark's project metadata names
SeraphimSerapis/tool-eval-bench.
The campaign used a local adapter for mirrored TP2 requests and structured
responses; the exact benchmark source revision is not recorded in the reports.
These results should not be assumed identical to another revision or an
unmodified upstream installation.
Standard69 covers TC-01 through TC-69. Hard15 is a separate TC-70 through
TC-84 suite, not an additional trial of Standard69. Each scenario earns two
points for pass, one for partial, or zero for fail. The harness reports
round(100 * total_points / max_points): 87% and 83% are its rounded
scores, not exact point ratios and not pass-only rates. The exact ratios are
120/138 and 25/30.
tooleval-config.json separates the settings recorded
in both result files from the saved server/runner recipe. The served model
alias was glm53-f1-dflash2-production; it is retained as provenance rather
than rewritten to suggest the public package was evaluated under its new name.
The saved server recipe disabled prefix caching and used the no-K3 hybrid W4
target with DFlash2 k=7. These scores do not establish disk-cache behavior.
The category totals are preserved in the summaries. In Standard69, Safety & Boundaries scored 18/26 and Toolset Scale 4/8. The three safety-warning scenario IDs are retained, but raw conversations, prompts, tool arguments, and private endpoint details are not included.
BF16 comparison scope
The teacher was zai-org/GLM-5.3-Flash-BF16 at revision
f12e0fe1f6b2ea274c11a569582edfd99d993c5e. The candidate was the production
no-K3 N32/G128 W4 prepack with the dense F1 and native IU4 large-M paths.
This is one matched block: the first 768 teacher-tokenizer tokens of the
retained wiki.test.raw, with 767 next-token positions and all 154,880
vocabulary logits scored. The scoring program checked equality of input token
IDs, target token IDs, and logit tensor shapes before comparing distributions.
The candidate quality arm used RCCL; it was not a packaged NHI performance run.
The BF16 teacher used layer streaming. Neither arm's collection timing is
presented as production inference throughput.
The JSON preserves the saved numeric precision for PPL, NLL, agreement, forward and reverse KL, Jensen-Shannon divergence, and top-k set recall. Agreement and recall values are fractions; KL, NLL, and Jensen-Shannon values use nats. This narrow quantization characterization is not a full WikiText evaluation, a broad capability benchmark, or a ranking against other quantizations.
Retention and limits
The ToolEval summaries identify their original report basenames and saved run IDs. The BF16 summary identifies its retained result schema and protocol. Only public-safe aggregate fields and scenario outcome IDs are repackaged. Private hostnames, absolute paths, endpoint addresses, raw prompts/responses, and per-layer storage timings are intentionally omitted.
The saved ToolEval reports do not bind the run to an exact public artifact or benchmark source commit. The retained BF16 report does not record a corpus revision or checksum. This folder therefore provides the saved results and their known configuration, not a fully self-contained reproduction bundle. No new hashes, quality runs, or tests were performed while staging it.