jcbtc's picture
Add model card
a1ab828 verified
|
Raw History Blame Contribute Delete
8.8 kB

Packaged gate and retained quality evidence

256K DFlash depth and cache-safety investigation

The provisional production recipe uses DFlash2 k=5 with prefix caching disabled. A separate local inference workload was later found to have remained resident during the original target-only, k=3, k=5, and HumanEval speed measurements. Those rows are retained as directional, non-isolated evidence. The HumanEval 0–9 gate's 10/10 pass@1 result remains valid.

Two follow-up repetitions of the 65,680-token k=5 stress request cleared the campaign's minimum generation floor without unloading or restarting the GLM pair. This is a long-context stress result, not representative short-prompt chat throughput. A fully isolated depth comparison is deferred to a maintenance window. Full protocol and findings.

The packaged 64K NHI functional gate passed with useful disk-prefix reuse on dev8. Its one-token cache replay is separate from the historical quality results below; it is not a new TG or quality benchmark. The 128K production default passed the short-prompt sanity check below on dev8 and the separate 131,000-token generation/restart/reuse gate on packaged dev9. Both ranks reused 129,024 cached tokens with actual GPU loads and no crash or rank divergence. These are bounded results, not a full quality sweep. The 256K profile is experimental.

Production-chat HumanEval 0–9 sanity, 128K profile

The packaged NHI/DFlash2 k=7 production path ran HumanEval tasks 0–9 through the mirrored chat frontend with the official chat template, temperature 1.0, top-p 0.95, reasoning effort max, and an 8,192-token completion cap. The v3 run was scored once: 10/10 passed, all natural stops, with no caps, exact-loop aborts, retries, or selection among repeated answers. This run used dev8, before the subsequent dev9 cache repairs.

Measurement Result
Tasks / passed 10 / 10
Total prompt / new output tokens 1,537 / 7,892
Sum TTFT / mean TTFT 13.293937734 s / 1.3294 s
Total decode time 319.626816043 s
Weighted decode throughput 24.6600 tok/s
Accepted / proposed draft tokens 6,352 / 10,745 (59.1159%)
Verification steps / cadence 1,535 / 4.80248 per second
Total client wall time 333.041892 s

Weighted decode throughput is (7892 - 10) / 319.626816043, excluding each request's first token and its prefill interval. Cadence is 1535 / 319.626816043; neither is an unweighted average of request rates. The saved responses and one-pass scorer output retain the evidence.

This is a production-chat-adapted 10-task sanity check, not full HumanEval 164 or a canonical leaderboard score. These are short prompts in a 128K configured engine, not a near-128K input test. The separate capacity and persistent-prefix write/restart/read gate below passed and carries no quality-score claim.

Near-128K generation and persistent-cache reuse

The 131,000-token cold request completed 128 output tokens: TTFT 377.2719 s, effective prompt throughput approximately 347.23 tok/s, and client wall time 390.6911 s. This is a capacity/performance observation, not a quality score or successful persistent-cache restart/reuse result: one rank's disk writes reached ENOSPC.

The initial restore fault was fixed in dev9's Mamba cached-state indexing and paired admission paths. After repairing incomplete cache writes and restarting both packaged dev9 ranks, the exact 131,000-token request and salt reused 129,024 cached tokens per rank (98.49%), with 7.568049367 s TTFT and 10.713744593 s client wall time. The answer naturally stopped at READY with 24 output tokens, including 21 reasoning tokens. Both ranks recorded filesystem reads and actual GPU loads, without crash or rank divergence. The versioned gate report links the saved response and rank metrics. This single-prompt functional result is not a quality score or steady-state TG benchmark.

Free-space changes indicated cache storage on the order of 100 GiB or more per rank; the exact cache footprint was not separately measured. The 8 GiB setting limits RAM staging only. The filesystem tier has no built-in quota or automatic eviction; allow substantial dedicated NVMe headroom and monitor its growth. See the disk-cache instructions.

Historical quality evidence

These public-safe summaries package existing campaign reports. No evaluation was rerun to create them. They characterize the historical production target shape and serving configuration, not a new measurement of the final packaged wheel, persistent prefix cache, new 128K default, or experimental 256K profile.

Evidence Retained result Summary
ToolEval Standard69 120/138 points; 57 pass, 6 partial, 6 fail standard69.json
ToolEval Hard15 25/30 points; 11 pass, 3 partial, 1 fail hard15.json
Matched BF16 comparison 767 positions; PPL 2.148624 vs 2.261154; top-1 agreement 91.917% bf16-quality.json

ToolEval scope and configuration

The saved reports identify tool-eval-bench 2.1.0, one trial per scenario, and completed runs on 2026-09-02 UTC. The benchmark's project metadata names SeraphimSerapis/tool-eval-bench. The campaign used a local adapter for mirrored TP2 requests and structured responses; the exact benchmark source revision is not recorded in the reports. These results should not be assumed identical to another revision or an unmodified upstream installation.

Standard69 covers TC-01 through TC-69. Hard15 is a separate TC-70 through TC-84 suite, not an additional trial of Standard69. Each scenario earns two points for pass, one for partial, or zero for fail. The harness reports round(100 * total_points / max_points): 87% and 83% are its rounded scores, not exact point ratios and not pass-only rates. The exact ratios are 120/138 and 25/30.

tooleval-config.json separates the settings recorded in both result files from the saved server/runner recipe. The served model alias was glm53-f1-dflash2-production; it is retained as provenance rather than rewritten to suggest the public package was evaluated under its new name. The saved server recipe disabled prefix caching and used the no-K3 hybrid W4 target with DFlash2 k=7. These scores do not establish disk-cache behavior.

The category totals are preserved in the summaries. In Standard69, Safety & Boundaries scored 18/26 and Toolset Scale 4/8. The three safety-warning scenario IDs are retained, but raw conversations, prompts, tool arguments, and private endpoint details are not included.

BF16 comparison scope

The teacher was zai-org/GLM-5.3-Flash-BF16 at revision f12e0fe1f6b2ea274c11a569582edfd99d993c5e. The candidate was the production no-K3 N32/G128 W4 prepack with the dense F1 and native IU4 large-M paths.

This is one matched block: the first 768 teacher-tokenizer tokens of the retained wiki.test.raw, with 767 next-token positions and all 154,880 vocabulary logits scored. The scoring program checked equality of input token IDs, target token IDs, and logit tensor shapes before comparing distributions. The candidate quality arm used RCCL; it was not a packaged NHI performance run. The BF16 teacher used layer streaming. Neither arm's collection timing is presented as production inference throughput.

The JSON preserves the saved numeric precision for PPL, NLL, agreement, forward and reverse KL, Jensen-Shannon divergence, and top-k set recall. Agreement and recall values are fractions; KL, NLL, and Jensen-Shannon values use nats. This narrow quantization characterization is not a full WikiText evaluation, a broad capability benchmark, or a ranking against other quantizations.

Retention and limits

The ToolEval summaries identify their original report basenames and saved run IDs. The BF16 summary identifies its retained result schema and protocol. Only public-safe aggregate fields and scenario outcome IDs are repackaged. Private hostnames, absolute paths, endpoint addresses, raw prompts/responses, and per-layer storage timings are intentionally omitted.

The saved ToolEval reports do not bind the run to an exact public artifact or benchmark source commit. The retained BF16 report does not record a corpus revision or checksum. This folder therefore provides the saved results and their known configuration, not a fully self-contained reproduction bundle. No new hashes, quality runs, or tests were performed while staging it.