Skip to main content
MZNLLM Systems
Navigate
MZN · LLM Systems
Phase 2 · Evidence-aware capability review

21 capability slots.
554 canonical nodes.
One challengeable position map.

The Canonical LLM Company Anatomy is a neutral reference baseline. MZN's position layer is separate: an evidence-informed estimate of capability coverage, maturity and validation status across that baseline—not a claim of frontier-lab parity, production readiness or independent certification.

21canonical capability slots
554total canonical nodes
533descendant endpoints
4separate review dimensions
Read before scoring

This page is a review surface, not the full corpus.

MZN's technical and research corpus is materially larger than can be reproduced inside one page, PDF, pitch deck or review session. The current assessment therefore uses the structured corpus, inventories, representative deep artifacts, selected implementation/internal-test records and backing material identified for controlled review. Evidence is surfaced according to the question being tested.

Not reproduced here does not mean absent.A public surface cannot carry every source file, test record, security-sensitive mechanism or historical artifact. Reviewers can request the relevant capability family or backing file set.
Unsurfaced or restricted evidence is not independent validation.Availability for controlled review and external reproduction are separate states. Phase 3 is where selected claims are challenged, reproduced, rejected, refined or validated.
Assessment method

Four dimensions prevent one percentage from pretending to mean everything.

Capability Coverage is an indicative synthesis value. It is not a literal audited ratio of completed endpoints and should not be read as “percent of a frontier lab built.”

01

Capability Coverage

How much of the canonical domain is represented across the mapped corpus and known backing material.

02

Maturity

Documented → architecture → implemented → internally tested → operating. Maturity remains asset-specific.

03

Validation

Internal evidence is separated from independent benchmark, reproduction, legal/IP review and institutional test.

04

Disclosure

Public, controlled, restricted and reserved evidence states determine access—not technical truth by themselves.

Current 21-slot position map · 2026-08-31

Coverage is high in several domains; validation remains deliberately separate.

These values supersede the older Strong / Partial / Gap visualization for the current executive review surface. They are designed to route diligence, not to close it.

CodeCapabilityCoverageVisualMaturity / boundaryReview route
A · Pre-training Stack
A1Data66 canonical descendants90%
High architecture breadth; data, governance and provenance depth. Phase 1 source-system history remains contextual, not Phase 2 solo proof.Request evidence →
A2Tokenizer72 canonical descendants92%
Implemented and internally tested; independent frontier-scale compression/downstream-quality benchmarking remains open.Request evidence →
A3Architecture44 canonical descendants86%
Deep documented model/system architecture; no claim of a frontier foundation-model build.Request evidence →
A4Training35 canonical descendants84%
Reviewer-grade training recipe across eight decision areas; full-scale execution remains compute-dependent.Request evidence →
A5Compute21 canonical descendants67%
Substantial architecture plus GPU monitoring implementation; frontier-scale training-cluster execution is not demonstrated.Request evidence →
B · Post-training & Alignment
B1SFT19 canonical descendants79%
Structured fine-tuning strategy and training decisions; a large reproducible SFT campaign is not established on this public surface.Request evidence →
B2Preference Optimization29 canonical descendants76%
Documented architecture/method coverage; implementation and comparative outcome evidence remain selective.Request evidence →
B3Constitutional Methods18 canonical descendants78%
Safety/governance architecture with explicit policy boundaries; independent efficacy validation remains separate.Request evidence →
B4Red-Teaming23 canonical descendants86%
Broad threat/security corpus, structured tests and explicit gap-closure priorities; maturity varies by sub-family.Request evidence →
C · Evaluation & Safety
C1Capability Evaluation20 canonical descendants88%
Structured test architecture, comparison logic and review surfaces; external benchmark replication remains open.Request evidence →
C2Safety Evaluation20 canonical descendants91%
High documented coverage with hostile-review and failure-analysis logic; independent validation remains separate.Request evidence →
C3Robustness16 canonical descendants87%
Recovery, failover, anomaly, self-healing and stress/failure logic materially strengthen the legacy Partial view.Request evidence →
C4Output Safety11 canonical descendants93%
Strong output gating, egress, provenance and pre-commit control architecture; implementation depth varies by sub-family.Request evidence →
D · Inference & Production
D1Serving13 canonical descendants77%
Substantial architecture and runtime planning; broad production serving at frontier scale is not demonstrated.Request evidence →
D2Inference Optimization19 canonical descendants91%
DCA, UIOP, Multi-Brain, Suprompt, OFRP and memory/routing candidates; measured gains remain workload-dependent.Request evidence →
D3Monitoring15 canonical descendants94%
Strong architecture plus GPU Sentinel implementation/internal testing; independent performance validation remains open.Request evidence →
D4Deployment12 canonical descendants75%
Operational/deployment architecture exists; Phase 3 institutional deployment is intentionally future work.Request evidence →
E · Cross-Cutting
E1Data Governance17 canonical descendants91%
Object-first, lineage, reuse, retention, consent, export and dataset-gating architecture; hearing-grade internal material exists.Request evidence →
E2Security27 canonical descendants96%
Very broad cross-layer architecture with selected implementation/internal tests and a large restricted research corpus.Request evidence →
E3Privacy15 canonical descendants89%
Strong architectural treatment of classification, minimization, identity, lineage and reuse; legal validation is separate.Request evidence →
E4Compliance21 canonical descendants82%
Governance, audit, approval and evidence-routing maturity; jurisdiction-specific legal certification is not claimed.Request evidence →
Version migration: the earlier public MZN map used a 529-sub-endpoint snapshot and a 7 Strong / 13 Partial / 1 Gap position layer. The current canonical baseline is v1.0 (2026-05-09): 21 root slots and 554 total nodes, including 533 descendants. The old snapshot is historical context; it is not the current baseline or scoring method.
Representative evidence behind the recalibration

Depth is sampled, then the review expands where the question requires it.

Training

HUAI Training Recipe

Eight decision areas from model selection and fine-tuning through optimizer, schedule, batching, stability, parallelism and checkpoint management. It materially upgrades A4 architecture coverage while preserving the execution boundary.

Data / Privacy

Hearing-grade governance pack

Object-first handling, reuse separation, lineage, consent, retention, export logic, worked cases, reviewer simulations and failure injections strengthen E1/E3 without pretending to be legal certification.

Security

Security as a system

Mother / Genesis, ZOE, ISBP and related materials span root trust, identity/privilege, runtime controls, audit, recovery, lifecycle, output control and adaptive defense, with heterogeneous implementation maturity.

Monitoring

GPU Sentinel

Implementation/internal-test evidence for accelerated-infrastructure telemetry, anomaly, operations, audit and hardware-trust surfaces makes D3 materially stronger than architecture-only coverage.

Tokenizer

Implementation-level asset

Multiple tokenizer families, multilingual handling, fragmentation/context-budget work and internal test ladders support high A2 maturity while independent frontier-scale benchmarking remains open.

Gap discipline

Known gaps are kept visible

Runtime alignment monitoring, adversarial fuzzing, output provenance, data poisoning detection, sensitive-operation MFA and supply-chain security remain explicit closure targets where applicable.

Review rule

Ask the question that matters. Then open the evidence required to answer it.

A reviewer does not need to read hundreds of artifacts before asking a useful question. Choose a slot, capability family, maturity claim, chronology point or implementation claim. MZN can then surface the relevant public material, controlled evidence or restricted dossier. If the conclusion requires a broader sample, the review can expand to the full relevant corpus.

Challenge a slot

Example: “Show the evidence behind A4 Training or E2 Security.”

Challenge maturity

Ask what is architecture, implemented, internally tested or actually operating.

Challenge provenance

Route Phase 1 team work and Phase 2 solo formation to their separate records.

Challenge validation

Define the benchmark or reproduction test instead of treating documentation as proof.

Challenge the map

The goal is not predetermined agreement.

The goal is an evaluation process large enough for the object being evaluated. Request the relevant evidence, reproduce what requires validation, and reject or revise what does not survive review.