Independent models can widen the review surface. They do not become the jury.
The historical MZN evaluation protocol used multiple frontier models to stress-test definitions, identify blind spots and compare interpretations. The useful output is disagreement, question quality and reasoning trace—not synthetic validation.
Framework & scope
Give models the same phase boundary, claim definition and review question before comparing conclusions.
Evidence-aware questioning
Ask models to separate public evidence, restricted material, missing evidence and independent-validation requirements.
Compare disagreement
Model variance is useful. It can reveal prompt sensitivity, category ambiguity and assumptions that need human review.
Cross-model agreement is not independent validation. It is an AI-assisted review signal that can improve diligence design.