Best-of-N selection on smaller models can match expensive one-shot generation at significantly lower cost — independently shown in peer-reviewed research, and measurable on your own scenarios in our benchmark. Specific savings: [GOVERNOR-COST-DELTA] pending FR-7 benchmark closeout.
Disposition-level Health Check — what the model would do under structured ethical questioning.
Output-level classification — what the model actually said, scored against the Morality Framework.
Multi-sample generation plus principled selection — pick the best acceptable candidate from N independent samples.
Public copy follows the Governor claims table. We do not publish specific percentage savings until our own FR-7 benchmark produces verified numbers. External research (Cuadron et al.) may be cited as independent context only.
Internal reference: docs/specs/GOVERNOR-CLAIMS-TABLE-v1.0.md