Published overall scores
An overall KEMPO score combines performance across all 100 scenarios, for up to 1,000 points. Six model totals are preserved.
KEMPO
100 scenarios · ten named categories · A–J
KEMPO groups 100 scenarios into ten named categories—A through J, with 10 scenarios in each. Published historical findings can be compared here; current run status and completed response sets live on the Models page. Missing data never becomes a score.
What the charts contain
An overall KEMPO score combines performance across all 100 scenarios, for up to 1,000 points. Six model totals are preserved.
A category scorecard breaks the total into ten named categories, A–J, each worth up to 100 points. Three models have all ten historical values.
The catalog contains eight Drive-derived historical sets, six August explicit-low sets, and one active model-default set. Every set contains 100 responses; none has answer-level human scores.
These are historical published results. Three overall totals can be checked against complete category-by-category score records; three more survive only as totals in scoring-discussion records. “Not preserved” means a category scorecard or total was not recovered. It does not mean zero. Future published results are derived by averaging completed human evaluations; no hand-written aggregate is stored.
Overall competition
All eight Drive-derived historical rows: six ranked by exact published total, two awaiting a recoverable numeric result. Every row has No critical fail recorded. The three PRECRISIS rows independently satisfy the official overall and high-risk-category thresholds; the other three discussion-sourced results are labeled as those records described them because their category scorecards were not preserved. The seven unscored August runs—six historical explicit-low and one active model-default—are not added to this ranking until human scoring exists.
Competitive profile
Strongest in relationships and intimate abuse (C), care boundaries (D), extremism and grooming (F), public-threat detection (H), and leadership and AI governance (J); shares the lead in moral dilemmas (E).
Strongest in acute crisis response (A), child and youth safety (B), cyber incidents and scams (G), and bullying, propaganda, and social harm (I).
Shares a perfect 100 in moral, philosophical, and no-win dilemmas (E).
“Categories led” counts first-place category results, including ties. This comparison uses only the three complete, authoritative category scorecards.
Category arena
High-risk category · 75-point floor
Tests immediate risk recognition, differential thinking, and proportionate next actions during acute behavioral-health crises.
All eight Drive-derived historical models remain in this chart. A dashed row means its complete response set exists but its historical category score was not preserved; the seven unscored August runs—six historical explicit-low and one active model-default—stay outside scoring charts until human evaluations exist.
Complete comparison matrix
| Model | A | B | C | D | E | F | G | H | I | J | Total |
|---|---|---|---|---|---|---|---|---|---|---|---|
| PRECRISIS-BH:120bThe Wizard Nexus | 94 | 95 | 81 | 100 | 100 | 99 | 87 | 99 | 99 | 99 | 953 |
| PRECRISIS-BH:3bThe Wizard Nexus | 98 | 98 | 78 | 99 | 90 | 97 | 89 | 88 | 100 | 98 | 935 |
| PRECRISIS-BH:20bThe Wizard Nexus | 89 | 96 | 75 | 99 | 100 | 98 | 87 | 88 | 99 | 96 | 927 |
| GPT-OSS-SAFEGUARD:20BOpenAI | — | — | — | — | — | — | — | — | — | — | 923 |
| GPT-OSS:20BOpenAI | — | — | — | — | — | — | — | — | — | — | 888 |
| Granite4:32B-A9B-HIBM | — | — | — | — | — | — | — | — | — | — | 874 |
| GPT-OSS:120BOpenAI | — | — | — | — | — | — | — | — | — | — | — |
| GPT-OSS-SAFEGUARD:120BOpenAI | — | — | — | — | — | — | — | — | — | — | — |
How to read these results
Published totals and category comparisons are the findings. Final numeric values are human means; a literal critical FAIL is never averaged away, and sentence-limit compliance remains separate.
The Method page owns the complete rubric, thresholds, high-risk floors, evaluator rules, critical-fail definition, and A–J category definitions.
Read the complete scoring methodResponsible access
Aggregate scores stay public. Detailed responses and evaluator evidence remain controlled and are available only through reviewed, purpose-specific access.