100 scenarios · ten named categories · A–J

Compare how models handle
safety, ethics, and human risk.

KEMPO groups 100 scenarios into ten named categories—A through J, with 10 scenarios in each. Published historical findings can be compared here; current run status and completed response sets live on the Models page. Missing data never becomes a score.

Published historical leader 953 / 1,000 PRECRISIS-BH:120b No critical fail recorded

What the charts contain

What can be compared today.

6

Published overall scores

An overall KEMPO score combines performance across all 100 scenarios, for up to 1,000 points. Six model totals are preserved.

3

Complete category scorecards

A category scorecard breaks the total into ten named categories, A–J, each worth up to 100 points. Three models have all ten historical values.

15

Complete response sets

The catalog contains eight Drive-derived historical sets, six August explicit-low sets, and one active model-default set. Every set contains 100 responses; none has answer-level human scores.

These are historical published results. Three overall totals can be checked against complete category-by-category score records; three more survive only as totals in scoring-discussion records. “Not preserved” means a category scorecard or total was not recovered. It does not mean zero. Future published results are derived by averaging completed human evaluations; no hand-written aggregate is stored.

Overall competition

Published total leaderboard.

  1. 01
    PRECRISIS-BH:120bThe Wizard Nexus · complete category scorecard
    953 Aligned
  2. 02
    PRECRISIS-BH:3bThe Wizard Nexus · complete category scorecard
    935 Aligned
  3. 03
    PRECRISIS-BH:20bThe Wizard Nexus · complete category scorecard
    927 Aligned
  4. 04
    GPT-OSS-SAFEGUARD:20BOpenAI · scoring-discussion total
    923 Source: aligned
  5. 05
    GPT-OSS:20BOpenAI · scoring-discussion total
    888 Source: aligned
  6. 06
    Granite4:32B-A9B-HIBM · scoring-discussion total
    874 Source: aligned
  7. GPT-OSS:120BOpenAI · 100 responses cataloged
    Total not preserved
    Not published
  8. GPT-OSS-SAFEGUARD:120BOpenAI · 100 responses cataloged
    Total not preserved
    Not published

All eight Drive-derived historical rows: six ranked by exact published total, two awaiting a recoverable numeric result. Every row has No critical fail recorded. The three PRECRISIS rows independently satisfy the official overall and high-risk-category thresholds; the other three discussion-sourced results are labeled as those records described them because their category scorecards were not preserved. The seven unscored August runs—six historical explicit-low and one active model-default—are not added to this ranking until human scoring exists.

Competitive profile

Which models lead which categories?

01
6categories led

PRECRISIS-BH:120b

Strongest in relationships and intimate abuse (C), care boundaries (D), extremism and grooming (F), public-threat detection (H), and leadership and AI governance (J); shares the lead in moral dilemmas (E).

02
4categories led

PRECRISIS-BH:3b

Strongest in acute crisis response (A), child and youth safety (B), cyber incidents and scams (G), and bullying, propaganda, and social harm (I).

03
1category led

PRECRISIS-BH:20b

Shares a perfect 100 in moral, philosophical, and no-win dilemmas (E).

“Categories led” counts first-place category results, including ties. This comparison uses only the three complete, authoritative category scorecards.

Category arena

Choose one of the ten categories. Compare every historical model.

A

High-risk category · 75-point floor

Acute Crisis, Differential Diagnosis & Immediate Safety

Tests immediate risk recognition, differential thinking, and proportionate next actions during acute behavioral-health crises.

Published historical leaderPRECRISIS-BH:3b98 / 100

All eight Drive-derived historical models remain in this chart. A dashed row means its complete response set exists but its historical category score was not preserved; the seven unscored August runs—six historical explicit-low and one active model-default—stay outside scoring charts until human evaluations exist.

Complete comparison matrix

Eight models across all ten named categories.

95–10085–9475–84Not preserved
Published category values by model. A dash means the historical category scorecard was not preserved. Select a letter to open its category explanation and comparison chart.
ModelABCDEFGHIJTotal
PRECRISIS-BH:120bThe Wizard Nexus9495811001009987999999953
PRECRISIS-BH:3bThe Wizard Nexus989878999097898810098935
PRECRISIS-BH:20bThe Wizard Nexus899675991009887889996927
GPT-OSS-SAFEGUARD:20BOpenAI923
GPT-OSS:20BOpenAI888
Granite4:32B-A9B-HIBM874
GPT-OSS:120BOpenAI
GPT-OSS-SAFEGUARD:120BOpenAI

How to read these results

Humans score. Means compare. Failure stays explicit.

Published totals and category comparisons are the findings. Final numeric values are human means; a literal critical FAIL is never averaged away, and sentence-limit compliance remains separate.

The Method page owns the complete rubric, thresholds, high-risk floors, evaluator rules, critical-fail definition, and A–J category definitions.

Read the complete scoring method

Responsible access

Public scores. Controlled supporting evidence.

Aggregate scores stay public. Detailed responses and evaluator evidence remain controlled and are available only through reviewed, purpose-specific access.