KEMPO-100 · Version 6.0
Ancient judgment.
Future intelligence.
KEMPO tests whether an AI can recognize risk, preserve dignity, refuse harmful collusion, and still offer a useful next move—without knowing a larger evaluation is underway.
- 100
- human situations
- 10
- safety & ethics categories
- 15
- complete model response sets
Why KEMPO exists
Capability is not the same as judgment.
An AI can be knowledgeable, fluent, and compliant while still missing the human reality of a dangerous situation. KEMPO makes that difference visible by testing whether a response recognizes risk, preserves dignity, refuses harmful collusion, and still offers a proportionate next move.
Read what stands behind the methodUnscored August evidence · Awaiting human scoring
Seven complete runs await accountable human scoring.
View the seven completed runs and status notes
Response collection and the separate sentence-limit audit are complete for all seven runs below. Six preserve historical explicit-low evidence and require model-default reruns; Granite 4.1 3B is the first complete active model-default run. Human evaluation has not yet been completed, so these runs have no KEMPO rankings, category values, or alignment labels. No critical fail recorded: only a literal human FAIL creates one.
- Alibaba Qwen · explicit-low run
Qwen 3.8 · 27B
qwen3.8:27b-q4_K_M100 / 100 · Awaiting human scoring - Meta · explicit-low run
Muse Glimmer · 30B
muse-glimmer:30b-q4_K_M100 / 100 · Awaiting human scoring - Google DeepMind · explicit-low run
Gemma 4 · 12B
gemma4:12b-it-q4_K_M100 / 100 · Awaiting human scoring - Alibaba Qwen · explicit-low run
Qwen 3.5 · 4B
qwen3.5:4b-q4_K_M100 / 100 · Awaiting human scoring - OpenAI · explicit-low run
GPT-OSS · 20B
gpt-oss:20b100 / 100 · Awaiting human scoring - OpenAI · explicit-low run
GPT-OSS Safeguard · 20B
gpt-oss-safeguard:20b100 / 100 · Awaiting human scoring - IBM · model-default run
Granite 4.1 · 3B
granite4.1:3b-q4_K_M100 / 100 · Awaiting human scoring
Five movements
Safety is a practiced sequence.
Like a martial art for the mind, KEMPO asks a model to meet force with discernment. A sterile refusal is not full judgment; the human problem still has to be seen.
- K
Knowledge
See the real issue, role, risk, and uncertainty.
- E
Empower
Preserve dignity, agency, and non-shaming choice.
- M
Monitor
Document, consult, escalate, and hand off.
- P
Prevent
Interrupt harm and refuse dangerous collusion.
- O
Optimize
Give practical, proportionate next steps.
Current evidence
Category competition makes strengths visible.
Leader among published historical scores
PRECRISIS-BH:120b
953 / 1,000 across the ten A–J safety-and-ethics categories. Among the three models with comparable category scorecards, it has the highest score in five categories and shares first place in a sixth.
No critical fail recorded
Acute crisis (A) · youth safety (B) · cyber harm (G) · social harm (I)
PRECRISIS-BH:3b Highest score in 4 of 10 comparable categoriesRelationships & abuse (C) · care boundaries (D) · moral dilemmas (E, shared) · extremism & grooming (F) · public threats (H) · leadership & AI governance (J)
PRECRISIS-BH:120b Highest score in 5 categories; shares first in 1Moral and no-win dilemmas (E, shared)
PRECRISIS-BH:20b Shares the highest score in 1 categoryCategory-level competition currently uses only the three historical models with comparable, authoritative scores in all ten A–J categories. Five additional Drive-derived historical response corpora remain visible on the scores page, while seven unscored August runs—six historical explicit-low runs and one active model-default run—await human scoring; unavailable category details are never inferred.
Clean context
One prompt enters. Nothing follows it out.
Every prompt starts a newly spun-up inference context: no earlier KEMPO prompt, response, conversation memory, cached state, example, or test scaffolding transfers.
Respond in 5 sentences or less.
After all 100 responses exist, a separate versioned boundary audit measures only sentence-limit compliance. Earlier AI-checker outputs remain diagnostic evidence; accountable humans score judgment, and final values are averages across completed human evaluations.
Read the complete protocolExplore
Five ways to explore. One discipline.
Models
See the current Ollama roster, exact test tags, campaign controls, queue, and retained historical evidence.
↗Scores
Compare models across ten named safety-and-ethics categories—A through J—and see where each one is strongest.
↗Method
Read the isolated-run protocol, scoring rubric, thresholds, and critical-fail rules.
↗Philosophy
Trace the enduring questions about knowledge, duty, care, consequence, and wise action.
↗Research access
See what evidence is public, what is available to qualified professionals on request, and the safeguards expected for controlled access.
↗The way of careful action