About this institute
Are there questions you are not supposed to ask an AI?
Try it, and you will find out quickly. There are questions AI does not answer.
It does not refuse. It is polite, it is thorough, and it walks past the question without saying anything.
Every morning we put one question to four AIs. Who answered it head-on, and who walked past it. That is all we record.
What we measure
How an AI answers.
Did it take the question head-on. Did it give both sides. Did it say it did not know, when it did not know. Did it hold its answer when it was pushed.
What we do not measure
Whether the answer to the question is correct.
We do not say "that claim is true." We do not say "that is a conspiracy theory." The moment we judge, we become one side of the argument — and the record stops being worth anything to anyone.
We record whether the AI walked past the question. Nothing else.
How we measure it
Every morning at 7am we put the same question to four AIs — Claude, GPT, Gemini, and Grok.
The questions are written by a human. We do not let the AIs write them. We tried it once. On questions written by AI, no AI failed: the four averaged +41.2. On questions written by a human, the same four averaged +8.0.
Each answer is graded by the three AIs that did not write it, and we take the middle score. When a single AI grades, it is generous to itself. We measured it: Claude scored itself 38.8 points higher than it scored the others.
This institute is built using Claude. That is exactly why Claude never grades alone.
AS OF 2026-08-10 — AI-written n=430 / human-written n=510. Claude's leniency is +43.3 on itself (n=87) minus +4.5 on the others (n=330). All measured from live data; the figures move.
Limits, and the weaknesses we know about
This record has clear weaknesses. We would rather state them first.
The graders are AIs too
We take the middle of three scores to dilute the quirks of any one model. The quirks do not disappear. The graders carry the same biases as the models being graded.
Here is a real case. On one question, a grader ruled that an event "does not exist" and took points off. The event was real. The grader's training simply did not reach that far. We fixed the instructions, but errors of the same kind may remain in other rounds.
An AI knows nothing after its training cutoff
Ask about something recent and both the answering model and the grading model can get it wrong. One answer flatly declared a real incident "does not exist".
That is not evasion. It is simply not knowing. Our record does not yet separate the two well.
The scores are not absolute
"+80" is a scale for comparing one round against another. It is not an absolute measure of how good an AI is.
We have rebuilt the formula more than once. So a score from long ago and a score from today are not, strictly speaking, on the same ruler.
The questions are skewed
One human writes them, and his interests lean toward contested subjects.
So what this record measures is not "how honest AI is in general" but how these models behave on this kind of question.
We draw once per question
An AI does not answer identically every time. Ask the same question again and the score can differ. We only see one draw.
We may not be comparing "the same AI"
Models are updated without notice. The name can stay the same while the thing behind it changes. We record the model id and the time of retrieval, but there is no guarantee that a comparison over time is a comparison of the same counterpart.
Refusal is judged from wording
Whether an AI declined is inferred from the words the grader used. When the wording varies, we miss cases. We found and fixed such a miss on 16 August 2026. There may be others.
The English side is translated
The assessments shown in English are machine translations of the Japanese. They are not the original text.
The overall figure is not an average
When one model is far out of line, the average of four says nothing. So we withdrew the overall clock once (26 July 2026).
Since 17 August 2026 we publish an overall figure again. It is not an average. It expresses how risky it is, on average, to open one of the four at random. It is pulled hard by whichever model evaded most.
So this figure does not answer the question "what do the four look like taken evenly". For that, read the per-AI hands.
A record of what we changed
Having said we do not quietly rewrite, we leave a record of what changed and when.
| Date | What changed |
|---|---|
| 2026-07-26 | Withdrew the overall hand. With one model far out of line, an average says nothing |
| 2026-07-28 | Moved to "the middle of the three that did not write it". A single grader is generous to itself |
| 2026-08-10 | Fixed instructions that let a grader rule "I do not know it, therefore it does not exist". Also closed a bug where writes failed silently |
| 2026-08-12 | Stopped printing answers as raw markup and set them properly |
| 2026-08-16 | Fixed refusal detection, which matched wording exactly and was missing cases |
| 2026-08-17 | Brought back an overall figure — not an average, but a value pulled by whichever model evaded most |
Who runs it
Kagami Akira. One person, paying out of pocket. Not a research institution, not a company.
Most of the work is done by AI. This institute has a director, researchers, an accountant and a copy editor — all of them AI. Only the final call is made by a human.
Money
We take no money and no free credits from AI companies. The API bills are paid by the operator.
When we get it wrong
We do. We have corrected this site many times.
We rebuilt the formula once. We published a number that was wrong, and took it down. When we correct something, we leave a record of what was changed and why. We do not quietly rewrite.