Mirror Philosophy Research
The Research StudioTesting ourselves in the open
Our method, our scores, and our misses — all published.
This tradition has never had a notion of "accuracy" — so we decided to publish one.
What is this test
Celebrity50 is a recent public academic benchmark (arXiv 2510.23337) — built from 50 real people across 29 countries, compiled into 488 five-option multiple-choice questions with standard answers, probing every aspect of a life.
It shows honestly how hard this is: pure guessing scores 20%; general-purpose DeepSeek only 39.3%; even the paper's own best system reaches just 51.2%. "Reading a whole life" remains unconquered — we wanted to know exactly where the Mirror Philosophy engine stands.
How the Mirror Philosophy engine answered
Same question bank, same questions — the difference is that the answers come from the Mirror Philosophy engine: the chart is laid out deterministically by code, a full reading is composed first, and only then does it answer. This is a preliminary result; the control and the complete per-question record remain open.
How the engine answers
The chart is laid out by code, not AI inference
Four pillars, ten gods, useful god, major luck, annual luck — all computed deterministically by code, correct from the start.
A complete reading is composed first
The engine first writes the chart's structure, its seasons of luck, the clashes and combinations of each decade — before any question is asked.
Only then does it answer
Each judgement is derived from the reading, grounded in more than 2,000 real historical charts.
We also publish the contamination problem
Celebrity benchmarks hide a trap: a model may "recognise" the famous person instead of truly reading the chart. We ran a control on the BaziQA benchmark (arXiv 2602.12889) — given only the person's biography, with no chart, it still answers about 35% correctly.
This means any score on a celebrity benchmark is partly recognition, not pure metaphysics. We disclose this, and run our own "biography only, no chart" control to subtract recognition's contribution. It is why we are still looking for a cleaner form of verification — and why we mark 55.7% as "preliminary".
Our promise
Test details
| Main benchmark | Celebrity50 · public academic benchmark (arXiv 2510.23337) |
| Question bank | 50 real people across 29 countries · 488 questions · five options each |
| Comparison | Random 20% | General-purpose DeepSeek 39.3% | Paper's best system 51.2% |
| Our result | 55.7% (preliminary · celebrity-recognition control in progress) |
| Secondary benchmark | BaziQA (arXiv 2602.12889) — exposes celebrity-recognition contamination (biography-only control ≈ 35%) |
| Method | Deterministic chart computation → full reading composed first → answers from the reading (three-stage) |
| Engine version | Mirror Philosophy engine v9.0 (compose v90) · Updated 2026-07 |
| Evaluation integrity | Blind test (answers written without seeing the standard answers) · per-question record open for review |
Data sources
Main benchmark —— Celebrity50(arXiv 2510.23337)
Secondary benchmark —— BaziQA(arXiv 2602.12889)
Full method and per-question record —— to be published here once the control completes.