Mirror Philosophy Research

The Research Studio

Testing ourselves in the open

Our method, our scores, and our misses — all published.
This tradition has never had a notion of "accuracy" — so we decided to publish one.

Engine version v9.0 · Core compose v90 · Updated 2026-07
Preliminary · still under verification. The 55.7% below is our measured preliminary result on the Celebrity50 benchmark, higher than the paper's own best system (51.2%) and general-purpose DeepSeek (39.3%). We are still running the "celebrity recognition" control — until that control passes and at least two runs agree, this page is positioned as an "open, transparent self-evaluation", not a "world-first / beyond SOTA" claim.

What is this test

Celebrity50 is a recent public academic benchmark (arXiv 2510.23337) — built from 50 real people across 29 countries, compiled into 488 five-option multiple-choice questions with standard answers, probing every aspect of a life.

It shows honestly how hard this is: pure guessing scores 20%; general-purpose DeepSeek only 39.3%; even the paper's own best system reaches just 51.2%. "Reading a whole life" remains unconquered — we wanted to know exactly where the Mirror Philosophy engine stands.

How the Mirror Philosophy engine answered

Random guessing
20%
General-purpose DeepSeek
39.3%
The paper's best system
51.2%
Mirror Philosophy engine v9.0
55.7%

Same question bank, same questions — the difference is that the answers come from the Mirror Philosophy engine: the chart is laid out deterministically by code, a full reading is composed first, and only then does it answer. This is a preliminary result; the control and the complete per-question record remain open.

How the engine answers

I

The chart is laid out by code, not AI inference

Four pillars, ten gods, useful god, major luck, annual luck — all computed deterministically by code, correct from the start.

II

A complete reading is composed first

The engine first writes the chart's structure, its seasons of luck, the clashes and combinations of each decade — before any question is asked.

III

Only then does it answer

Each judgement is derived from the reading, grounded in more than 2,000 real historical charts.

We also publish the contamination problem

Celebrity benchmarks hide a trap: a model may "recognise" the famous person instead of truly reading the chart. We ran a control on the BaziQA benchmark (arXiv 2602.12889) — given only the person's biography, with no chart, it still answers about 35% correctly.

This means any score on a celebrity benchmark is partly recognition, not pure metaphysics. We disclose this, and run our own "biography only, no chart" control to subtract recognition's contribution. It is why we are still looking for a cleaner form of verification — and why we mark 55.7% as "preliminary".

Our promise

We publish our own misses. Some questions in these banks no chart can answer. On those, we honestly admit: we are no more accurate than anyone else. An engine that tells you how certain it is, and where it cannot see, can be checked against its own record.

Test details

Main benchmarkCelebrity50 · public academic benchmark (arXiv 2510.23337)
Question bank50 real people across 29 countries · 488 questions · five options each
ComparisonRandom 20% | General-purpose DeepSeek 39.3% | Paper's best system 51.2%
Our result55.7% (preliminary · celebrity-recognition control in progress)
Secondary benchmarkBaziQA (arXiv 2602.12889) — exposes celebrity-recognition contamination (biography-only control ≈ 35%)
MethodDeterministic chart computation → full reading composed first → answers from the reading (three-stage)
Engine versionMirror Philosophy engine v9.0 (compose v90) · Updated 2026-07
Evaluation integrityBlind test (answers written without seeing the standard answers) · per-question record open for review

Data sources

Main benchmark —— Celebrity50(arXiv 2510.23337)
Secondary benchmark —— BaziQA(arXiv 2602.12889)
Full method and per-question record —— to be published here once the control completes.

Mirror Philosophy · The Research Studio — The chart is the mirror; the Way completes in the heart.