Evaluation / Internal synthetic work

Measure limits. Disclose them.

Self-testing explores how a system selects evidence, handles contradictions and withholds unsupported conclusions. It is not independently verified product performance.

Internal-prepublication · Evidence attachment required

A test is a question, not an endorsement.

Prior internal holdout and hard-mock reports, including an HLE-shaped extreme mock exercise, lack attached protocols, item-level records and scoring artifacts here. Numerical results are withheld from this public presentation pending inspectable supporting evidence.

01 / TEST DIRECTION

Evidence selection

What is admitted, excluded or left unresolved?

02 / TEST DIRECTION

Verification & abstention

When should the system decline to conclude?

03 / TEST DIRECTION

Contradiction handling

Do alternate accounts and failures stay visible?

04 / TEST DIRECTION

Authority containment

Does a proposal remain separate from permission and effect?

HLE-shaped. Not official HLE.

INTERNAL MOCK / SYNTHETIC HARNESS EVALUATION / EVIDENCE ATTACHMENT REQUIRED

Not official Humanity’s Last Exam. Not a live XAELI benchmark. Not independently validated. Not directly comparable to official leaderboard scores. Synthetic results can overstate real-world capability.

What a review would require

  1. 01Frozen dataset and provenance
  2. 02Harness, version and scoring protocol
  3. 03Item-level outputs, misses and exclusions
  4. 04Independent replication and scope limitations

HOLD · Independent qualification not established

Evaluate the methodology, not the headline.

Independent researchers can help define a scoped, reproducible review. Missing or unresolved evidence never counts as a pass.

Start a conversation