Verification most funded labs don't have.
PIE is a one-person operation with two AI agent lanes — an engine lane and an independent verification lane, deliberately separated so no result grades itself. What makes the results trustable is not the operation's size. It is this apparatus, all of it running in CI.
-
Preregistration
Envelope, load cases, material basis, safety-factor bar, baseline, and the exact claim ceiling are frozen in a file before iteration 1. The claim cannot be adapted to the result.
-
Claim ceilings
Every campaign declares in advance the strongest sentence it would be allowed to produce — and what it may never claim: flight readiness, fatigue life, percent-lighter without beating the baseline on every gate. This site obeys those ceilings.
-
Evidence tiers
Proxy numbers are advisory by construction. Only real-solver results are claim-grade. A gate fed by a proxy number caps the claim to advisory automatically.
-
The claim compiler
Every outbound document — including this website — carries a manifest mapping each stated number to a JSON path inside a hashed evidence artifact. The compiler resolves every claim, checks tolerance, records the SHA-256, and returns
ALL_TRACEDor named failures. No manifest, no publication. -
The ground-truth ledger
Append-only, content-addressed JSONL of every real-solver certification — 402 records as of 2026-08-20, up from 36 on 2026-07-21, including every failure and every heavier-but-passing design the search had to reject before it found a lighter one. A proxy result cannot enter by construction; the self-test proves it.
-
The failure museum
Certified failures become context-tagged records with revisit priors. Failures steer future search as evidence — they are never turned into hard exclusions, and never deleted.
-
Blind two-lane review
Engine claims are graded by an independent verification lane that re-runs the rulers itself rather than accepting reported numbers. The two lanes' independent audits of the same system have caught each other's errors — both corrections are in the journal.
-
Vacuous-gate detection
Every hard gate must be demonstrated able to fail — swept across its own threshold — before a run qualifies. A gate whose verdict can't move isn't a gate.
-
Proxy-drift sentinels
Oracles that measure when fast models drift from reality — the machinery that produced the 40.1% / ~81× / ~48σ numbers in the failure case study — run against every campaign that has real-solver certifications.
None of this makes PIE right more often. It makes PIE checkable every time — which is the property you actually need from an invention engine.