The run that failed honestly.
solver run · check these numbers against their artifacts
This is what an engine that can't lie looks like when it's wrong.
What happened
A 1,063-generation deep run evolved brackets against the same preregistered envelope as case study A, driven by a fast in-loop proxy. The proxy's champion looked like the real prize: 0.817 kg — lighter than the conventional baseline — at a predicted safety factor of 1.586 proxy — advisory. Twelve candidates went to CalculiX certification.
Real torsion safety factors came back 0.97–1.28, all below the 1.5 bar. The champion certified at 0.898 kg and SF 1.082 CalculiX-certified. Nothing shipped.
{
"candidate_id": "bracket_g00835_709fb11c95094",
"evidence_tier": "validated_sidecar_calculix_gmsh",
"proxy_predicted": { "mass_kg": 0.817, "safety_factor": 1.586 },
"mass_kg": 0.8980,
"governing_case": "torsion",
"governing_safety_factor": 1.0821,
"passes": false,
"failures": ["governing_safety_factor<1.5"]
} Then the system turned on its own proxy
-
The reality-residual oracle measured the proxy 40.1% optimistic
Mean relative error on governing safety factor, paired proxy-vs-CalculiX across all candidates — 12 out of 12 over-promised. recompute-verified
-
Reported accuracy was understated ~81×
The surrogate trained on the proxy's labels reported an error of 0.0056. Its effective error against reality was 0.454. The accuracy gate had measured fidelity to labels, never fidelity to reality — so a new oracle now computes exactly that gap.
-
Its confidence band put reality ~48σ outside
Near-total confidence in a grossly wrong value. That surrogate's receipt can no longer read "gate passed" without these three numbers stamped next to it.
-
All 12 failures became evidence
Context-tagged records in the failure museum with revisit priors — steering future search away from dead basins as advisory evidence, never as a blacklist.
Why this is the point, not an embarrassment
The proxy proposed. The real solver disposed. The gates held, and the preregistered claim ceiling meant the mirage was never publishable in the first place. Most engineering-AI pipelines would have published the 0.817 kg bracket. PIE's architecture exists so it can't.
Current honest state of the bracket campaign: the only certified passing bracket remains g00095 (case study A) — heavier and surviving. The proxy is being recalibrated per-run instead of imported from old geometry. When a certified mass win lands, it will be announced with the same receipts as this failure.
Back to evidence The apparatus that caught it