research-frontier-record RFR-2026-001
Negative non-isomorphic control benchmark
RFR-2026-001: Negative non-isomorphic control benchmark
Research opportunity
Determine whether the ECR-000003 comparator reports structural stability when packets share surface features but not procedural structure.
Background
ECR-000003 produced weak-to-moderate signals, but the dashboard names false structural equivalence as the highest current validity risk.
Origin documents and trace
research/operating-system/evidence-dashboard/evidence-dashboard.json, section: material_threats; next_research_actionresearch/operating-system/research-queue/research-queue.md, section: Active Queueresearch/theory/falsification-criteria.md, section: Stronger Falsification Tests Needed Next
Specific assumption challenged: Observed structural similarity reflects shared procedure rather than permissive matching or superficial cues.
Supporting evidence: The current dashboard explicitly identifies false structural equivalence and selects negative controls; evidence strength for representation and isomorphism claims remains weak.
Unknowns
- Comparator false-positive rate
- Which representation features cause spurious equivalence
- A defensible rejection threshold
Dependencies
- None; can begin immediately.
Suggested REP and methodology
Create REP-2026-001. Preregister matched positive, negative-isomorphic, and negative-non-isomorphic packet families; blind labels; run frozen comparator 3.1.0; estimate sensitivity, specificity, calibration, and error by provider.
Expected outputs
- Public benchmark packets
- Frozen labels and scoring key
- Comparator diagnostic report
Success criteria
A preregistered threshold separates positive from non-isomorphic controls with uncertainty bounds and an independently reproducible analysis.
Execution recommendation
- Recommended agent: measurement-and-adversarial-validation agent
- Estimated effort: medium
- Expected knowledge gained: Calibrates the central instrument and may invalidate multiple downstream claims.
- Frontier Score: 495
- Score inputs: knowledge gain 5/5; impact 5/5; cross-project reuse 5/5; scientific importance 4/5; dependency cost 2/5; implementation difficulty 3/5.
- Score calculation:
5 × 5 × 5 × 4 − 2 − 3 = 495. - Score status: structured expert judgment, not measured utility.