research-frontier-record RFR-2026-001

Negative non-isomorphic control benchmark

RFR-2026-001: Negative non-isomorphic control benchmark

Research opportunity

Determine whether the ECR-000003 comparator reports structural stability when packets share surface features but not procedural structure.

Background

ECR-000003 produced weak-to-moderate signals, but the dashboard names false structural equivalence as the highest current validity risk.

Origin documents and trace

Specific assumption challenged: Observed structural similarity reflects shared procedure rather than permissive matching or superficial cues.

Supporting evidence: The current dashboard explicitly identifies false structural equivalence and selects negative controls; evidence strength for representation and isomorphism claims remains weak.

Unknowns

  • Comparator false-positive rate
  • Which representation features cause spurious equivalence
  • A defensible rejection threshold

Dependencies

  • None; can begin immediately.

Suggested REP and methodology

Create REP-2026-001. Preregister matched positive, negative-isomorphic, and negative-non-isomorphic packet families; blind labels; run frozen comparator 3.1.0; estimate sensitivity, specificity, calibration, and error by provider.

Expected outputs

  • Public benchmark packets
  • Frozen labels and scoring key
  • Comparator diagnostic report

Success criteria

A preregistered threshold separates positive from non-isomorphic controls with uncertainty bounds and an independently reproducible analysis.

Execution recommendation

  • Recommended agent: measurement-and-adversarial-validation agent
  • Estimated effort: medium
  • Expected knowledge gained: Calibrates the central instrument and may invalidate multiple downstream claims.
  • Frontier Score: 495
  • Score inputs: knowledge gain 5/5; impact 5/5; cross-project reuse 5/5; scientific importance 4/5; dependency cost 2/5; implementation difficulty 3/5.
  • Score calculation: 5 × 5 × 5 × 4 − 2 − 3 = 495.
  • Score status: structured expert judgment, not measured utility.