research-frontier-record RFR-2026-002

Independent human procedural-extraction baseline

RFR-2026-002: Independent human procedural-extraction baseline

Research opportunity

Measure whether trained humans recover structures similar to one another and to model outputs from the same blinded packets.

Background

The repository repeatedly limits internal same-model evidence and has no independent human baseline.

Origin documents and trace

Specific assumption challenged: Model convergence is meaningful without knowing the human reliability ceiling or floor.

Supporting evidence: H003 and H015 name human extraction as the next test; the validation program requires independent replication.

Unknowns

  • Human inter-rater reliability
  • Human-model agreement
  • Training and expertise effects

Dependencies

Suggested REP and methodology

Create REP-2026-002. Recruit independent analysts; stratify expertise; blind artifact identity; train only on the frozen codebook; double-code packets; adjudicate after computing agreement.

Expected outputs

  • De-identified coding dataset
  • Reliability and disagreement audit
  • Human-model comparison

Success criteria

Sample-size justification is met and agreement estimates with confidence intervals are reported before adjudication.

Execution recommendation

  • Recommended agent: human-factors replication agent
  • Estimated effort: high
  • Expected knowledge gained: Establishes whether current measurement is reproducible beyond model self-consistency.
  • Frontier Score: 491
  • Score inputs: knowledge gain 5/5; impact 5/5; cross-project reuse 5/5; scientific importance 4/5; dependency cost 4/5; implementation difficulty 5/5.
  • Score calculation: 5 × 5 × 5 × 4 − 4 − 5 = 491.
  • Score status: structured expert judgment, not measured utility.