research-frontier-record RFR-2026-006

Blind corpus reliability and sampling audit

RFR-2026-006: Blind corpus reliability and sampling audit

Research opportunity

Estimate classification reliability and sampling bias across the 100-artifact corpus.

Background

The corpus is established in size but all characterizations are draft and single-analyst.

Origin documents and trace

Specific assumption challenged: Frequency tables from provisional single-analyst labels represent the target population.

Supporting evidence: Corpus statistics explicitly state single-analyst draft status, provisional capability assignments, and sparse composition.

Unknowns

  • Inter-rater reliability
  • Coverage bias
  • Stability of identity/capability labels

Dependencies

  • None; can begin immediately.

Suggested REP and methodology

Create REP-2026-006. Define target population and sampling frame; draw a stratified blind re-review sample; estimate agreement and prevalence-adjusted uncertainty; audit missingness.

Expected outputs

  • Sampling frame
  • Double-coded sample
  • Reliability and bias report

Success criteria

Reliability and uncertainty are published by artifact family, with disagreements retained rather than overwritten.

Execution recommendation

  • Recommended agent: corpus-statistics agent
  • Estimated effort: high
  • Expected knowledge gained: Turns descriptive counts into interpretable evidence.
  • Frontier Score: 391
  • Score inputs: knowledge gain 4/5; impact 4/5; cross-project reuse 5/5; scientific importance 5/5; dependency cost 4/5; implementation difficulty 5/5.
  • Score calculation: 4 × 4 × 5 × 5 − 4 − 5 = 391.
  • Score status: structured expert judgment, not measured utility.