research-document CORPUS-100-FRONTIER

Reference corpus and classification evidence frontier

Reference corpus and classification evidence: research frontier

Scope and source artifacts

These files are treated as one analysis unit because they share an evidence base or form one accepted/current package. Generated derivatives, templates, prompts, private working packets, and superseded runs are not counted as independent conclusions.

Knowledge extraction

  • Primary objective: Characterize 100 knowledge artifacts across identity, capability, lifecycle, and composition.
  • Primary claims: The corpus covers diverse artifact identities and capabilities. Most current entries are established artifacts.
  • Methodology: Batch-based single-analyst characterization under FEMS-1.
  • Accepted hypotheses/findings: The corpus contains 100 indexed artifacts.
  • Rejected or unsupported hypotheses: Current frequency distributions are validated population estimates.
  • Assumptions: The sample is sufficiently representative. Labels are reliable despite single-analyst coding.
  • Limitations: All characterizations are draft; capability assignments are provisional; composition is sparse.
  • Known uncertainties: Inter-rater reliability; Sampling bias; Label stability; Primitive saturation.
  • Confidence: bounded by the source package; no confidence is promoted by this frontier analysis.

Adversarial challenge

The corpus is operationally treated as an evidence base while its statistics explicitly describe all characterizations as draft. Evidence that would materially change the assessment includes independent replication, preregistered negative controls, matched causal comparison, or a demonstrated failure of the current measurement instruments. The analysis deliberately treats planned validation as a dependency, not as completed evidence.

Top five opportunities

Rank Record Opportunity Category Frontier score
1 RFR-2026-006 Blind corpus reliability and sampling audit Statistics 391
2 RFR-2026-007 Primitive-vocabulary saturation study Measurement 314
3 RFR-2026-002 Independent human procedural-extraction baseline Human Factors 491
4 RFR-2026-005 Reasoning-versus-coordination grammar test Theory 392
5 RFR-2026-013 Accessible representation equivalence study Accessibility 138

Semantic duplicates were merged into the linked repository-wide records. The five selections maximize uncertainty reduction and repository reuse, rather than creating five unique labels for each source package.