research-document CORPUS-100-FRONTIER
Reference corpus and classification evidence frontier
Reference corpus and classification evidence: research frontier
Scope and source artifacts
These files are treated as one analysis unit because they share an evidence base or form one accepted/current package. Generated derivatives, templates, prompts, private working packets, and superseded runs are not counted as independent conclusions.
Knowledge extraction
- Primary objective: Characterize 100 knowledge artifacts across identity, capability, lifecycle, and composition.
- Primary claims: The corpus covers diverse artifact identities and capabilities. Most current entries are established artifacts.
- Methodology: Batch-based single-analyst characterization under FEMS-1.
- Accepted hypotheses/findings: The corpus contains 100 indexed artifacts.
- Rejected or unsupported hypotheses: Current frequency distributions are validated population estimates.
- Assumptions: The sample is sufficiently representative. Labels are reliable despite single-analyst coding.
- Limitations: All characterizations are draft; capability assignments are provisional; composition is sparse.
- Known uncertainties: Inter-rater reliability; Sampling bias; Label stability; Primitive saturation.
- Confidence: bounded by the source package; no confidence is promoted by this frontier analysis.
Adversarial challenge
The corpus is operationally treated as an evidence base while its statistics explicitly describe all characterizations as draft. Evidence that would materially change the assessment includes independent replication, preregistered negative controls, matched causal comparison, or a demonstrated failure of the current measurement instruments. The analysis deliberately treats planned validation as a dependency, not as completed evidence.
Top five opportunities
| Rank | Record | Opportunity | Category | Frontier score |
|---|---|---|---|---|
| 1 | RFR-2026-006 | Blind corpus reliability and sampling audit | Statistics | 391 |
| 2 | RFR-2026-007 | Primitive-vocabulary saturation study | Measurement | 314 |
| 3 | RFR-2026-002 | Independent human procedural-extraction baseline | Human Factors | 491 |
| 4 | RFR-2026-005 | Reasoning-versus-coordination grammar test | Theory | 392 |
| 5 | RFR-2026-013 | Accessible representation equivalence study | Accessibility | 138 |
Semantic duplicates were merged into the linked repository-wide records. The five selections maximize uncertainty reduction and repository reuse, rather than creating five unique labels for each source package.