research-document
Evidence Ledger
Evidence Ledger
Status: Working draft
Purpose: Track claims about Framework Engineering and the current state of evidence supporting those claims.
| Claim | Supporting Evidence | Opposing Evidence | Confidence | Status |
|---|---|---|---|---|
| Longitudinal Reference Cases reveal framework behavior that cannot be observed through single-state validation. | LRC-001 | None so far | High | Supported |
| Evidence Isolation materially reduces hindsight contamination. | LRC-001 | None so far | High | Supported |
| Diagnostic Stability is measurable. | LRC-001 | Needs additional cases | Medium | Provisionally Supported |
| Research Question Sets improve longitudinal comparability. | LRC-001 | Needs additional cases | Medium-High | Supported |
| Different components of a framework evolve at different rates as evidence changes. | LRC-001 | Needs additional cases | Medium | Emerging |
| Single aggregate scores can hide materially different weaknesses. | EDF rubric discussion | Not yet empirically tested | Medium | Supported as design concern |
| Evidence profiles better preserve actionable evaluation detail than single scores. | EDF rubric discussion; artifact maturity discussion | Needs validation during IFVS/DAR use | Medium | Provisionally Supported |
| Framework Engineering should avoid artificial precision in qualitative judgments. | Cross-framework evaluation discussion | None so far | Medium-High | Supported |
| Framework Engineering can downshift appropriately for trivial problems. | ARC-004 | Needs additional adversarial cases | Medium | Provisionally Supported |
| Framework Engineering can preserve uncertainty under sparse evidence. | ARC-001 | Needs additional adversarial cases | Medium | Provisionally Supported |
| Framework Engineering can stop when evidence is insufficient. | ARC-005 | Needs additional adversarial cases | Medium | Provisionally Supported |
| Calibration behavior appears to include complexity, confidence, and conclusion calibration. | ARC-001, ARC-004, ARC-005 | Needs broader validation | Medium | Emerging |
| Many frameworks appear to cluster by grammar rather than domain. | Framework Census Batch 001 | Needs larger census and blind clustering | Medium | Provisionally Supported |
| SWOT is better understood as Situational Assessment than Strategy. | Batch 001 grammar reduction | Needs comparison with more strategy frameworks | Medium | Provisional |
| Primary Question is a useful discriminator for framework archetypes. | Batch 001 | Needs more frameworks | Medium | Under investigation |
| Framework Engineering requires a boundary distinction between frameworks and adjacent artifact types. | FE-005B prediction pass; BPMN/OSI/TCP controls | Needs larger artifact review | Medium | Supported as design need |
| Many artifacts commonly called frameworks may be better classified as models, notations, standards, or tools. | FE-005B controls | Needs review | Medium | Under investigation |
| Artifact characterization should precede artifact classification. | FE-008 design rationale | Needs blind review evidence | Medium | Supported as design rule |
| Separating artifact identity from capabilities handles overlap better than layered classification labels. | FE-008 overlap analysis | Needs applied case review | Medium | Provisionally Supported |
| Separating Foundation, Research, and Applications improves epistemic clarity. | Repository architecture review; Knowledge Promotion discussion | Needs contributor usability testing | Medium | Provisionally Supported |
| Separating identity from capabilities improves knowledge artifact classification. | FE-008 Pilot Batch 001; FE-008 Pilot Batch 002 | Needs larger reviewer comparison | Medium | Provisionally Supported |
| Identity and Capability are independent dimensions. | FE-008 Pilot | None so far | Medium | Provisionally Supported |
| Framework Engineering currently has evidence as a characterization and improvement methodology, not yet as a full engineering discipline. | FE-008 identity-capability work; FE-011A redesign pilot; FE-012A reconstruction run; FE-012B synthesis run | No independent validation across the full theory stack yet | Low | Provisionally Supported |
| The Theory of Framework Engineering is mature enough to enter adversarial validation. | FE-008, FE-011A, FE-012A, FE-012B, theory traceability work | No independent replication yet | Medium | Ready for validation |
| Independent model families can converge on primitive grammar extraction using a shared packet instrument. | FE-012C manual replication run | Blinding weak; no human validation; provided vocabulary may constrain agreement | Moderate | Under Investigation |
| Artifact-first corpus structure improves traceability over CSV-only corpus structure. | Reference Corpus architecture discussion | Needs use during Batch 001 | Medium | Under Investigation |
| Framework Engineering redesigns may improve task output quality compared with original frameworks. | FE-010 redesign exercise; FE-011A pilot kit | Not yet tested with blinded outputs | Low | Under Investigation |
| A finite primitive reasoning vocabulary may underlie diverse knowledge artifacts. | FE-012A internal extraction run; FE-011A reasoning-layer pattern | No independent extractor convergence yet | Low | Provisionally Supported |
| A frozen primitive reasoning vocabulary may be expressive enough to synthesize coherent procedural frameworks for novel problems. | FE-012B internal synthesis run; FE-012A internal extraction run | Coordination-heavy cases created missing primitive pressure around synchronization, scheduling, and sequencing | Low | Provisionally Supported |
| Procedural frameworks may exhibit language-like structures recoverable as procedural ASTs | FE-012C revealed stable reasoning-shape agreement and primitive grammar convergence | FE-013 not yet run; programming-language analogy may bias analysis | Low | Under Investigation |
| A research operating system organized around hypotheses and bundled evidence runs may improve evidence collection efficiency without increasing theory inflation | Need ECR-000001 and follow-on runs | Operating overhead may exceed benefit; multi-hypothesis runs may blur interpretation | Low | Under Investigation |
| Hypothesis dependency tracking may reduce research drift and accelerate multi-hypothesis evidence runs | FE-012C revealed comparator and abstraction drift; research operating system introduced hypothesis tracking | Dependency model not yet tested in live evidence runs | Low | Under Investigation |
| Recognition sensitivity testing is required before stronger claims about procedural extraction can be made | ECR-000001 showed high recognition and dashboard marked H013 not ready | ECR-000002 not yet run | Moderate | Active Test Planned |
| Recognition may depend on control-flow topology rather than lexical cues alone | ECR-000002 P001D reduced recognition for GPT/Gemini but not Claude | Single graph-only packet; small sample; Claude-specific behavior unclear | Low | Active Test Planned |
Notes
- Claims should accumulate evidence over time rather than replace earlier observations.
- Confidence should increase only when additional independent evidence supports the same claim.