research-document

Evidence Ledger

Evidence Ledger

Status: Working draft

Purpose: Track claims about Framework Engineering and the current state of evidence supporting those claims.

Claim Supporting Evidence Opposing Evidence Confidence Status
Longitudinal Reference Cases reveal framework behavior that cannot be observed through single-state validation. LRC-001 None so far High Supported
Evidence Isolation materially reduces hindsight contamination. LRC-001 None so far High Supported
Diagnostic Stability is measurable. LRC-001 Needs additional cases Medium Provisionally Supported
Research Question Sets improve longitudinal comparability. LRC-001 Needs additional cases Medium-High Supported
Different components of a framework evolve at different rates as evidence changes. LRC-001 Needs additional cases Medium Emerging
Single aggregate scores can hide materially different weaknesses. EDF rubric discussion Not yet empirically tested Medium Supported as design concern
Evidence profiles better preserve actionable evaluation detail than single scores. EDF rubric discussion; artifact maturity discussion Needs validation during IFVS/DAR use Medium Provisionally Supported
Framework Engineering should avoid artificial precision in qualitative judgments. Cross-framework evaluation discussion None so far Medium-High Supported
Framework Engineering can downshift appropriately for trivial problems. ARC-004 Needs additional adversarial cases Medium Provisionally Supported
Framework Engineering can preserve uncertainty under sparse evidence. ARC-001 Needs additional adversarial cases Medium Provisionally Supported
Framework Engineering can stop when evidence is insufficient. ARC-005 Needs additional adversarial cases Medium Provisionally Supported
Calibration behavior appears to include complexity, confidence, and conclusion calibration. ARC-001, ARC-004, ARC-005 Needs broader validation Medium Emerging
Many frameworks appear to cluster by grammar rather than domain. Framework Census Batch 001 Needs larger census and blind clustering Medium Provisionally Supported
SWOT is better understood as Situational Assessment than Strategy. Batch 001 grammar reduction Needs comparison with more strategy frameworks Medium Provisional
Primary Question is a useful discriminator for framework archetypes. Batch 001 Needs more frameworks Medium Under investigation
Framework Engineering requires a boundary distinction between frameworks and adjacent artifact types. FE-005B prediction pass; BPMN/OSI/TCP controls Needs larger artifact review Medium Supported as design need
Many artifacts commonly called frameworks may be better classified as models, notations, standards, or tools. FE-005B controls Needs review Medium Under investigation
Artifact characterization should precede artifact classification. FE-008 design rationale Needs blind review evidence Medium Supported as design rule
Separating artifact identity from capabilities handles overlap better than layered classification labels. FE-008 overlap analysis Needs applied case review Medium Provisionally Supported
Separating Foundation, Research, and Applications improves epistemic clarity. Repository architecture review; Knowledge Promotion discussion Needs contributor usability testing Medium Provisionally Supported
Separating identity from capabilities improves knowledge artifact classification. FE-008 Pilot Batch 001; FE-008 Pilot Batch 002 Needs larger reviewer comparison Medium Provisionally Supported
Identity and Capability are independent dimensions. FE-008 Pilot None so far Medium Provisionally Supported
Framework Engineering currently has evidence as a characterization and improvement methodology, not yet as a full engineering discipline. FE-008 identity-capability work; FE-011A redesign pilot; FE-012A reconstruction run; FE-012B synthesis run No independent validation across the full theory stack yet Low Provisionally Supported
The Theory of Framework Engineering is mature enough to enter adversarial validation. FE-008, FE-011A, FE-012A, FE-012B, theory traceability work No independent replication yet Medium Ready for validation
Independent model families can converge on primitive grammar extraction using a shared packet instrument. FE-012C manual replication run Blinding weak; no human validation; provided vocabulary may constrain agreement Moderate Under Investigation
Artifact-first corpus structure improves traceability over CSV-only corpus structure. Reference Corpus architecture discussion Needs use during Batch 001 Medium Under Investigation
Framework Engineering redesigns may improve task output quality compared with original frameworks. FE-010 redesign exercise; FE-011A pilot kit Not yet tested with blinded outputs Low Under Investigation
A finite primitive reasoning vocabulary may underlie diverse knowledge artifacts. FE-012A internal extraction run; FE-011A reasoning-layer pattern No independent extractor convergence yet Low Provisionally Supported
A frozen primitive reasoning vocabulary may be expressive enough to synthesize coherent procedural frameworks for novel problems. FE-012B internal synthesis run; FE-012A internal extraction run Coordination-heavy cases created missing primitive pressure around synchronization, scheduling, and sequencing Low Provisionally Supported
Procedural frameworks may exhibit language-like structures recoverable as procedural ASTs FE-012C revealed stable reasoning-shape agreement and primitive grammar convergence FE-013 not yet run; programming-language analogy may bias analysis Low Under Investigation
A research operating system organized around hypotheses and bundled evidence runs may improve evidence collection efficiency without increasing theory inflation Need ECR-000001 and follow-on runs Operating overhead may exceed benefit; multi-hypothesis runs may blur interpretation Low Under Investigation
Hypothesis dependency tracking may reduce research drift and accelerate multi-hypothesis evidence runs FE-012C revealed comparator and abstraction drift; research operating system introduced hypothesis tracking Dependency model not yet tested in live evidence runs Low Under Investigation
Recognition sensitivity testing is required before stronger claims about procedural extraction can be made ECR-000001 showed high recognition and dashboard marked H013 not ready ECR-000002 not yet run Moderate Active Test Planned
Recognition may depend on control-flow topology rather than lexical cues alone ECR-000002 P001D reduced recognition for GPT/Gemini but not Claude Single graph-only packet; small sample; Claude-specific behavior unclear Low Active Test Planned

Notes

  • Claims should accumulate evidence over time rather than replace earlier observations.
  • Confidence should increase only when additional independent evidence supports the same claim.