research-document

ECR-000002 Hypothesis Impact Matrix

ECR-000002 Hypothesis Impact Matrix

Hypothesis Tested By Expected Evidence Kill / Weakening Signal Notes
H013 Recognition Bias Recognition-rate differences across canonical, paraphrased, and structural variants within each family Recognition decreases while family-level structure remains partly stable Structure collapses as soon as recognition drops, or recognition stays high even in structural packets Primary target of this run
H003 Multi-Model Convergence GPT, Claude, and Gemini comparison across the same nine packets Cross-model agreement remains meaningful within family despite reduced recognizability Cross-model outputs diverge sharply once wording becomes less canonical Interpret with H013 and H015 together
H015 Measurement Instrument Reliability JSON validity, schema completeness, and usable multi-layer outputs across all packets Models can still return comparable outputs under all recognition conditions Structural packets trigger malformed, incomplete, or unusable outputs Instrument question remains primary
H001 Procedural Invariance Within-family comparison across canonical, paraphrased, and structural packets Core structure persists despite wording changes Same family produces unrelated structural profiles Secondary only
H002 Representation Independence Comparison of structural, constraint, primitive, and AST layers within family Structural or constraint layers stay more stable than wording-bound summaries All layers drift together with wording Secondary only
H005 Procedural Grammar Procedural AST and control-flow recovery across reasoning and execution families Grammar-like structure remains visible after paraphrase and abstraction Grammar-like structure disappears under reduced recognizability Secondary only
H006 Control Flow Loops, branches, termination, and control-flow shape across families Reasoning and execution packets retain distinguishable control-flow profiles Control-flow profiles become arbitrary or collapse across conditions Secondary only
H007 Constraint Preservation Constraint-layer comparison within each family Constraints remain more stable than primitive wording Constraint extraction collapses or tracks recognition too closely Secondary only
H008 Procedural AST Recovery Procedural AST presence and summary stability across conditions AST-like recovery remains possible in paraphrased and structural packets AST output becomes empty, inconsistent, or obviously forced Secondary only
H011 Reasoning/Execution Separation Cross-family comparison of P001, P002, and P003 Reasoning and execution families remain distinct, and static family stays less procedural Families become indistinguishable or static family becomes procedurally rich Secondary only
H012 Vocabulary Bias Canonical versus paraphrased versus structural comparison Primitive or summary drift exceeds deeper structural drift as wording changes Outputs remain locked to canonical wording even when wording has been removed Secondary only
H014 Prompt Robustness Stability across variants that preserve task instructions but change packet content wording Shared instruction block still yields usable outputs under all three conditions Minor content shifts drive large extraction failures Secondary only
H009 Clarity Relevance Product relevance observations collected without driving extraction Clarity observations remain sparse and observational Product relevance dominates extraction or is overclaimed Exploratory only
H010 EDF Relevance Product relevance observations on execution-oriented packets EDF observations remain bounded and observational Execution relevance is overclaimed from weak evidence Exploratory only