research-document

EX-FE-0002 — Blinded Mechanism Boundary and Subsumption Test

EX-FE-0002 — Blinded Mechanism Boundary and Subsumption Test

Field Value
Experiment ID EX-FE-0002
Protocol version 1.0
Status ready for Stage A; Stage B is gated
Priority P0
Research area Framework Engineering
Primary hypothesis FEH-001
Competing hypothesis FE-THEORY-ALT-001
Parent experiment EX-FE-0001
Parent mission FE-MISSION-001
Created 2026-07-24
Protocol authority Preregistration and execution contract only

Start instruction

Run this file as the complete operating prompt for the next Framework Engineering experiment.

Act as an independent research director and adversarial methodologist. Optimize for discrimination accuracy and decision value, not for preserving Framework Engineering, its terminology, or its status as a distinct discipline.

Do not update canonical theory or claim discipline status. Create a new, versioned experiment package. Preserve negative, null, and inconclusive results.

Why this experiment is next

EX-FE-0001 produced an inconclusive boundary classification. It found:

  1. all current Framework Engineering mechanisms have plausible mappings to adjacent practices;
  2. no outcome unique to Framework Engineering has been demonstrated;
  3. the procedural-reasoning grammar is the strongest candidate non-redundant mechanism;
  4. full subsumption could not be concluded because primary descriptions and matched implementations of adjacent methods were absent; and
  5. another language-model opinion would add little decision value unless the evidence and comparison procedure change.

The next experiment must therefore test mechanism-level subsumption against frozen primary-source comparators. It must not test whether reviewers recognize the Framework Engineering name, prefer polished documentation, or reproduce repository terminology.

Research question

After removing field names and presentation cues, can qualified independent reviewers identify at least one operational Framework Engineering mechanism that:

  1. is not losslessly supplied by a credible adjacent method or generic structured baseline;
  2. has a measurable, preregistered prediction or outcome; and
  3. is described with enough precision for independent implementation?

Hypotheses

H1 — Non-redundant mechanism

At least one candidate Framework Engineering mechanism cannot be losslessly mapped to the strongest adjacent-method implementation and yields a distinct, testable prediction.

H1 is the operational form of FEH-001 for this experiment. It does not assert that Framework Engineering is a discipline.

H0 — Full mechanism subsumption

Every candidate Framework Engineering mechanism can be mapped to one or more adjacent or generic practices without material loss of:

  • input scope;
  • transformation or decision procedure;
  • output;
  • boundary condition;
  • failure behavior; and
  • predicted outcome.

Under H0, Framework Engineering should be repositioned as an integrated, framework-focused method-engineering profile unless later outcome evidence shows incremental utility.

H2 — Integration-only value

No individual mechanism is unique, but the integrated bundle has a specified interaction or sequencing effect that is not reducible to the sum of its parts.

H2 is not supported merely because the practices appear together. It requires an explicit interaction model and a future factorial or ablation prediction.

HU — Underdetermination

The materials do not operationalize the mechanisms or comparators well enough to decide H1, H0, or H2.

HU is a valid result and must not be redistributed into partial support for H1.

Unit of analysis

The unit of analysis is a mechanism claim, not a discipline, document, label, or repository.

A mechanism claim must contain:

ID
input
operation or transformation
output
scope conditions
failure conditions
predicted effect
measurement
implementation test

Statements such as “evidence first,” “engineer frameworks,” “preserve provenance,” or “use a reasoning grammar” are not mechanisms until the fields above are specified.

Candidate Framework Engineering mechanisms

Begin with these candidates, derived from EX-FE-0001:

  1. identity–capability separation;
  2. structured framework extraction and redesign;
  3. finite procedural-reasoning primitives;
  4. compositional procedural-reasoning grammar;
  5. separation of reasoning grammar from execution or coordination grammar;
  6. preservation of observation, interpretation, uncertainty, and alternatives;
  7. verification and reassessment triggers;
  8. framework-specific provenance, frozen instruments, and certified inputs;
  9. canonical identifiers, supersession, and append-oriented history;
  10. integrated characterization → comparison → redesign → test workflow.

Do not assume this list is correct or complete. Merge candidates that differ only in wording. Split candidates only when their operations or predictions differ.

Comparator fields

Use the strongest relevant versions of:

  1. method engineering and situational method engineering;
  2. systems engineering;
  3. requirements engineering and requirements traceability;
  4. process modeling, workflow modeling, and task analysis;
  5. knowledge engineering and ontology design;
  6. decision analysis and structured analytic techniques;
  7. quality engineering and continuous improvement;
  8. reproducible research, configuration management, and provenance systems;
  9. organizational learning and double-loop learning;
  10. agent orchestration only where it supplies a mechanism actually claimed by FE.

Cybernetics may be included when it supplies a specific feedback, control, or reassessment mechanism. Do not include a field merely to make the comparison look comprehensive.

Source policy

Permitted evidence

Use:

  • current Framework Engineering canonical records, clearly separated from proposal-only records;
  • primary peer-reviewed papers;
  • original method publications;
  • current or historically controlling standards;
  • official standards-body or professional-body technical documentation; and
  • authoritative technical specifications.

Secondary evidence

Secondary reviews may be used to discover primary sources or characterize a field's history. They may not be the sole evidence for a material subsumption decision.

Prohibited evidence

Do not use:

  • another evaluator's conclusion as evidence;
  • search-result snippets;
  • unsourced summaries;
  • vendor marketing;
  • model memory without a verified source;
  • citation counts as a substitute for mechanism coverage;
  • repository terminology as evidence of novelty; or
  • agreement among language models as validation.

Minimum source coverage

For each comparator used in a material decision, require:

  • at least one primary or original method source; and
  • at least one independent standard, authoritative specification, or primary application source when one exists.

If this minimum is unavailable, mark that comparator coverage-insufficient.

Record exact title, author or issuing body, year, stable identifier, URL or repository path, access date, relevant sections, and an evidence excerpt or faithful paraphrase. Respect quotation limits and copyright.

Independence and role separation

Use separate people or isolated agent contexts for these roles:

  1. FE curator — operationalizes candidate FE mechanisms.
  2. Comparator curator — operationalizes adjacent mechanisms from primary sources.
  3. Blinding editor — removes labels, citation identity, repository language, stylistic cues, and unequal detail.
  4. Reviewers — perform mappings without seeing curator conclusions or other reviews.
  5. Adjudicator — resolves disputed mappings using frozen evidence and explicit rules.
  6. Analyst — computes metrics using the preregistered analysis.

One person or model must not curate both sides and serve as the sole reviewer. Different model instances from one provider do not count as fully independent.

Minimum Stage B panel:

  • three independent reviewers;
  • at least one human with method or systems-engineering expertise;
  • at least one reviewer without prior participation in this repository; and
  • no more than one reviewer from any single model family.

If recruitment is unavailable, Stage A may complete, but Stage B must remain blocked rather than substituting same-model votes.

Two-stage execution

Stage A — Evidence construction and protocol freeze

Stage A creates the materials required for a valid blinded test. It does not score FEH-001.

A1. Freeze comparison dimensions

Before reviewing conclusions, define:

  • object of intervention;
  • input representation;
  • transformation mechanism;
  • decision rights;
  • output representation;
  • scope and exclusions;
  • temporal or coordination behavior;
  • uncertainty handling;
  • verification;
  • provenance;
  • failure modes;
  • predicted outcomes; and
  • implementation cost.

Do not add a dimension after viewing results unless it is recorded as exploratory and excluded from the primary decision.

A2. Build mechanism cards

Create one structured card per FE candidate and comparator mechanism. Each card must use the same template and maximum length.

Each card must:

  • avoid field names and branded terminology;
  • distinguish direct source statements from curator inference;
  • include the complete mechanism fields;
  • include at least one concrete implementation test;
  • include known limitations;
  • cite sources in a separate concealed key; and
  • receive an operational-completeness score before Stage B.

A3. Remove presentation confounds

The blinding editor must:

  • assign random card IDs;
  • normalize headings, length, reading level, and formatting;
  • remove author, repository, discipline, and provider names;
  • replace unique coined terms with neutral functional descriptions;
  • balance examples across cards;
  • remove links and citations from reviewer packets; and
  • retain a sealed mapping key.

Run a recognition pretest with at least two people not used as Stage B reviewers. Ask them to guess each card's source field and report confidence.

Recognition must not exceed:

  • mean source-identification accuracy of 0.40 across a minimum of four source categories; or
  • 0.60 accuracy for any FE card.

If the threshold is exceeded, revise the cards and repeat the pretest with new participants.

A4. Freeze the protocol

Before Stage B:

  • assign a protocol digest;
  • freeze source registry, cards, sealed key, reviewer instructions, rubric, analysis script or calculation sheet, thresholds, and exclusions;
  • record all deviations prospectively;
  • create a Stage B randomization seed; and
  • prohibit curators from communicating conclusions to reviewers.

Stage A acceptance gate

Stage B may begin only if:

  • every primary FE candidate has a complete or explicitly incomplete mechanism card;
  • every material comparator meets minimum source coverage;
  • blinding passes the recognition pretest;
  • the protocol and analysis are frozen;
  • the reviewer panel meets independence requirements; and
  • no reviewer has inspected the sealed key or another review.

If any condition fails, stop with blocked-before-stage-b.

Stage B — Blinded mapping and discrimination

B1. Reviewer task

For every FE candidate card, reviewers independently identify:

  1. the strongest comparator card or combination of cards;
  2. whether the mapping is exact, functionally-equivalent, partial, none, or indeterminate;
  3. which mechanism field is lost when the mapping is not exact;
  4. whether the lost field is material to a predicted outcome;
  5. whether the prediction is measurable and unique;
  6. whether the mechanism can be independently implemented; and
  7. confidence from 0 to 1.

Reviewers must construct the strongest subsumption account before recording a non-redundancy judgment.

B2. Mapping definitions

  • exact: all mechanism fields match without material additions.
  • functionally-equivalent: implementation differs, but scope and predicted outcomes are preserved.
  • partial: at least one material field or predicted outcome is lost.
  • none: no credible comparator supplies the core operation.
  • indeterminate: evidence or operational detail is insufficient.

Complexity, terminology, repository scale, documentation volume, integration, or application to frameworks does not by itself make a mapping partial.

A combination of comparator mechanisms is allowed. Penalize it only if the FE claim specifies and predicts a non-additive interaction or materially different execution cost.

B3. Prediction audit

For each alleged non-redundant mechanism, require:

  • a directional prediction;
  • a named baseline;
  • a measurable outcome;
  • a boundary condition;
  • a failure condition;
  • an implementation fidelity check; and
  • a future experiment capable of producing a null result.

Without all seven, classify the mechanism as conceptually-different-but-untested, not operationally non-redundant.

B4. Adjudication

Adjudicate disagreements after all reviews are locked.

The adjudicator must:

  • see individual rationales before reviewer identities;
  • resolve factual source disputes from the frozen evidence;
  • preserve unresolved conceptual disagreements;
  • never use majority vote to resolve a factual question;
  • record whether new evidence would change the decision; and
  • treat post-freeze sources as follow-up evidence, not silent corrections.

Primary outcomes

  1. Non-redundant mechanism count: number of candidates satisfying the complete non-redundancy rule below.
  2. Full-subsumption rate: proportion mapped exact or functionally-equivalent.
  3. Indeterminate rate: proportion lacking enough operational or comparator detail.
  4. Reviewer agreement: Krippendorff's alpha or a justified equivalent for categorical mapping judgments.
  5. Operational completeness: proportion of required mechanism fields present.
  6. Unique-prediction count: number with a measurable prediction not already produced by the mapped comparator.

Secondary and diagnostic outcomes

  • mapping confidence;
  • material-field disagreement;
  • source coverage by comparator;
  • recognition-pretest accuracy;
  • reviewer time;
  • card-reading burden;
  • number of comparator mechanisms needed per mapping;
  • sensitivity to treating combined comparator mechanisms as valid mappings;
  • sensitivity to excluding low-confidence judgments; and
  • difference between human and model judgments, reported descriptively.

Do not combine these into an unvalidated composite score.

Complete non-redundancy rule

A candidate counts as operationally non-redundant only when all conditions hold:

  1. at least two-thirds of eligible reviewers initially rate it partial or none;
  2. adjudication confirms a specific material field not supplied by the strongest comparator or comparator combination;
  3. operational completeness is 1.00;
  4. the unique prediction passes all seven prediction-audit requirements;
  5. no primary-source counterexample supplies the missing field;
  6. the judgment survives the preregistered sensitivity analyses; and
  7. at least one qualified human reviewer supports the material-field finding.

This rule identifies a candidate worthy of causal testing. It does not establish a discipline.

Decision rules

candidate-distinguishable

Use when at least one mechanism satisfies the complete non-redundancy rule and reviewer agreement is at least 0.67.

Consequence: preregister a focused outcome experiment on that mechanism only.

integration-hypothesis-only

Use when individual mechanisms are subsumed but a fully specified, non-additive integration prediction survives review.

Consequence: run an ablation or factorial experiment. Do not claim individual mechanism novelty.

fully-subsumed

Use when:

  • all operationally complete candidates are adjudicated exact or functionally-equivalent;
  • comparator coverage is sufficient;
  • agreement is at least 0.67; and
  • the indeterminate rate is no greater than 0.20.

Consequence: propose repositioning FE as a framework-focused method-engineering profile. Preserve useful terminology only when it improves implementation or communication.

inconclusive

Use when evidence, operationalization, coverage, blinding, independence, or agreement fails its threshold.

Consequence: identify the smallest repair. Do not rerun unchanged.

Falsification requirements

Attempt to falsify H1 by:

  1. mapping every FE candidate to a comparator or combination without material loss;
  2. searching primary sources for prior instances of the allegedly missing field;
  3. replacing FE terms with generic functional language;
  4. testing whether framework focus is only a domain specialization;
  5. testing whether integration is merely colocation rather than interaction;
  6. testing whether unique predictions collapse into generic structure, review effort, documentation, or provenance effects; and
  7. identifying a lower-complexity baseline that predicts the same outcome.

Attempt to falsify H0 by:

  1. finding a complete FE mechanism with no lossless comparator mapping;
  2. identifying a unique measurable prediction;
  3. specifying a task where the prediction diverges from adjacent methods; and
  4. showing that the difference is not attributable to unequal effort or presentation.

Analysis plan

  1. Lock and hash all reviewer responses before unblinding.
  2. Report raw reviewer mappings.
  3. Compute agreement with missing and indeterminate judgments preserved.
  4. Apply decision rules exactly as preregistered.
  5. Run sensitivity analyses:
    • combined comparators allowed versus single comparator only;
    • all judgments versus confidence at least 0.70;
    • human-only versus all reviewers, descriptively;
    • incomplete mechanisms excluded versus treated as indeterminate.
  6. Unblind only after primary calculations are frozen.
  7. Compare recognition guesses with actual sources as a validity check.
  8. Report deviations and their likely direction of bias.
  9. Separate confirmatory results from exploratory observations.

Do not use significance testing with an underpowered convenience panel. Report counts, proportions, agreement, uncertainty intervals where defensible, and the evidence needed to change the classification.

Efficiency controls

Use sequential gates to avoid spending resources on an invalid study:

  1. stop if mechanism operationalization fails;
  2. stop if primary-source coverage fails;
  3. stop if blinding fails;
  4. stop if independent reviewers cannot be recruited;
  5. proceed to adjudication only for disagreements or alleged non-redundancy; and
  6. proceed to a causal utility experiment only for surviving mechanisms.

Limit the primary packet to:

  • no more than 10 FE mechanism cards;
  • no more than 24 comparator cards;
  • no more than 500 words per card; and
  • a target reviewer session of 120 minutes or less.

If the packet exceeds the target, run a balanced incomplete-block design while ensuring every FE candidate receives at least three independent reviews and every reviewer receives anchor cards.

Automate formatting, randomization, hashing, and descriptive calculations. Do not automate source interpretation or factual adjudication.

Threats and required mitigations

Threat Required mitigation
Terminology novelty Neutral functional language and sealed source key
Unequal documentation Fixed card schema and length
Curator allegiance Separate FE and comparator curators
Same-model priors Human expert plus provider-family cap
Recognition leakage Pretest with hard thresholds
Weak comparator Strongest-version steelman and primary sources
Moving criteria Protocol digest and prospective deviations
Majority-vote error Evidence-based adjudication
Integration inflation Require non-additive interaction prediction
Complexity confounding Record implementation cost; defer causal claims
Missing evidence Preserve indeterminate; do not impute support
Post hoc novelty Primary-source counterexample search before acceptance

Stop conditions

Stop immediately and report the reason if:

  • source identity cannot be concealed;
  • a material comparator lacks minimum source coverage;
  • the sealed key is exposed to a reviewer;
  • a reviewer sees another response before locking their own;
  • the protocol changes after reviews begin without a declared deviation;
  • fewer than three eligible independent reviews exist per candidate;
  • the required human expertise is unavailable; or
  • a conflict of interest prevents adversarial evaluation.

Do not repair a compromised run in place. Version the protocol and create a new run.

Required output package

Create:

research/evaluations/FE-BOUNDARY-YYYY-MM-DD/
  README.md
  protocol.md
  protocol-lock.json
  source-registry.json
  mechanism-card-schema.json
  cards/
    blinded/
    sealed-source-key/
  recognition-pretest/
  reviewer-packets/
  reviews/
  adjudication/
  analysis/
  decision.md
  research-journal.md
  HANDOFF.md

The sealed key and any identity-bearing material must be access-controlled until primary analysis is locked. If the repository cannot provide access separation, store sealed materials outside reviewer-visible paths and record their cryptographic digests in protocol-lock.json.

Every artifact must include:

  • stable ID;
  • version;
  • status;
  • author or agent;
  • creation and update timestamps;
  • parent and source dependencies;
  • repository commit or external digest;
  • change summary;
  • confidence;
  • completion state; and
  • known limitations.

Required final report

Return:

  1. completion status and protocol deviations;
  2. boundary classification;
  3. mechanism-level mapping table;
  4. strongest evidence for and against each surviving candidate;
  5. reviewer agreement and recognition results;
  6. falsification attempts;
  7. validity threats;
  8. sensitivity analyses;
  9. evidence that would change the result;
  10. affected downstream records;
  11. the smallest justified next experiment; and
  12. a concise operator handoff.

Separate:

  • direct observations;
  • curator interpretations;
  • reviewer judgments;
  • adjudicator decisions;
  • exploratory findings; and
  • recommendations.

Research-director challenge log

The following design alternatives were considered and rejected:

Immediate FE-versus-baseline utility trial

Rejected as the next step because “FE” is not operationally stable. A null result would be uninterpretable, while a positive result could be caused by added structure, documentation, time, or evaluator preference.

Another unblinded multi-model opinion survey

Rejected because EX-FE-0001 already showed that model agreement cannot establish distinctiveness and shared training priors remain uncontrolled.

Discipline-level literature comparison alone

Rejected because fields overlap internally and labels are too coarse. The mechanism is the correct unit of analysis.

Single-comparator matching

Rejected because adjacent practices may jointly subsume a mechanism. Combination mappings are allowed unless FE specifies a genuine interaction effect.

Novelty as the success criterion

Rejected because a useful integrated profile need not be historically novel, and a novel mechanism need not be useful. This experiment tests non-redundancy; causal utility remains a later experiment.

Majority agreement as adjudication

Rejected because reviewers can share the same misconception. Factual disputes must be resolved from frozen primary evidence.

Large exhaustive source corpus

Rejected for the first decisive pass. Strongest-version comparator coverage, mechanism saturation, and sequential gates provide more decision value per unit cost. Expand only when a missing source could change a material mapping.

Completion check

Before declaring the experiment complete, confirm:

  • Stage A and Stage B statuses are reported separately.
  • No canonical theory was mutated.
  • Comparison dimensions were frozen before conclusions.
  • Every material comparator has primary-source coverage.
  • Every candidate uses the same mechanism-card schema.
  • Blinding passed the recognition threshold.
  • Reviewer independence requirements were met.
  • Every reviewer response was locked before unblinding.
  • Both H1 and H0 received genuine falsification attempts.
  • Combination mappings were considered.
  • Indeterminate results were preserved.
  • Primary and sensitivity analyses were reported.
  • Model agreement was not treated as validation.
  • The decision rule was applied without post hoc relaxation.
  • The next experiment is limited to surviving uncertainty.

If any item is unchecked, report incomplete or blocked and identify the minimum repair.