research-document

FE-012B Methodological Review

FE-012B Methodological Review

1. What failed?

No total synthesis collapse occurred in the internal run, but three problems produced clear missing-primitive pressure:

  • coordinating autonomous underwater robots
  • coordinating lunar habitat maintenance
  • managing restoration of a historic cathedral

The weakest cases lost specificity where the vocabulary needed to express timing, dependency, or synchronization rather than only judgment and action.

2. Which primitives were overused?

  • Observe
  • Evaluate
  • Decide
  • Reassess

These primitives appeared in many syntheses because they are broadly applicable.

That helps transfer, but it also risks flattening different procedural frameworks into similar-looking loops.

3. Which primitives were unnecessary?

No primitive was globally unnecessary, but some were only lightly used in this run:

  • Classify
  • Explain
  • Compare
  • Reflect

Their lower usage does not show they are dispensable.

It shows that this particular novelty set favored coordination, prioritization, and intervention more than taxonomy-heavy or retrospective reasoning.

4. Which primitives appear indispensable?

  • Bound
  • Observe
  • Evaluate
  • Decide
  • Act

Most syntheses became noticeably weaker when one of these functions was absent.

Prioritize also appeared close to indispensable in constrained resource problems.

5. Which problems resisted synthesis?

The most resistant problems were:

  • coordinating autonomous underwater robots
  • coordinating lunar habitat maintenance
  • managing restoration of a historic cathedral

These cases strained the vocabulary because they required long-range coordination, sequencing discipline, or synchronization pressure without allowing direct borrowing from known operational frameworks.

6. What would falsify the hypothesis next?

  • Repeated synthesis failure across a larger and independently reviewed novelty set
  • Persistent reviewer findings that the frameworks are coherent only in generic form and break under domain detail
  • Frequent need for missing primitive requests
  • Independent analysts producing inconsistent compositions for the same problem
  • Evidence that synthesized structures are merely thin rephrasings of remembered frameworks rather than genuine primitive compositions

Overall Assessment

The internal run did not falsify the hypothesis.

It also did not validate it strongly.

The main next falsification risk is not total failure.

It is that the primitive vocabulary may be expressive only at a generic procedural level and may not preserve enough structure to generate distinctive frameworks reliably, especially when timing and coordination become first-class requirements.