research-document FE-EXPERIMENT-SYSTEM-ASSESSMENT-2026-07-23
Repository and research-state assessment
Repository and research-state assessment
Repository basis
research/evaluations/FE-EVAL-2026-07-23/framework-engineering-executive-brief.md(FE-EVAL-BRIEF-2026-07-23) is the best concise current evaluation, not constitutional authority.research/evaluations/FE-EVAL-2026-07-23/RP-2026-07-23-framework-engineering-evaluation.md(FE-EVAL-REP-2026-07-23) is the parent execution package.research/theory/theory-of-framework-engineering-v0.1.md(FE-THEORY-0.1) is the best current theory and remains provisional.research/operating-system/research-constitution.mdgoverns evidence behavior; internal model agreement is evidence, not validation.CURRENT_STATE.mdcontrols operational state subject to its human-decision section.- ECR-000003 board updates and
framework-engineering-registry-updates.mdare proposal-only.
Existing ECR directories already separate packets, provider responses, comparison, review, and dashboards, but they do not supply a reusable provider-neutral experiment contract or write-once runner. The new system extends that separation under research/framework-engineering/experiments/; it does not rename or migrate historical ECR records.
Scientific state
The repository supports structured traceability and a narrow comparison pipeline, but not distinctiveness or causal advantage. FEH-001 (defensible boundary), FEH-002/003/004/007/010 (continuation and provenance value), and FEH-005/006/008 (safe concurrency/governance) remain open. The largest unknown is incremental utility against complexity-matched alternatives.
The immediate portfolio is dependency ordered:
EX-FE-0001, boundary-discrimination infrastructure pilot: high information value, low cost, LLM/human judgment appropriate only as protocol input.FE-EXP-BOUNDARY-001: primary literature plus blinded expert classification; high value, moderate cost.FE-EXP-COLDSTART-001: continuation/provenance A/B/C test; high value, blocked on authority/ID policy.FE-EXP-UTILITY-001: matched utility experiment; highest causal value, blocked on boundary work.FE-EXP-HUMAN-001: independent human baseline; empirical and recruitment dependent.FE-EXP-CONCURRENCY-001: failure-injection governance test; blocked on provenance evidence and minimal contract.
The pilot was selected because it exercises schemas, prompt generation, input allowlisting, isolation, manual import, checksums, comparison templates, and synthesis lineage without fabricating provider evidence or prematurely running the blocked concurrency mission.
Conflicts and resolved assumptions
- “Latest” does not mean “accepted”; status and review gates control authority.
- Markdown called immutable is not inherently immutable. The implementation combines create-only runner writes with Git-retrievable history.
- Stable IDs alone do not prove continuity benefit.
- Provider executables on PATH do not prove authentication, model version, command compatibility, or permission to spend.
- A multi-provider result cannot establish empirical validity through agreement.