research-document

FE-011A Pilot Run Log

FE-011A Pilot Run Log

Run 1

  • run_id: FE-011A-SWOT-001
  • date: 2026-06-28
  • model: GPT-5 Codex
  • framework: SWOT
  • version_label: Version B
  • packet_path: research/experiments/FE-011A-llm-blind-pilot/framework-packets/swot-version-b.md
  • scenario_path: research/experiments/FE-011A-llm-blind-pilot/scenarios/swot-scenario.md
  • participant_prompt_path: research/experiments/FE-011A-llm-blind-pilot/prompts/participant-prompt.md
  • output_path: research/experiments/FE-011A-llm-blind-pilot/results/swot/output-1.md
  • evaluator: GPT-5 Codex blinded evaluator pass
  • evaluator_output_path: research/experiments/FE-011A-llm-blind-pilot/results/swot/evaluator-review.md
  • notes: Output was stored as Output 1 for blinded comparison.

Run 2

  • run_id: FE-011A-SWOT-002
  • date: 2026-06-28
  • model: GPT-5 Codex
  • framework: SWOT
  • version_label: Version A
  • packet_path: research/experiments/FE-011A-llm-blind-pilot/framework-packets/swot-version-a.md
  • scenario_path: research/experiments/FE-011A-llm-blind-pilot/scenarios/swot-scenario.md
  • participant_prompt_path: research/experiments/FE-011A-llm-blind-pilot/prompts/participant-prompt.md
  • output_path: research/experiments/FE-011A-llm-blind-pilot/results/swot/output-2.md
  • evaluator: GPT-5 Codex blinded evaluator pass
  • evaluator_output_path: research/experiments/FE-011A-llm-blind-pilot/results/swot/evaluator-review.md
  • notes: Output was stored as Output 2 for blinded comparison.

Run 3

  • run_id: FE-011A-5W-001
  • date: 2026-06-28
  • model: GPT-5 Codex
  • framework: Five Whys
  • version_label: Version B
  • packet_path: research/experiments/FE-011A-llm-blind-pilot/framework-packets/five-whys-version-b.md
  • scenario_path: research/experiments/FE-011A-llm-blind-pilot/scenarios/five-whys-scenario.md
  • participant_prompt_path: research/experiments/FE-011A-llm-blind-pilot/prompts/participant-prompt.md
  • output_path: research/experiments/FE-011A-llm-blind-pilot/results/five-whys/output-1.md
  • evaluator: GPT-5 Codex blinded evaluator pass
  • evaluator_output_path: research/experiments/FE-011A-llm-blind-pilot/results/five-whys/evaluator-review.md
  • notes: Output was stored as Output 1 for blinded comparison.

Run 4

  • run_id: FE-011A-5W-002
  • date: 2026-06-28
  • model: GPT-5 Codex
  • framework: Five Whys
  • version_label: Version A
  • packet_path: research/experiments/FE-011A-llm-blind-pilot/framework-packets/five-whys-version-a.md
  • scenario_path: research/experiments/FE-011A-llm-blind-pilot/scenarios/five-whys-scenario.md
  • participant_prompt_path: research/experiments/FE-011A-llm-blind-pilot/prompts/participant-prompt.md
  • output_path: research/experiments/FE-011A-llm-blind-pilot/results/five-whys/output-2.md
  • evaluator: GPT-5 Codex blinded evaluator pass
  • evaluator_output_path: research/experiments/FE-011A-llm-blind-pilot/results/five-whys/evaluator-review.md
  • notes: Output was stored as Output 2 for blinded comparison.

Run 5

  • run_id: FE-011A-OODA-001
  • date: 2026-06-28
  • model: GPT-5 Codex
  • framework: OODA
  • version_label: Version A
  • packet_path: research/experiments/FE-011A-llm-blind-pilot/framework-packets/ooda-version-a.md
  • scenario_path: research/experiments/FE-011A-llm-blind-pilot/scenarios/ooda-scenario.md
  • participant_prompt_path: research/experiments/FE-011A-llm-blind-pilot/prompts/participant-prompt.md
  • output_path: research/experiments/FE-011A-llm-blind-pilot/results/ooda/output-1.md
  • evaluator: GPT-5 Codex blinded evaluator pass
  • evaluator_output_path: research/experiments/FE-011A-llm-blind-pilot/results/ooda/evaluator-review.md
  • notes: Output was stored as Output 1 for blinded comparison.

Run 6

  • run_id: FE-011A-OODA-002
  • date: 2026-06-28
  • model: GPT-5 Codex
  • framework: OODA
  • version_label: Version B
  • packet_path: research/experiments/FE-011A-llm-blind-pilot/framework-packets/ooda-version-b.md
  • scenario_path: research/experiments/FE-011A-llm-blind-pilot/scenarios/ooda-scenario.md
  • participant_prompt_path: research/experiments/FE-011A-llm-blind-pilot/prompts/participant-prompt.md
  • output_path: research/experiments/FE-011A-llm-blind-pilot/results/ooda/output-2.md
  • evaluator: GPT-5 Codex blinded evaluator pass
  • evaluator_output_path: research/experiments/FE-011A-llm-blind-pilot/results/ooda/evaluator-review.md
  • notes: Output was stored as Output 2 for blinded comparison.

Execution Notes

  • This pilot was executed as an internal single-agent run.
  • The same model family generated participant outputs and evaluator reviews.
  • Findings should therefore be treated as pilot evidence about the instrument, not as independent validation.