Self-Improvement
Self-improvement uses run results, findings and human feedback to propose changes to how the factory works. Each improvement has its own evidence, evaluation and approval process before it changes future work.
Improvement targets
Section titled “Improvement targets”A candidate identifies the configuration to change and the result it should improve: quality, cost, completion time or human effort.
| Target | Possible changes |
|---|---|
| Skills | Primary instructions, additional skills and evaluation criteria. |
| Task configuration | Harness, model, effort, context selection and tools. |
| Workflows | Steps, parallel work, routing and revision limits. |
| Standards and checks | Engineering rules, validators and acceptance requirements. |
| Policies | Required reviews, reviewer eligibility and approval requirements. |
Improvement lifecycle
Section titled “Improvement lifecycle”Improvement work has a separate lifecycle from the delivery run that revealed the problem. Completing that run provides evidence; it does not approve a change to the factory.
The intended lifecycle separates these stages. They describe the improvement process, rather than a single record’s status values.
| Stage | Result |
|---|---|
| Evidence | A documented problem or opportunity, linked to run results, findings or review history. |
| Candidate | A specific proposed change, its expected effect and an evaluation plan. |
| Evaluation | A comparison with the existing configuration, including regressions and trade-offs. |
| Review | A decision to adopt, revise, reject or defer the candidate. |
| Activation | The approved version becomes available for future work through the target’s publication or activation controls. |
| Follow-up | Subsequent results establish whether to retain, revise or replace the change. |
A successful evaluation supports approval; approval authorises adoption. The relevant workflow, skill, standard or policy still has its own versioning and activation controls.
Retrospectives
Section titled “Retrospectives”The defined retrospective workflow takes a completed workflow run as its subject and authors a retrospective record. Five evaluations assess the record before its findings are collated.
| Evaluation | Focus |
|---|---|
| Signal quality | Whether the observations provide useful improvement signals. |
| Miss classification | How failures and missed expectations are classified. |
| Guardrail quality | The proposed controls for preventing recurrence. |
| Evidence grounding | Whether the conclusions are supported by run evidence. |
| Actionability | Whether the recommendations are specific enough to act on. |
Findings can return the record for revision within a defined limit. The accepted artifact supplies evidence for follow-up improvement work; it does not update a workflow, skill or policy. The current retrospective definition does not automatically launch that follow-up work.
Candidate evaluation
Section titled “Candidate evaluation”Compare the candidate with the existing version on representative work, using the same acceptance criteria. Runtime Evaluations provide the quality, cost and time evidence for that comparison; review history helps assess the human effort involved.
Keep the source evidence, proposed configuration and comparison results together. A cheaper model may require more revisions; an additional check may catch defects while increasing completion time. The candidate should make those trade-offs explicit before review.
Adoption and history
Section titled “Adoption and history”Adoption follows the controls for the thing being changed: publishing a workflow version, updating a skill, publishing a standard or activating a policy version. Improvement work does not grant itself permission to weaken checks or remove approvals.
The governance ledger preserves governed review and decision history. Results from subsequent runs provide evidence for the next improvement, including changes to the retrospective workflow itself.