EVOLUTION / AN ENGINEERING DIRECTION

From doing the work
to improving how it gets done.

Models improve over time. Memory, skills, tools and workflows can improve too. EvoForge is developing ways to carry useful experience forward, with evaluation and approval keeping each change accountable and under your control.

01 / THE TRAJECTORYACTION → PERSISTENCE → LEARNING

Agent evolution has a history.
Each step opens a new possibility.

Tool use, persistent memory, reusable skills and evaluation loops form several parallel paths toward agents that learn from real work.

01
ACTION

From conversation to action

Tool use and coding agents connected reasoning to files, terminals and tests, letting results be checked against the real environment.

Reasoning meets reality
02
PERSISTENCE · OPENCLAW

An agent gets a place to live

Resident agents such as OpenClaw brought channels, memory and extensible skills together, allowing work and context to continue across time.

Work gains continuity
03
PROCEDURAL MEMORY · SKILLS

Experience becomes procedure

Skills package instructions, references and scripts as reusable methods. Review and maintenance keep those methods relevant as the system grows.

Experience becomes reusable
04
LEARNING LOOPS · HERMES

Candidates enter evaluation

Systems such as Hermes connect persistent memory with skill revision. Self-evolution research extends evaluation to prompts, skills, tool descriptions and code.

Change becomes measurable
05
METACOGNITION · EVOFORGE

The system learns where to improve

EvoForge is exploring a metacognitive framework that recognizes recurring gaps, identifies whether the next improvement belongs in knowledge, procedure, tools or models, and evaluates the result.

The system reasons about its own growth
02 / THE KERNELFOUR LAYERS · ONE CODEBASE

One integrated system,
from context to action.

Written in Rust. Replaceable models connect at the top; your computer, apps and tools form the working environment below. Context, memory, permissions and orchestration connect the two.

FIG.01 / KERNEL · GENERAL ARRANGEMENTSECTION VIEW
LAYER 0 · MODELS · REPLACEABLE CLAUDE GPT GEMINI DEEPSEEK · QWEN · KIMI LOCAL · OLLAMA ONE DIALECT FILE PER VENDOR · CIRCUIT BREAKER PER CREDENTIAL LAYER 1 · KERNEL · BUILT HERE CONTEXTBUDGETED · COMPACTED TOOLS ×98NATIVE · LINTED MEMORYUSER / PROJECT · EXPIRY APPROVALSRISK BY IMPACT · SANDBOX ORCHESTRATIONLEAD · SWARM · GOALS EVOLVEGATES · LEDGER SKILLS ×54 · PLUGINS · CONNECTORS MEDIA: IMAGE · VOICE · MUSIC · VIDEO · DIRECTOR'S CONSOLE LAYER 2 · BODY · THE WHOLE MACHINE BROWSER DESKTOP APPS SHELL · PTY SERVER FLEET MAIL · CHANNELS LAYER 3 · YOU · THE CONTROL SURFACE DESKTOP APP · APPROVE · INTERRUPT · UNDO · SWITCH MODEL · TURN EVOLUTION ON OR OFF
K-01

Context engine

Attachments and tool results are organized within a context budget, helping long tasks stay focused and continuous.

K-02

Built-in tools

Ninety-eight tools cover document processing, code editing, browser and desktop interaction, terminal work and media creation.

K-03

Memory

Facts, preferences and project context are kept in the right scope, with expiry and recall controls that keep future work relevant.

K-04

Permissions and sandbox

Risk is assessed using the target, data sensitivity, recovery options and external impact, guided by the autonomy level you choose.

K-05

Orchestration

A lead agent or shared task board coordinates specialists, long-running work and scheduled tasks, with progress gathered in one place.

K-06

Model layer

Connect cloud or local models and choose the one that best fits each task, with provider health and credentials managed separately.

03 / MODEL INDEPENDENCEINTELLIGENCE ABOVE THE PROVIDER

Models change.
What you have built does not reset.

Foundation models supply the reasoning. EvoForge keeps the memory, skills, specialists and evidence that make that reasoning useful in your world, then carries them to whichever model you choose next.

FIG.02 / MODEL-INDEPENDENT COMPOUNDINGSYSTEM VIEW
MODEL AMODEL BMODEL CLOCAL MODELYOUR ACCUMULATED INTELLIGENCEMEMORY · SKILLS · SPECIALISTS · POLICY · EVIDENCEREAL WORK ACROSS TIME
04 / THE LOOPFROM EXPERIENCE TO IMPROVEMENT

Turn experience into tested improvement.

Observation, proposal, comparison, adoption and restoration remain distinct stages. Candidate changes are independently evaluated before they influence future work.

FIG.03 / GOVERNED EVOLUTION PIPELINETRACE MODE
EXECUTEREAL TASKEVIDENCEWHAT HAPPENEDPROPOSECANDIDATE ASSETCOMPAREISOLATED HOSTDECIDEPROMOTE / REJECT REVIEW CHANGES · RESTORE SUPPORTED VERSIONS #01 · a91f#02 · 3c0e#03 · d7b2#04 · e41a#05 · 08fc#06 · b6d9#07 · 52e7#08 · f0a3#09 · 9c1b#10 · 7e44 INPUT: EXECUTION TRAJECTORYOUTPUT: DURABLE ASSETHASH-CHAINED LEDGER · EVERY DECISION ATTRIBUTABLE · REPLAYABLE
05 / WHEN IT LEARNSTHREE RHYTHMS
01 / EVERY TURN

A micro-reflection after each exchange

A short pass over what just happened: a fact worth keeping, a preference you stated, a correction you made. Cheap enough to run every time, small enough to stay out of your way.

02 / ON SURPRISE

When the world disagrees with its expectation

Before an action it can write down what it expects to see. A mismatch is a surprise, and surprise is what triggers a deeper reflection. Calibration is measured, so confident-and-wrong is caught, not hidden.

03 / WHILE IDLE

Dreaming: trajectories become skills

In the background it re-reads recent work, merges duplicate memories, retires stale ones, and turns a repeated successful trajectory into a candidate skill. Candidates go to the loop above; nothing is promoted by dreaming alone.

06 / WHAT ACCUMULATESINSPECTABLE ASSETS

What it learns stays visible.

Memory, skills, tools, specialists, policy and evaluation results remain inspectable assets you can open, test, version, share, disable or restore. They move with your work when you change models.

01

Memory

Durable facts, preferences, project context, and learned constraints, each with a scope and an optional expiry.

02

Skills

Fifty-four skills are built in. New skills use the same format and are checked for when and how they should be used.

03

Capabilities

Plugins, connectors and in-process tools add new ways to work, introduced through explicit review and your permissions.

04

Policy

Rules include self-tests that help detect problems early and present them clearly.

05

Specialist agents

Persistent identities with their own brief, memory, tools, permissions, and lifecycle. Colleagues, not disposable personas.

06

Selection evidence

Comparable cases, regression results, your decisions, attribution, and reversal history, in a ledger you can replay.

07 / EVALUATIONINDEPENDENT TESTING BEFORE ADOPTION

Improvements should
stand up to testing.

Candidates are independently tested and compared before adoption. Evidence, not confidence, determines what carries forward.

Isolated evaluation

Candidate skills and rules are compared on representative tasks in an evaluation environment separated from your working data.

Regression checks

A test used to validate a fix reproduces the original issue, confirms the change, and remains for future work.

Evidence ledger

Comparable cases, results, decisions and restoration history remain together in a record you can review.

AB02

Isolated evaluation

A candidate skill or rule is compared against the baseline on the same tasks inside an evaluation host that cannot touch your real data. The verdict is written to the ledger before anything is applied.

03GIT HISTORY → TASKS

Self-curriculum

It mines your own git history for tasks it once did wrong, replays them two ways, and turns the difference between the failed and the real fix into a lesson. The curriculum grows from your work, not from a benchmark.

08 / THE HORIZONRESEARCH VISION

From personal node
to a network of verified capability.

The longer-term vision is a distributed network where people exchange evaluated skills, tools and experience while deciding what to trust and adopt. This network remains in research. Private memory stays local by default, and sharing requires authorization.

09 / THE SWITCHESTWO TOGGLES · BOTH YOURS

Evolution runs only as far as you let it.

Two independent toggles in the app. Everything they produce is visible in the memory and skills panels, where any item can be disabled or rolled back with one action.

DEFAULT · ON

Organize memory in the background

Consolidates memory and proposes skills during idle time. A separate approval setting controls whether new skills are enabled automatically.

DEFAULT · OFF

Approve improvements automatically

When off, every new skill waits for your review. Turn it on only when you want evaluated candidates to be adopted automatically.

Every task leaves behind
something you keep.