From conversation to action
Tool use and coding agents connected reasoning to files, terminals and tests, letting results be checked against the real environment.
Reasoning meets realityEVOLUTION / AN ENGINEERING DIRECTION
Models improve over time. Memory, skills, tools and workflows can improve too. EvoForge is developing ways to carry useful experience forward, with evaluation and approval keeping each change accountable and under your control.
Tool use, persistent memory, reusable skills and evaluation loops form several parallel paths toward agents that learn from real work.
Tool use and coding agents connected reasoning to files, terminals and tests, letting results be checked against the real environment.
Reasoning meets realityResident agents such as OpenClaw brought channels, memory and extensible skills together, allowing work and context to continue across time.
Work gains continuitySkills package instructions, references and scripts as reusable methods. Review and maintenance keep those methods relevant as the system grows.
Experience becomes reusableSystems such as Hermes connect persistent memory with skill revision. Self-evolution research extends evaluation to prompts, skills, tool descriptions and code.
Change becomes measurableEvoForge is exploring a metacognitive framework that recognizes recurring gaps, identifies whether the next improvement belongs in knowledge, procedure, tools or models, and evaluates the result.
The system reasons about its own growthWritten in Rust. Replaceable models connect at the top; your computer, apps and tools form the working environment below. Context, memory, permissions and orchestration connect the two.
Attachments and tool results are organized within a context budget, helping long tasks stay focused and continuous.
Ninety-eight tools cover document processing, code editing, browser and desktop interaction, terminal work and media creation.
Facts, preferences and project context are kept in the right scope, with expiry and recall controls that keep future work relevant.
Risk is assessed using the target, data sensitivity, recovery options and external impact, guided by the autonomy level you choose.
A lead agent or shared task board coordinates specialists, long-running work and scheduled tasks, with progress gathered in one place.
Connect cloud or local models and choose the one that best fits each task, with provider health and credentials managed separately.
Foundation models supply the reasoning. EvoForge keeps the memory, skills, specialists and evidence that make that reasoning useful in your world, then carries them to whichever model you choose next.
Observation, proposal, comparison, adoption and restoration remain distinct stages. Candidate changes are independently evaluated before they influence future work.
A short pass over what just happened: a fact worth keeping, a preference you stated, a correction you made. Cheap enough to run every time, small enough to stay out of your way.
Before an action it can write down what it expects to see. A mismatch is a surprise, and surprise is what triggers a deeper reflection. Calibration is measured, so confident-and-wrong is caught, not hidden.
In the background it re-reads recent work, merges duplicate memories, retires stale ones, and turns a repeated successful trajectory into a candidate skill. Candidates go to the loop above; nothing is promoted by dreaming alone.
Memory, skills, tools, specialists, policy and evaluation results remain inspectable assets you can open, test, version, share, disable or restore. They move with your work when you change models.
Durable facts, preferences, project context, and learned constraints, each with a scope and an optional expiry.
Fifty-four skills are built in. New skills use the same format and are checked for when and how they should be used.
Plugins, connectors and in-process tools add new ways to work, introduced through explicit review and your permissions.
Rules include self-tests that help detect problems early and present them clearly.
Persistent identities with their own brief, memory, tools, permissions, and lifecycle. Colleagues, not disposable personas.
Comparable cases, regression results, your decisions, attribution, and reversal history, in a ledger you can replay.
Candidates are independently tested and compared before adoption. Evidence, not confidence, determines what carries forward.
Candidate skills and rules are compared on representative tasks in an evaluation environment separated from your working data.
A test used to validate a fix reproduces the original issue, confirms the change, and remains for future work.
Comparable cases, results, decisions and restoration history remain together in a record you can review.
A candidate skill or rule is compared against the baseline on the same tasks inside an evaluation host that cannot touch your real data. The verdict is written to the ledger before anything is applied.
It mines your own git history for tasks it once did wrong, replays them two ways, and turns the difference between the failed and the real fix into a lesson. The curriculum grows from your work, not from a benchmark.
The longer-term vision is a distributed network where people exchange evaluated skills, tools and experience while deciding what to trust and adopt. This network remains in research. Private memory stays local by default, and sharing requires authorization.
Two independent toggles in the app. Everything they produce is visible in the memory and skills panels, where any item can be disabled or rolled back with one action.
Consolidates memory and proposes skills during idle time. A separate approval setting controls whether new skills are enabled automatically.
When off, every new skill waits for your review. Turn it on only when you want evaluated candidates to be adopted automatically.