Evolving Agents Labs is an
open-source lab, and three of its projects are active.
lora-kernel serves small local models as trained specialists. agentvcs versions and merges
the harness those models run in, and attributes each change in a result to the change that caused it.
gene-evidence writes evidence reports on predicted genes in which every sentence cites its source. The idea
they share: what an AI system answers, changes or reports should be traceable to what produced it.
One connection is built today; two are planned — how they connect.
Can one small local model, given a trained specialist per job and a library of
notes it has learned to navigate, do an organisation’s routine work — and hand the rest
to a frontier model?
active · flagship An operating layer for
specialized local AI agents: one resident model, a LoRA per job, facts in notes a person can edit,
a router that abstains.
Built: specialists that beat the bare base on their job, a router that serves
no foreign text locally, edits answered without retraining — measured.
Not yet: no installable package, no real traffic.
Why it matters: if it works, an organisation’s routine AI work is
served from one small local GPU — private, cheap to run, its facts in notes a person can
read and edit — and only the rest is sent to a frontier model. The saving is not measured yet.
When an agent’s harness changes while it runs, can every result still be
traced to the change that caused it — and can two diverging harnesses be merged on the
evidence?
active · Rust core v0.1 Version control for
agent harnesses: every step stamped with the harness version that produced it, patches gated,
blame per metric, and a three-way merge whose real conflicts go to Claude Code —
its only LLM.
Built: the Rust core, CLI, MCP server and Python SDK; attribution on a real
model (below). Not yet: no PyPI package; no real Claude Code merge
session recorded in the repository.
Why it matters: harnesses now change while agents run — a prompt, a
model, a router. If it works, you know which change moved which result, and two harnesses that
diverged are merged on evidence rather than by guess. Attribution has run on one real model.
Can every sentence of a report on a predicted gene be traced to the tool output
it rests on?
prototype · stage 1 Auditable evidence
reports for gene predictions the reference annotation does not have: an evidence graph of tool
outputs, and a report that cites it in every sentence.
Built: stage 1, deterministic, no language model; a development run on rat.
Not yet: a model that drafts the report. Known limitation: it explains a candidate; it
does not pick which to validate first.
Why it matters: a wet-lab experiment on a candidate gene costs real money.
If it works, every claim in the evidence report can be checked against its source before that
money is spent. The checkable report is built; choosing what to test is not.