Evolving Agents Labs every number checked nightly

Evolving Agents Labs is an open-source lab, and three of its projects are active. lora-kernel serves small local models as trained specialists. agentvcs versions and merges the harness those models run in, and attributes each change in a result to the change that caused it. gene-evidence writes evidence reports on predicted genes in which every sentence cites its source. The idea they share: what an AI system answers, changes or reports should be traceable to what produced it. One connection is built today; two are planned — how they connect.

lora-kernel

Can one small local model, given a trained specialist per job and a library of notes it has learned to navigate, do an organisation’s routine work — and hand the rest to a frontier model?

active · flagship  An operating layer for specialized local AI agents: one resident model, a LoRA per job, facts in notes a person can edit, a router that abstains.

Built: specialists that beat the bare base on their job, a router that serves no foreign text locally, edits answered without retraining — measured. Not yet: no installable package, no real traffic.

Why it matters: if it works, an organisation’s routine AI work is served from one small local GPU — private, cheap to run, its facts in notes a person can read and edit — and only the rest is sent to a frontier model. The saving is not measured yet.

agentvcs

When an agent’s harness changes while it runs, can every result still be traced to the change that caused it — and can two diverging harnesses be merged on the evidence?

active · Rust core v0.1  Version control for agent harnesses: every step stamped with the harness version that produced it, patches gated, blame per metric, and a three-way merge whose real conflicts go to Claude Code — its only LLM.

Built: the Rust core, CLI, MCP server and Python SDK; attribution on a real model (below). Not yet: no PyPI package; no real Claude Code merge session recorded in the repository.

Why it matters: harnesses now change while agents run — a prompt, a model, a router. If it works, you know which change moved which result, and two harnesses that diverged are merged on evidence rather than by guess. Attribution has run on one real model.

gene-evidence

Can every sentence of a report on a predicted gene be traced to the tool output it rests on?

prototype · stage 1  Auditable evidence reports for gene predictions the reference annotation does not have: an evidence graph of tool outputs, and a report that cites it in every sentence.

Built: stage 1, deterministic, no language model; a development run on rat. Not yet: a model that drafts the report. Known limitation: it explains a candidate; it does not pick which to validate first.

Why it matters: a wet-lab experiment on a candidate gene costs real money. If it works, every claim in the evidence report can be checked against its source before that money is spent. The checkable report is built; choosing what to test is not.

Three workbenches in one long room. On the left a card-catalogue cabinet, three of its drawers lit by a small brass radar dish; in the middle a brass machine whose paper ledger carries wax seals that change colour at a swapped gear; on the right a naturalist's bench with a printed report tied by red threads to pinned evidence cards. A solid green cord joins the first two benches; two dashed ochre chalk lines lead from the third bench to the other two.
Three benches, one room: lora-kernel, agentvcs, gene-evidence. The solid cord exists today; the dashed lines are planned.

What they share is one rule: a statement that does not cite its source does not ship. gene-evidence’s check-report fails a sentence whose citation does not resolve; lora-kernel’s endpoint withholds an answer whose citation fails the referee’s check.

The flagship · lora-kernel · Apache 2.0

The LoRA is not the textbook. It is the specialist who knows how to use the library.

One GPU, one small resident model, and a shelf of QLoRA adapters — each an expert defined by its corpus. An OpenAI-compatible API serves each request with the matching expert and sends everything else to a frontier model. Facts live in short notes a person can read and correct; the adapter holds the habit of navigating them.

What is measured, and what is not → The architecture article Repository

A reader at a desk before a wall of card-catalogue drawers labelled Procedures and Encyclopedia; a small radar dish on the desk lights three drawers; through an open door, a domed building on a hill labelled frontier.
The specialist, the library, and through the door the frontier for everything else.

agentvcs

Change your harness while it runs and still know what worked.

Every step of a run is stamped with the harness version that produced it, so a patch applied mid-run can be gated, attributed with blame, rolled back or resumed from a checkpoint. Diverged harnesses merge dimension by dimension; each real conflict goes, with its evidence, to Claude Code, agentvcs’s only LLM, and the resolution is checked before it is recorded.

Repository → The July 2026 write-up (earlier Python version)

Measured, and not

A Rust core, CLI, MCP server, Python SDK and a conformance suite.

ran  On Qwen2.5-1.5B on llama.cpp, a patch applied mid-run passed a paired gate with a recall gain of +0.444, and blame attributed +0.442 to exactly one patch — the one applied (RUN_REAL.md).

not built  No PyPI package; no real Claude Code merge session recorded; no bisect --exec, remotes or sync.

A brass machine on a bench with one freshly replaced gear; a paper ledger unrolls from it, each line sealed in wax, blue before the change and orange after, and a red ribbon ties the new gear to the first orange seal.
Every step stamped with the harness version that produced it.

gene-evidence

Auditable evidence for gene predictions the reference annotation does not have.

Some predicted genes are absent from the species’ reference annotation: missed genes, or artefacts. For each, gene-evidence analyze builds an evidence graph — locus context, retrocopy signature, Swiss-Prot homology, Pfam domains, ORF sanity — each node one tool output with its version and parameters. The report cites nodes in every sentence and check-report fails if one does not. No language model; the rules are in EVIDENCE.md.

Repository → The rat run

The first run, on rat

A development run, not a test of ranking quality.

947 candidates; 578 of 947 (61 %) sit on loci the reference calls pseudogenes, which is why locus context is read before homology. Two runs gave byte-identical graph and report, and every sentence passed the citation check; 35 unit tests. The top 100 are all tier 1 and repeat-rich families fill them, as written down before the run: 32 hit the Smok kinase family, 24 T-cell receptor or immunoglobulin variable segments, 2 retroviral Pol polyproteins (rank 1 among them).

known limitation  On plant genomes, against long-read transcript evidence, the tiers ranked supported candidates worse than the predictor’s own confidence score. Use the report to explain a candidate, not to choose which to validate first (README).

Bar chart of the 947 candidates of the rat development run by what the reference annotates at the locus: pseudogene 578, nothing 285, protein-coding gene on the other strand or in an intron 45, non-coding gene 39.
The rat run by what the reference annotates at each candidate’s locus.
One candidate's section of a gene-evidence report: evidence for, evidence against, context and a suggested validation experiment, each line ending in bracketed graph-node citations.
Rank 1 of the rat run: every line ends in the graph nodes it cites.

Twenty-five more are archived and read-only, listed in the archive. agentvcs is among the projects that passed the ANFAIA Research Lab.

Questions, reproductions that disagree with a number here, and proposals are welcome as issues: lora-kernel · agentvcs · gene-evidence · everything else under github.com/EvolvingAgentsLabs.