gene-evidence
Auditable evidence for gene predictions that the reference annotation does not have.
prototype · stage 1 Carbon-A, a public set of
gene predictions, covers thousands of GenBank assemblies, and some of its predicted genes are absent
from the species’ reference annotation (NCBI RefSeq). Each is either a gene the reference missed
or a prediction artefact. For each one, gene-evidence analyze collects what the reference
annotates at the locus, a retrocopy signature, Swiss-Prot homology, Pfam domains, ORF sanity and
optional RNA evidence into an evidence graph: every node is one tool output with its tool,
version, database hash, parameters and raw values. The report cites graph nodes in every sentence,
and check-report fails if one does not. There is no language model and no learned
model in this stage; the rules are fixed and written down in
EVIDENCE.md.
Repository →
Quickstart: see the README
The rat run
The evidence against comes first
Why locus context is read before homology.
A processed pseudogene is a copy of a spliced mRNA inserted back into the genome: no
introns, nearly identical to its parent, often with a reading frame intact enough for a predictor to
call it a gene. On those loci a near-identical homolog and the parent’s domains point the wrong
way. So a pseudogene at the locus or a retrocopy signature puts a candidate last, whatever the
evidence for it says. On the rat development run,
578 of 947 candidates (61 %) sit on loci the reference already calls pseudogenes.
What the first run showed
A development run on rat, not a test of ranking quality.
947 candidates analysed. Two independent runs gave byte-identical graph and
report (determinism.txt),
and every sentence passed the citation check
(check.txt); 35 unit tests on
synthetic fixtures, no network. The top 100 are all tier 1, and reading them found the failure mode
written down before the run: repeat-rich families fill them — 32 hit the Smok kinase family,
24 T-cell receptor or immunoglobulin variable segments, and 2 retroviral Pol polyproteins,
rank 1
among them. The tier rule was not changed after seeing it.
checked nightly Each of these numbers is read from
the run’s committed files by the same check as lora-kernel’s.
Known limitation
Use the report to explain a candidate, not to choose which to validate first.
Checked against independent long-read transcript evidence on plant genomes, the fixed
tier ordering ranked supported candidates better than chance but worse than the gene
predictor’s own confidence score, and demoting candidates on their evidence against did not
improve the top of that ordering. The tiers are a reading aid, not a prediction.
Nor does the project claim that tier 1 is enriched for real genes, that a candidate
ranked last is not a gene — retrocopies can be expressed, and pseudogene calls can be wrong
— or that “no Swiss-Prot hit” means “novel”. Its
README says so first.