MESSAI · Onboarding  /  Atlas

The onboarding atlas: five parts, one page.

Everything the overview routes to, in full. Part I is the atlas itself — the science, access, the pipeline, categorisation, tooling, priors, the held-out audit, the roadmap, the gaps and a glossary, with the interactive platform model at 07. Part II is the runbook for a corpus pass. Part III is the platform map in full. Part IV follows a prediction into operator economics and the public claims. Part V is the P&ID symbol library. Every count names the database it was measured on and when; re-measure before quoting.

5 parts 46 sections ↑ Onboarding overview

Part I of V · Onboarding atlas · 17 sections · ~35 min

MESSAI · Onboarding · 2026-09-07

The science, three databases, one bucket, six pipeline stages, and the week that changed the diagnosis.

The Onboarding Atlas, keeping the artifact’s numbering. Sections 01–05 are the science with no repo access needed; 06 is where the data lives and 11 the acquisition tooling; 13–14 are what the models can actually do and the held-out audit that measures it; 15–18 are the July 2026 week, the roadmap, the ranked gaps and the glossary. 07 and 12 are pointers: the platform map is Part III and the corpus runbook is Part II. The corpus reference that used to sit here as 08–10 — the pipeline, categorisation and ontology, growing the corpus — now lives in Part II as 08–10, next to the runbook it belongs to. Every number carries its date; re-measure before quoting.

3 + 1
Environments local · staging · prod, plus one shared R2 bucket
gate on project REF, never hostname
6
Pipeline stages acquire → screen → extract → sync → harmonize → promote
harmonization is the binding constraint
3
Managed agents screener · extractor · modeler
hold no credential; presigned URLs only
May 2026
last discovery run acquisition stalled; 2025–26 is 23% PDF-backed
see Part II · 10 Growing the corpus
4.1%
of the R2 corpus usable for modeling before the July work
316 of 7,782 · CORPUS-DIAGNOSIS.md
00Before you touch anything · what to load, and what never to do

This is the reference half of onboarding. If you are looking for where to start — the four routes, the six pipeline stages, what the platform can and cannot support, and your first task — that is the onboarding overview, and it takes about six minutes. This page is what you read once you know which part you need.

Part I is numbered on its own 00–18 scale, so a cross-reference like “see 08 Pipeline” means section 08 of this part. Every number here carries the database it was measured on and the date; re-measure before quoting. The other four parts, with their section counts and read times, are listed at the foot of this page.

For agents, load before touching anything
  1. The memory index at ~/.claude/projects/-Users-samfrons-repos-messai-ai/memory/MEMORY.md, then four entries in particular: feedback-corpus-lives-on-staging-not-local, feedback-grade-experiment-records-not-paper-completeness, feedback-dangling-node-modules-symlinks-false-green-typecheck, project-acquisition-path-status-2026-09-07.
  2. The six skills that carry the procedures: mes-paper-retrieval, mes-parameter-extraction, mes-data-harmonization, mes-db-migration-sync, mes-experiment-modeling, mes-knowledge-graph.
  3. The three gates that tell you whether a run actually worked: scripts/quality/coupling-completeness.ts, scripts/extraction/audit-disk-vs-db.ts, and pnpm verify:science.
Never
  • Never write to a remote database except through resolveDbTarget(), gated on the Supabase project ref. Staging and production share a pooler hostname, so a hostname check proves nothing.
  • Never start a paid extraction without a three-paper smoke run first.
  • Archive, never delete. Move rows and files; do not drop them.
  • Never quote a corpus number that is older than its last re-measurement. Re-measure the target database and carry the date.
  • Never trust a type check you have not watched fail. Forty-six dangling symlinks in the root node_modules make the type gate report false green.

Repo documents, in reading order

docs/onboarding/README.md is the ordered index: QUICKSTART, infra-access-r2-staging-prod, paper-pipeline, categorization-and-ontology, acquiring-recent-papers, streamlined-corpus-run, extractor-contract-and-gaps, the 2026-07-19–24 handoff, and CORPUS-DIAGNOSIS. Then docs/extraction/extractor-contract-and-gaps.md, docs/AGENT_HANDOFF.md, docs/PLATFORM_STATE.md, docs/TECH_STACK.md, docs/QUICKSTART.md, docs/multi-zone-deployment.md, CLAUDE.md.

01Microbes as self-replicating electrocatalysts

Read this first if the science is new to you. A microbial electrochemical system (MES), also called a bioelectrochemical system, uses the electrochemical activity of living microorganisms to interconvert chemical and electrical energy, or to drive synthesis. Every device in the taxonomy is a variation on one cell; what changes is the reaction you drive and the product you harvest. The full treatment is on /learn/science; the two diagrams below are the pair everything later in this page assumes.

The reference device

In the canonical case — a microbial fuel cell — electroactive bacteria colonise an anode and oxidise an organic substrate such as acetate or the organic load of wastewater. Oxidation releases electrons, which the bacteria deposit onto the electrode instead of onto a soluble acceptor. Those electrons travel through an external circuit to a cathode, where they reduce a terminal acceptor, often oxygen. The electron flow is a usable current; the proton flow that balances it closes the circuit through the electrolyte.

MICROBIAL FUEL CELL — THE REFERENCE DEVICE electrolyte / separator anode (−) biofilm substrate → CO₂ + H⁺ + e⁻ electroactive bacteria oxidise the organic load and deposit electrons onto the electrode cathode (+) O₂ + 4H⁺ + 4e⁻ → 2H₂O terminal acceptor is reduced H⁺ migrate to balance charge LOAD e⁻ → useful work
The biofilm on the anode oxidises substrate and deposits electrons on the electrode; they do useful work through the external load and return to the cathode, while protons migrate through the electrolyte to close the circuit. Redrawn from /learn/science.
The core mechanism
Oxidise, transfer, reduce
Bacteria at the anode oxidise substrate and release electrons. Electrons flow through the external circuit to the cathode, generating power while reducing an acceptor. Protons migrate through the electrolyte or a membrane to balance charge.
Why microbes
Self-replicating catalysts
Microbes are self-replicating, self-repairing catalysts that work at ambient temperature and pressure, and can process complex, dilute waste streams that precious-metal catalysts handle poorly or not at all.
Why the method looks like this
Two facts govern everything
The literature reports the same quantity in mutually incompatible ways — different normalisation bases, peak versus steady-state, undeclared reference electrodes — so raw numbers cannot simply be averaged. And the headline metrics are heavy-tailed across orders of magnitude, so a single number is almost always the wrong answer.
The disciplined response
Five verbs
Classify precisely, extract with provenance, harmonise ruthlessly, model on the log scale, and report calibrated distributions rather than point values. Sections 02 to 05 are the first two verbs; Part II and Part III are the third; Part IV is the last two.

The public methodology page, section by section

Everything on messai.io/learn/science is carried here, so either page answers the question and neither is the only copy. That page is the outward-facing account and is depth-layered behind “Read the full method” panels; this one adds what was measured internally, and says where the two disagree. Use whichever you are already in.

On messai.io/learn/scienceWhat it coversHere
OverviewWhat an MES is, the reference device, why microbes, the two facts that shape the method01 — this section
System taxonomy17 types, 5 families, subtypes, combinations, study focus, domains02, measured in 09
ThermodynamicsGibbs, Nernst, the redox ladder, the loss cascade, the three efficiencies03
Electrode kineticsButler–Volmer, Tafel, polarisation, internal resistance, transport, EIS03
Biofilm & EETFour transfer routes, model organisms, Monod, biofilm architecture04
Performance metricsWhat is measured, in what units, and the normalisation trap05
Data pipelineDiscover → resolve → extract → harmonise → sync, the quality gate, verification08; the runbook is 10–12
Bayesian priorsHierarchical, log-scale, partial pooling, the prior record13
From priors to predictionsRouting, split-conformal calibration, the output contract13, audited in 14
What makes us differentFour claims, the proof numbers, open data, known limitationsbelow, and 17
GlossaryKey terms18, extended with platform terms

Sibling pages that page links out to, not duplicated here: System types, Sustainability, History of MES, the proof dashboard (live database counts and training artifacts per request) and the whitepaper — the strategic why, where the science page and this Atlas are the technical how.

The four public claims, and where each one is tested here

ClaimWhat the public page saysWhere it is audited in this Atlas
honesty
Honest data-quality signals
Every prediction ships a data_status field. Missing artifacts surface as awaiting_artifact, never a silent zero. Power-density CoV ≈ 1,285%, so point estimates always travel with context.The output contract and its two refusal states, 13. The CoV and why no bare mean is quoted, 05.
coverage
Full taxonomy coverage
17 primary types + 12 combinations + multi-axis subtypes. No silent coercion of MES / MDC / MMRC papers to MFC-shaped output — every system routes through its physics family.The taxonomy itself, 02. What the corpus actually carries, 09: primarySystemType is NULL on 10,286 visible papers, and five vocabularies describe the same papers.
traceability
Provenance on every value
Every extracted number is linked to its source paper. About half also carry the verbatim snippet, with an extractor version and confidence — where that is missing they say so rather than implying a quote. Roughly a quarter of papers have no DOI.Provenance fields and the disk-vs-DB audit, 08. What “done” means for a row, 08. The 2,075 no-DOI papers, 17.
rigour
Statistics that respect the data
Hierarchical, log-scale priors that pool strength across classes — not arithmetic means of heavy-tailed distributions. The mathematically wrong v0 was retired, not shipped.The priors and partial pooling, 13. Whether they beat a class median on paper-disjoint data, 14 — the answer is no, for every model tried.

The proof numbers, as published: 23,568+ curated papers, 195,846 extracted value rows, 2,812 parameter-graph edges, and 97.98% out-of-sample interval coverage on MFC + MEC, with an honest-gaps section naming what is thin or unvalidated. Two of those need the footnote this Atlas adds: the coverage figure was computed on a split that 14 shows is separable at AUC 0.77–0.92, and the corpus counts are dated snapshots that differ slightly from the staging measurements below. Open data: the curated packages are public — canonical microbes, electrode materials with DFT-derived properties, the parameter ontology, and a catalog of external datasets — each shipping a SCIENTIFIC_INTEGRITY.md of known caveats, with large blobs mirrored to Hugging Face. Details in 07.

02System taxonomy · 14 devices, 3 meta classes, 5 physics families

also public  Live version: messai.io/learn/science → System taxonomy.

The canonical taxonomy (version 2) is 17 primary system types — 14 real devices plus 3 meta classes — with 12 combination types, multi-axis subtypes, an 11-value study-focus axis and 8 application domains, all routed for prediction through 5 physics families. It is the authoritative classifier that the extraction pipeline, the ML predictors and the UI all share; the enum lives in apps/web/src/lib/taxonomy/system-types.ts. Section 09 measures how much of the corpus actually carries each value.

Why the family, not the device, is the unit of prediction

Restricting the world to MFC / MEC / MDC — as most tools do — silently coerces desalination, metal-recovery, synthesis and sensing systems into fuel-cell-shaped output. The right abstraction is the physics family: the group whose transport and reaction physics actually govern the outcome. Predictions route to a family; subtypes become features, combinations become coupling rules. The rationale is statistical: branching on all 17 types would train rare classes on a handful of papers each and return noise.

Anodic oxidation
MFC · MESNORK · MRB
Organic substrate oxidation at a bio-anode. The energy-harvesting workhorse — power generation, and the electron supply for downstream reduction.
Cathodic reduction
MEC · MES · MEFS
A reaction driven at the cathode: hydrogen evolution, CO₂ fixation to organics, or Fenton-reagent generation. Often needs a small applied voltage.
Ion transport
MDC · MREC · MCDI · MNRC
Selective ion migration across ion-exchange membranes — desalination, salinity-gradient power, capacitive deionization, nutrient migration.
Selective reduction
MMRC · MERC
Targeted reduction of a specific species — metal recovery, and the electro-driven degradation or immobilisation of contaminants.
Sensor / photo
MBES · MSC
Current as a signal (biosensing) or light as a driver (biophotovoltaics), where the output tracks an analyte or a photosynthetic input rather than bulk power.

The full device roster

TypeFull nameFamilyWhat it does
MFCMicrobial Fuel CellanodicOxidises organics, delivers current spontaneously
MECMicrobial Electrolysis CellcathodicAdds a small voltage to make H₂ at the cathode
MESMicrobial ElectrosynthesiscathodicReduces CO₂ at the cathode into organics
MDCMicrobial Desalination CellionField drives ions out of a middle chamber
MSCMicrobial Solar Cell / Biophotovoltaicsensor/photoCouples phototrophy to the electrodes
MEFSMicrobial Electro-Fenton SystemcathodicGenerates H₂O₂ / •OH to degrade pollutants
MESNORKMicrobial Electrochemical SnorkelanodicShort-circuited electrode accelerates oxidation
MERCMicrobial Electroremediating CellselectiveDegrades or immobilises a target contaminant
MNRCMicrobial Nutrient Recovery CellionRecovers N / P (ammonium, struvite)
MMRCMicrobial Metal Recovery CellselectiveReduces dissolved metals at the cathode
MRBMicrobial Rechargeable BatteryanodicBioelectrochemical energy storage
MBESBioelectrochemical Sensorsensor/photoCurrent tracks an analyte (BOD, toxicity, …)
MRECMicrobial Reverse Electrodialysis CellionHarvests salinity-gradient energy
MCDIMicrobial Capacitive Deionization CellionCapacitive ion storage for deionization

Three meta classes complete the 17: REVIEW (review / meta-analysis), FUNDAMENTAL (mechanistic electron-transfer studies) and OTHER (genuinely novel architectures). They are tracked so the predictor never mistakes a review’s cited numbers for fresh measurements — the same reason 26.3% of values sourced from references is a headline finding in 15 July 2026 agents. Abiotic fuel cells (PEM, SOFC, PAFC) are a separate taxonomy: the predictor stack has tuned physics for them, but they are not bioelectrochemical and are kept out of the BES classifier so a proton-exchange-membrane paper never contaminates the microbial priors.

Careful: MES means two different things
  • On this page and on the public science page, MES is the umbrella — microbial electrochemical system, the whole field.
  • In the extraction schema and in primarySystemType, MES narrowly means Microbial Electrosynthesis, cathodic CO₂ reduction: one of the 14 devices. audit-non-bes-candidates.py uses the second sense throughout.
  • Physics families are for calibration strata and out-of-distribution detection, not as predictors in their own right, and there is still no TypeScript constant for the family mapping — it lives in prose in docs/system-class-aware-predictor-architecture.md. See the ontology rules in 09.

Subtypes: orthogonal axes, not one flat list

Subtypes are stored as an object of orthogonal axes because a token like acetate means something different for a synthesis product than for a substrate. Each primary type populates the axes that make sense for it.

AxisUsed byExample values
genericMFC, MEC, MDC, MSC…single_chamber_air_cathode · dual_chamber · tubular · stacked · sediment · plant · constructed_wetland · membraneless
by_productMESacetate · methane · ethanol · butyrate · medium_chain_fatty_acids
by_carbon_sourceMESpure_co2 · flue_gas · bicarbonate · direct_air_capture
by_metalMMRCcopper · cobalt · chromium · silver · gold · rare_earth
by_contaminantMERCheavy_metal · perchlorate · chlorinated_solvent · nitrate · sulfate · petroleum_hydrocarbon
by_analyteMBESbod · toxicity · pathogen · dissolved_oxygen · glucose · lactate · ph · xenobiotic
by_powerMMRC, MBES, MCDIself_powered · externally_powered / supplemental_voltage
by_mechanismFUNDAMENTALdirect_ET_cytochrome · direct_ET_nanowire · mediated_endogenous · mediated_exogenous · interspecies_ET

Combinations: coupling rules, not new predictors

Real reactors are often two systems at once. Rather than train a separate model per combination — which would be starved of data — each is treated as a coupling rule over the base families. The best-studied combinations carry their own variants: CW_MFC splits into horizontal_subsurface_flow, vertical_flow and floating.

MFC_MEC MFC–MEC coupledMFC_AD anaerobic digestionMFC_MBR membrane bioreactorCW_MFC constructed wetlandOSMFC osmotic / forward osmosisSOLAR_MEC solar-assisted MECMFC_ECOG electrocoagulationMFC_ALGAE algae bioreactorMEC_DF dark fermentationMFC_MDC MFC–MDC integratedMES_AD electrosynthesis–ADMFC_EF_CW electro-Fenton + wetland

Two orthogonal axes: study focus and application domain

The primary type says what reactor a paper is about. Two further axes say what kind of study it is and what it is for, so a query can ask “which MFC papers are about new anode materials for wastewater?” without full-text search.

study focus11 values

materials_science · process_optimization · scale_up · biofilm_characterization · community_composition · electrode_engineering · kinetics · modeling · review_synthesis · characterization_methods · general.

application domain8 values, multi-tag

power_generation · wastewater_treatment · resource_recovery · electrosynthesis · bioremediation · desalination · biosensing · general. A paper can carry several. This axis is the most specific stratum key in prior lookup, so a missing domain silently falls back to a coarser prior.

03Thermodynamics & kinetics · What a cell can do, and what it actually does

also public  Live version: messai.io/learn/science → Thermodynamics + Electrode kinetics.

What a cell can do is bounded by thermodynamics; what it actually does is set by kinetics and losses. Thermodynamics gives the ceiling, and the ceiling is the gap between the two half-reaction potentials.

Cell potential and free energy E_cell = E_cathode − E_anode
ΔG = −n · F · E_cell
E = E° − (RT / nF) · ln(Q)   (Nernst — potential at real concentrations) n = electrons transferred · F = Faraday constant, 96 485 C/mol · R = gas constant · T = temperature · Q = reaction quotient.

The redox ladder

Each half-reaction has a standard potential on the standard-hydrogen-electrode (SHE) scale, shifted for real concentrations by Nernst. The cathode couple sits high, the anode couple sits low, and the vertical gap between them is the theoretical cell voltage. Swap the cathode couple and you have a different device.

Half-reactionE° (V vs SHE)Role
O₂ + 4H⁺ + 4e⁻ → 2H₂O+0.82 (pH 7)Air-cathode acceptor (MFC)
Fe(CN)₆³⁻ + e⁻ → Fe(CN)₆⁴⁻+0.36Common lab catholyte
2H⁺ + 2e⁻ → H₂−0.41 (pH 7)MEC cathode, hydrogen evolution
acetate / CO₂ couple≈ −0.28Acetate oxidation at the anode
CO₂ → acetate (8e⁻)≈ −0.28 to −0.5Electrosynthesis target (MES) — below the anode, which is why it needs an energy input

An air cathode (O₂/H₂O, +0.82 V) against an acetate-oxidising anode (≈ −0.28 V) gives E°cell ≈ 1.1 V. Real cells deliver 0.3–0.8 V after losses.

Where the 1.1 V goes — the loss cascade

WHERE THE 1.1 V GOES
The 1.1 V ceiling is spent, not delivered. Activation overpotential (the kinetic price of turning the reactions on), ohmic loss (IR drop through electrolyte, membrane and contacts), and concentration overpotential (reactant depletion at the surface) each take a bite, leaving 0.3–0.8 V under load. Illustrative split, redrawn from messai.io/learn/science.
η_act
Activation
The kinetic price of driving each electrode reaction at a useful rate. Dominates at low current.
η_ohmic = IR
Ohmic
Resistance of the electrolyte, membrane, electrodes and connections. Grows linearly with current.
η_conc
Concentration
Depletion of reactant, or build-up of product, at the electrode surface. Dominates at high current and sets the ceiling on current.

Kinetics: Butler–Volmer and Tafel

Kinetics turns a thermodynamic possibility into a rate. The governing relationship between overpotential and current density is the Butler–Volmer equation, with the biofilm acting as the anode’s catalyst layer.

Butler–Volmer, and its Tafel limit j = j₀ · [ exp(α_a · F · η / RT) − exp(−α_c · F · η / RT) ]
Tafel (large |η|):  η = a + b · log₁₀(j)   b = 2.303 · RT / (α · n · F) j = current density · j₀ = exchange current density · η = overpotential · α_a / α_c = charge-transfer coefficients · b = Tafel slope in mV/decade.

The exchange current density j₀ captures how intrinsically fast the electrode reaction is at equilibrium; the transfer coefficient α sets how symmetric the anodic and cathodic branches are; and the Tafel slope — the straight-line region on a log-current plot — is what most papers actually report, from which α can be back-fitted. A good bio-anode has a high j₀: the reaction turns on for less overpotential, so more voltage reaches the load. Every millivolt of η is voltage you do not deliver.

The polarisation curve, and the four levers

Sweeping the load traces voltage against current. The three regions map one-to-one onto the three losses — activation (the steep initial drop), ohmic (the linear middle), mass transport (the final collapse) — and the power curve peaks where the external load matches the internal resistance. That peak, not the open-circuit voltage, is what a design is optimised toward, which is why extracting a single quoted “maximum power density” without its derivation method is lossy (see the derivation_method field in 12 Corpus run).

Reactant must diffuse across a thin diffusion layer to reach the electrode. Push the current higher and the surface concentration falls; when it reaches zero the current is diffusion-limited and cannot rise further, no matter the voltage. That is the collapse at the right of the curve, and the limiting current.

LeverMechanismWhere it stops working
Shrink electrode spacingLower ohmic resistance → higher powerUntil mass transport or short-circuiting intervenes
Raise electrolyte conductivityLower ohmic resistanceWithin osmotic and biological limits
Grow a better biofilmRaise j₀, lower activation overpotential at the anodeBiofilm thickness has a sweet spot — see 04
Improve the cathode catalystThe cathode is frequently the limiting electrode in air-cathode MFCsCost and durability

Internal resistance is often the single largest lever on power: the sum of ohmic resistance (electrolyte, membrane, electrode spacing) and the charge-transfer resistances of both electrodes. By the maximum-power-transfer theorem, power delivered to the load is maximised when the external resistance matches the internal resistance — which is why it is one of the parameters the knowledge graph links to power with a causal edge.

Impedance (EIS). Electrochemical impedance spectroscopy applies a small AC perturbation across a frequency range and fits the Nyquist spectrum to an equivalent circuit — typically a Randles circuit — separating losses a single DC polarisation curve blends together: Rs solution resistance (the high-frequency intercept, the ohmic term), Rct charge-transfer resistance (the semicircle diameter, inversely related to j₀), Cdl double-layer capacitance, and the Warburg low-frequency diffusion tail. EIS spectra are Tier C on the extraction roadmap (12 Corpus run) — the schema does not capture them today.

Why the platform does not compute these from first principles
  • These levers interact non-linearly and trade off against each other — shrinking the electrode gap that lowers ohmic loss can also starve mass transport.
  • The losses are coupled and depend on variables papers report unevenly, so a fitted analytic model generalises poorly across the corpus.
  • MESSAI therefore learns their net effect empirically, conditioned on system class. The served physics layer is a 0-D mean function, not a spatially-resolved multiphysics model — stated plainly so it is not oversold. 14 Held-out benchmark measures what that buys.
04Biofilm & EET · How bacteria put electrons on a rock

also public  Live version: messai.io/learn/science → Biofilm & EET.

The defining trick of MES is extracellular electron transfer (EET): electroactive bacteria route respiratory electrons onto a solid electrode instead of a soluble acceptor. Understanding EET is understanding the anode.

Route 1
Direct transfer
Outer-membrane multiheme cytochromes make near-contact with the electrode, within about 1–2 nm, and hand off electrons directly. Fast, but limited to cells touching the surface.
Route 2
Conductive nanowires (pili)
Conductive protein filaments extend microns from the cell, wiring distant cells into the electrode and letting thick biofilms stay electrochemically active.
Route 3
Mediated transfer
Soluble redox shuttles — flavins secreted by Shewanella, or added mediators — ferry electrons across the gap. Effective, but shuttles can wash out of flow-through reactors.
Route 4
Biofilm-matrix conduction & DIET
Conduction through the matrix itself, so current scales with active biomass. Direct interspecies electron transfer lets one species hand electrons to another through pili or minerals — the basis of syntrophic consortia and electromethanogenesis.

The four routes are usually mixed in one biofilm, which is why by_mechanism (direct_ET_cytochrome, direct_ET_nanowire, mediated_endogenous, mediated_exogenous, interspecies_ET) is a subtype axis on FUNDAMENTAL papers rather than a device class.

The model organisms

OrganismSignatureRole
Geobacter sulfurreducensConductive nanowires + cytochromesGold-standard anode electrogen; dense conductive biofilms
Shewanella oneidensisFlavin-mediated + directFacultative model electrogen; versatile respiration
Sporomusa / Clostridium (acetogens)Wood–Ljungdahl CO₂ fixationElectrosynthesis of acetate at the cathode
Methanogens (e.g. Methanosarcina)DIET acceptorElectromethanogenesis — CO₂ → CH₄
Cyanobacteria (Synechocystis)Oxygenic photosynthesisPhotosynthetic anodes and biocathodes (MSC)

The mess-microbes package (07 Platform map) is the curated catalog behind this: 28 microbes, plus 181 MicrobeKineticConstant rows.

Growth kinetics — Monod

Saturating substrate kinetics μ = μ_max · S / (K_s + S)   (specific growth rate)
q = q_max · S / (K_s + S)   (specific substrate-utilisation rate) μ_max = maximum growth rate · K_s = half-saturation constant · S = substrate concentration.

Uptake is proportional to concentration when substrate is scarce and flattens to a maximum when it is abundant. K_s is substrate-specific, which is why substrate identity is kept as a canonical feature rather than pooling glucose and acetate kinetics together — and why substrate being the worst-linked input in the corpus (0.7% coupled, see 12) is a modelling problem, not a cosmetic one. On 2026-09-08 the identity μ_max = ln 2 / doubling time validated three values the paid LLM verifier had rejected: a free invariant beating a paid check.

Two facts about biofilms that shape the data model
  • Biofilm thickness has a sweet spot. Too thin and there is not enough catalyst; too thick and substrate cannot diffuse to the inner layers while protons cannot escape, so the interior goes acidic and inactive. The knowledge graph encodes this as a causal edge (biofilm thickness → power density) rather than a monotonic rule.
  • Community composition matters as much as any single species. Real anodes are mixed communities where fermenters, syntrophs and electroactive bacteria form a food web. Enrichment and inoculum are among the least reproducible variables in the literature — which is why the extractor records the community-analysis method (16S versus metagenomics) as a confidence signal, not just the organism name. Inoculum is reported on 0% of benchmark rows (14).
05Performance metrics · What is measured, and the normalisation trap

also public  Live version: messai.io/learn/science → Performance metrics.

MES performance is reported through a compact set of metrics, but with inconsistent normalisation across the literature. Knowing exactly what a number is normalised to is the difference between a fair comparison and a meaningless one.

MetricTypical unitWhat it tells you
Power density (areal)mW/m²Power per electrode area — the most-cited MFC metric
Power density (volumetric)W/m³Power per reactor volume — the scale-up metric
Current densitymA/cm² · A/m³Electron flux per area or volume
Coulombic efficiency%Fraction of substrate electrons recovered as current
COD / BOD removal%Organic load removed — treatment performance
Cell voltage / OCVVWorking voltage under load / open-circuit voltage
Internal resistanceΩ · Ω·cm²From the polarisation slope or an EIS fit
Energy efficiency%Useful energy out / energy in
H₂ production rate / yield (MEC)m³/m³/d · mol/molRate and recovery for electrolysis cells
Product titre / rate (MES)g/L · g/L/dElectrosynthesis output
Salt removal (MDC)% · mg/LDesalination performance
Sensitivity (MBES)signal / analyteBiosensor response
The normalisation trap, and why it is the platform’s central discipline
  • Power density is the worst offender. The same reactor can be reported at wildly different numbers depending on whether power is divided by anode area, cathode area, membrane area, or reactor volume.
  • Areal and volumetric densities are therefore routed to separate canonical kinds so they never pool into one prior. Mixing W/m² and W/m³ is a category error that contaminates the posterior roughly 1000× — the first harmonization hard rule in 08 Pipeline.
  • Every row records normalization_basis (anode vs cathode vs membrane vs volume) and value_kind / phase (peak vs steady-state), because those distinctions are worth 1.5–3× and up to 10–1000×. Peak-versus-steady-state alone produces 2–3× errors.
  • Reference electrodes are the quiet one. A 30–110 mV offset between reference types makes raw voltages incomparable; everything is normalised to vs-SHE. This is the single genuinely missing field in the v2 extraction schema — voltage_reporting_convention, whether a value is vs Ag/AgCl, vs SHE, or a whole-cell voltage. See 12.
  • The three efficiencies are not interchangeable. Coulombic efficiency is the fraction of available substrate electrons recovered as current, bounded 0–100%. Cathodic capture is the fraction of arriving electrons that end up in the desired product (H₂, acetate) rather than side reactions. Energy efficiency combines voltage and coulombic losses into the bottom line.

Why a bare mean is never quoted

Power density is right-skewed over orders of magnitude, so the arithmetic mean is dragged far past the median by a handful of high performers. The median describes a typical system; the mean describes almost none. Across the corpus the coefficient of variation is ≈ 1,285%. This is the concrete reason the priors are fitted on the log scale (13), and the reason the served interval on power density is four decades wide (14).

Coulombic efficiency is the diagnostic one

It is bounded and interpretable: a high current with low coulombic efficiency means electrons are leaking to alternative acceptors — oxygen crossover, methanogenesis, other respiration — rather than reaching the electrode. It is often more informative about the biology than power density is. On the held-out benchmark it is also the target with the widest between-paper spread (2.04 logit).

Always read the interval, the system class and the normalisation before comparing two numbers. Section 09 is how a raw name becomes a canonical slug and an SI value; section 14 is what happens when you try to predict these metrics across papers.

06Where the data lives and who may write to it

Staging and prod share the same Supabase pooler host in eu-central-1, so a hostname check proves nothing. Every remote write in scripts/ passes through resolveDbTarget(), which pins staging to its project ref and demands --expect-ref for prod. Promotion is one direction, additive, schema before data.

CLOUDFLARE R2 · SOURCE OF TRUTH FOR PDF BYTES SINCE 2026-08-14 bucket messai-papers · key pdfs/<aa>/<sha256>.pdf · private, served via signed 302 from /api/papers/[id]/pdf bucket backups · DB dumps, papers-archive/2026-08-11, curve-extractions · env R2_ACCOUNT_ID + R2_ACCESS_KEY_ID/SECRET in .env.local shared by all three DBs · orphan counts differ per environment pnpm papers:hydrate (rclone) r2Key on ResearchPaper presigned URLs to agents Local SUPABASE CLI · :54322 .env.development.local → DATABASE_URL / DIRECT_URL apps/web/.env.development.local for pnpm dev Every stage runs here first. Dual-flag writes: --apply + --i-understand-this-mutates-local pnpm db:up · pnpm db:migrate:local Staging REF · in .env (staging) .env.local → SUPABASE_STAGING_DIRECT_URL session pooler :5432 · never :6543 Backs local dev (since 2026-10-01). pnpm db:sync:staging (types SYNC) insert-only, strips unknown columns org "Frons digital" · eu-central-1 Production REF · in .env (prod) .env.production.local → DIRECT_URL (:5432) DATABASE_URL (:6543) = Prisma runtime only --target=prod needs --expect-ref. Backup → migrate → merge additive → NULL-guarded backfill → validate RLS deny-all · Prisma role bypasses → → PROMOTION local → staging → prod · additive only · migrate schema BEFORE data · run remote --apply with nohup · epd_total must not move on a harmonization run BRANCH TIERS (2026-10-01) feature PRgh pr create --base development developmentintegration · Vercel preview mainproduction · 4 Vercel projects promote by PR
Refs and ports verified against scripts/lib/db-target.ts and .env.local on 2026-09-07. No secret values.

Day-one checklist

#GetFromVerify with
1GitHub access to samfrons/messai-aithe repository ownergit clone
2.env.local + .env.development.localthe repository owner, out of band. Never chat, never git.bash scripts/quality/preflight-corpus-work.sh
3R2 token, bucket-scoped to messai-papersthe account owner · Cloudflare → R2 → Manage API Tokenspnpm tsx scripts/storage/r2-test-creds.ts
4Supabase org "Frons digital" membershipthe org owner invitesopen messai-staging in the dashboard
5Vercel team: messai-ai · messai-lab · messai-api · messai-sitethe team ownerpnpm tsx tools/diff-vercel-env.ts
6rclone + PostgreSQL 17 clientbrew install rclone postgresql@17rclone version
—Prod credentialsNot on day one. Handed over per task; every prod --apply is announced first.

Each rule is a past incident

  1. Never DROP, TRUNCATE, reset or force-overwrite. Archive by moving or soft flags.
  2. Migrate schema before syncing data. Sync intersects columns and silently drops the rest. Lost 22 patent rows, 2026-05-13.
  3. Sync is insert-only. Backfilling existing rows is a separate NULL-guarded UPDATE. The most common "my sync failed".
  4. Never pg_dump over :6543. The 2026-05-12 dump was 9 MB with zero data rows; DIRECT_URL gave 400 MB.
  5. Verify a backup by row count. grep -c '^COPY public\.' ≈ 60–80.
  6. Never put DB URLs in .env. The Prisma CLI reads it directly.
Dead variables that still answer
  • STAGING_DATABASE_URL — old Prisma host, empty since 2025-08-09.
  • PRODUCTION_DATABASE_URL, DEVELOPMENT_DATABASE_URL — retired 2026-08-11. They respond with a stale 3,892-paper snapshot; pnpm backup dumped the wrong DB for months.
  • Root cause: a duplicate accessor for a string an env file already owns. Never add another.
  • dotenv.config() writes into process.env; resolving two targets in one process leaks the first URL into the second. Use dotenv.parse.
07Platform map · One repo, four zones, one Postgres

Part III is the platform map  The measured platform — zones, schema, pipeline, 3D, P&ID, open source, gaps — lives once, in the Platform map tab. This section only opens the interactive model so it is one click from the atlas.

Architecture: one model, eight UML views

Eight UML 2 views over one model of the whole platform. C1 is the altitude everything else hangs from: three kinds of actor, four Vercel zones, the shared packages, the state layer and the external systems. Every box with a + corner opens its own diagram — the four zones, the package graph, the persistence model and the corpus pipeline — and Esc comes back up. Structural counts were measured against this repository on 2026-09-09; corpus and science numbers deliberately stay out of the diagram and live in the dated tiles.

C1System contextComponent / deployment diagram
100%

One repository, four Vercel projects, one Postgres. Everything a browser touches enters through apps/web, which owns the rewrite table; everything that writes to the database goes through apps/api. Click any node to open it.

ActorsPresentation · VercelApplicationStateInference & externalresearcher → proxy · call / request · HTTPSHTTPSweb → proxy · compositionproxy → auth · call / request · sessionsessionconsumer → mirrors · call / request · gitgitconsumer → registries · call / request · installinstallconsumer → hfhub · call / request · downloaddownloadconsumer → api · call / request · REST v1REST v1operator → batch · call / request · by handby handlocaldev → web · call / request · dev:zonesdev:zonesresearcher → chat · call / request · chatchatoperator → routine · call / request · MondayMondayoperator → actions · call / request · dispatchdispatchweb → site · call / request · site · 23site · 23web → lab · call / request · lab · 11lab · 11web → api · call / request · /api/* · 64/api/* · 64lab → api · call / request · fetchfetchweb → libs · dependency «use»lab → libs · dependency «use»site → libs · dependency «use»api → libs · dependency «use»web → pg · read · server readserver readlab → pg · readapi → pg · write · read+writeread+writeapi → r2 · write · signed URLssigned URLsapi → redis · dormant · unhostedunhostedapi → artifacts · read · DB-firstDB-firstapi → gateway · call / request · llm callsllm callsapi → auth · dependency «use»auth → oauth · call / request · OAuthOAuthauth → resend · call / request · emailemailapi → sentry · call / request · errorserrorschat → gateway · call / request · streamstreamchat → pg · read · toolstoolsapi → hf · call / request · embed queryembed querybatch → pg · write · syncsyncbatch → r2 · writebatch → artifacts · write · refitsrefitsbatch → openalex · call / request · acquireacquirebatch → gateway · call / request · extractextractbatch → hf · call / request · embedembedbatch → hfhub · write · datasetsdatasetsactions → mirrors · write · mirror syncmirror syncactions → registries · write · publishpublishbatch → gpscm · call / request · fitfitbatch → mlengine · call / request · parse · fitparse · fitroutine → artifacts · write · npe-healthnpe-healthroutine → gpscm · call / request · refitrefitactions → artifacts · write · fit-priorsfit-priors«actor»Researcherbrowser · messai.io«actor»Operator /Claude agentCLI · scripts · routine«actor»Downstream consumernpm · PyPI · HF · REST«actor»Local dev stacksupabase start · dev:zones«Vercel project · messai-site»apps/siteAstro 5 + Vite · static~8 s build«Vercel project · messai-ai»apps/webNext.js 16 · webpack95 pagesowns the rewrites«middleware»Edge proxyapps/web/src/proxy.tsauth on«Vercel project · messai-lab»apps/labNext.js 16 · webpack17 pagesR3F + three«Vercel project · messai-api»apps/apiNext.js 16 · Turbopack229 routesonly writer«package»@messai/* shared libslibs/ · 17 packages«scripts»Batch pipelinescripts/ · ml-engineno worker host«package»next-authGitHub · Google OAuth«component»AI chat & agents/api/chat · tool registry«scripts»Weekly Claude routineweekly-ml-audit.shweekly«ci»GitHub Actionsworkflow_dispatch only$0 spend«database»Supabase Postgres 17pgvector · 121 Prisma models«object store»Cloudflare R2messai-papers«queue»Upstash Redis + BullMQapps/api/src/lib/jobsdormant«artifact»Computed artifactspriors · calibration · DAG«external»AI GatewayAnthropic · Gemini · Groq«external»OpenAlex · Crossref ·Unpaywallmetadata + open-access PDFs«external»HF Inference Routerbge-large-en-v1.5 · 1024dper query«external»Hugging Face Hubdatasets · 23 records«external»Messai-io mirrorsgithub.com/Messai-io · 9read-only«external»npm · PyPIMESS-* · mess-methods«external»GP-SCMFly.io · messai-gp-scmenv-gated«service»ML engineDocker · local/batch · :8001pnpm ml:dev«external»GitHub · Google OAuthidentity providers«external»Sentry@sentry/nextjs · apps/api«external»Resendtransactional emailprod send untested
actor Vercel zone / route group shared package state batch script external system dormant
C1 as an outline (30 elements · 48 relationships)
  • Researcher «actor» — browser · messai.io → Edge proxy, AI chat & agents
  • Operator / Claude agent «actor» — CLI · scripts · routine → Batch pipeline, Weekly Claude routine, GitHub Actions
  • Downstream consumer «actor» — npm · PyPI · HF · REST → Messai-io mirrors, npm · PyPI, Hugging Face Hub, apps/api
  • Local dev stack «actor» — supabase start · dev:zones → apps/web
  • apps/site «Vercel project · messai-site» — Astro 5 + Vite · static → @messai/* shared libs
  • apps/web «Vercel project · messai-ai» — Next.js 16 · webpack → Edge proxy, apps/site, apps/lab, apps/api, @messai/* shared libs, Supabase Postgres 17
  • Edge proxy «middleware» — apps/web/src/proxy.ts → next-auth
  • apps/lab «Vercel project · messai-lab» — Next.js 16 · webpack → apps/api, @messai/* shared libs, Supabase Postgres 17
  • apps/api «Vercel project · messai-api» — Next.js 16 · Turbopack → @messai/* shared libs, Supabase Postgres 17, Cloudflare R2, Upstash Redis + BullMQ, Computed artifacts, AI Gateway, next-auth, Sentry, HF Inference Router
  • @messai/* shared libs «package» — libs/ · 17 packages
  • Batch pipeline «scripts» — scripts/ · ml-engine → Supabase Postgres 17, Cloudflare R2, Computed artifacts, OpenAlex · Crossref · Unpaywall, AI Gateway, HF Inference Router, Hugging Face Hub, GP-SCM, ML engine
  • next-auth «package» — GitHub · Google OAuth → GitHub · Google OAuth, Resend
  • AI chat & agents «component» — /api/chat · tool registry → AI Gateway, Supabase Postgres 17
  • Weekly Claude routine «scripts» — weekly-ml-audit.sh → Computed artifacts, GP-SCM
  • GitHub Actions «ci» — workflow_dispatch only → Messai-io mirrors, npm · PyPI, Computed artifacts
  • Supabase Postgres 17 «database» — pgvector · 121 Prisma models
  • Cloudflare R2 «object store» — messai-papers
  • Upstash Redis + BullMQ «queue» — apps/api/src/lib/jobs
  • Computed artifacts «artifact» — priors · calibration · DAG
  • AI Gateway «external» — Anthropic · Gemini · Groq
  • OpenAlex · Crossref · Unpaywall «external» — metadata + open-access PDFs
  • HF Inference Router «external» — bge-large-en-v1.5 · 1024d
  • Hugging Face Hub «external» — datasets · 23 records
  • Messai-io mirrors «external» — github.com/Messai-io · 9
  • npm · PyPI «external» — MESS-* · mess-methods
  • GP-SCM «external» — Fly.io · messai-gp-scm
  • ML engine «service» — Docker · local/batch · :8001
  • GitHub · Google OAuth «external» — identity providers
  • Sentry «external» — @sentry/nextjs · apps/api
  • Resend «external» — transactional email

Prerequisites. The memory entries to load, the six skills that carry the procedures, the three gates that tell you a run actually worked, and the rules that are never negotiable are in 00 Before you touch anything. The five-layer state table is in Part II · 00.

What Part III carries

The rest of the platform is drawn and inventoried in Part III, once: the interactive UML model of the whole infrastructure, the site map of every page and the zone that serves it, the Prisma schema and its hub-and-spoke relation map, the data pipeline as it actually runs, the 3D stack and the four registries a new reactor model must join, the P&ID library, the nine open-source packages, and the dependency order between the gaps. This section is the orientation; that is the reference.

08Design system · One component library, three generations of it

Every surface MESSAI ships is supposed to be built from @messai/ui and the Tailwind preset next to it. Most are. This section is the inventory a newcomer builds against — what exists, which generation is live, the rules that are not negotiable, and the measured distance between the rule and the tree. Counted against development on 2026-09-11; re-measure before quoting. Two things have moved since that count: the stock-colour rule became a lint ratchet (2026-09-12), and the palette became warm monochrome (2026-10-05) — both noted below.

48
live components what @messai/ui exports
the v1 tier
33
v3 components behind @messai/ui/v3
reachable from 1 page
20
app-local duplicates of library components
own code, not shims · 8 files left after 2026-09-12
2,823
hard-coded palette classes outside the token set
was 4,489 at v3 branch start
1
Tailwind preset every zone extends it
tailwind-preset.cjs

The three hard rules

Not negotiable — two are lint-enforced
  1. No border-radius, anywhere. Sharp corners are the intended aesthetic. Two exceptions exist and no third one does: rounded-chip (2 px) for the badge / tag / pill family, and the radio button, which uses an inline borderRadius: 9999 rather than widening the Tailwind enum. Enforced by no-restricted-syntax.
  2. Native <select> is banned — use Select from @messai/ui. Enforced by the same rule, with a 21-path allowlist that grandfathers surfaces mid-migration. Each follow-up removes its own entry; when the list empties the rule applies platform-wide on its own. Do not add to it.
  3. Colour comes from the preset, never from Tailwind's stock palette. text-blue-600 and bg-red-50 have no place in a MESSAI surface — the mes-* tokens carry the meaning, so a value swap in the preset moves every page at once. Enforced since 2026-09-12 by messai/no-stock-palette as a ratchet: tools/stock-palette-allowlist.js names the 183 files that already offended, the rule is on everywhere else, and pnpm tsx tools/stock-palette-allowlist.ts refuses to make the list longer. Since 2026-10-05 the stock scales are also remapped in the preset (amber → caution ochre, red/rose → critical rust, green/teal → positive moss, blue/gray → the warm neutral ramp), so a legacy class at least lands on-palette.

What exists, by tier

TierWhereCountWhat it is
Atoms & moleculessrc/components48Button, Input, Select, Checkbox, Radio, Switch, Slider, Label, Link, Code, Tag, Badge, Pill, Divider, Skeleton, Spinner, Eyebrow, FormField, Segmented, Card, AnchoredCard … This is the tier the @messai/ui barrel exports, so it is what an import gives you today.
Organismssrc/components—Modal, Tabs, Table, Toast, Tooltip, Popover, Accordion, Banner, Breadcrumb, Pagination, DropdownMenu, EmptyState, SearchFilters, PaperCard, PaperDetailModal, MetricWithProvenance, TrustBadge. Counted in the 48 above.
Chromesrc/chrome4UniversalHeader (7 files), UniversalFooter (4), AuthMenu, top-bar. The cross-zone furniture — a new zone mounts these rather than drawing its own.
Shellssrc/v3/shells3AppShell (6 files), LabShell (1), MarketingShell (0). A shell is how a page opts into the v3 typography and layout; the marketing one has no consumer yet.
Tokenstailwind-preset.cjs1The single build-time source every zone extends. Colour, spacing, type scale, the rounded-chip exception, motion durations.

The catalogue is drawn, not just listed: /design-system renders live token swatches and component examples, and /dev/v3 is the v3 proving page.

Three generations coexist, and only one is live

This is the thing to understand before adding a component, because the obvious guess is wrong twice.

GenerationImport pathConsumersStatus
v1 · src/components@messai/uievery zonelive The barrel exports this tier and nothing else. Build against it.
v3 · src/v3@messai/ui/v31 pageproven, unadopted 33 components across atoms / molecules / organisms / shells, opt-in by subpath, reachable from /dev/v3 alone. Its tokens did ship — see below — but its components have no production consumer.
app-local · apps/web/src/components/uirelative8 filesduplicate 22 files on 2026-09-11, of which only 2 were thin re-export shims. The 2026-09-12 collapse moved their importers onto @messai/ui; 8 files remain, each for a stated reason — badge and the shadcn-shaped tabs / dropdown-menu would change what a page renders if swapped, and icons.tsx (468 lines) should be promoted into the library rather than deleted.
The token swap landed; the component migration did not
  • Tokens: done, twice. The v3 palette was applied in place — the mes-* class names were kept and their hex values pivoted, so every page picked up the new values with no code change (old→new table: libs/shared/ui/src/v3/tokens/RECONCILIATION.md). On 2026-10-05 the same mechanism moved the platform to warm monochrome: one near-black ink in three shades (#1A1A17 / #44443D / #6A6A60) on cream #F4F1EA and white, rules #DCD9D0. The accent is the ink — buttons, active tabs and links are monochrome — and colour only means state (caution ochre, critical rust, positive moss) or data (the categorical set, charts only). Table and rules: docs/ui-conventions.md “Colour and type tokens”; the static decks and this atlas read the same values through libs/shared/ui/src/styles/deck.css.
  • Typography: one token system. Inter for UI and body, JetBrains Mono for labels and data, Source Serif 4 for titles (a webfont, so every OS draws the same face) — exposed as --messai-font-sans, --messai-font-mono and --messai-font-serif. The old split (DM Mono body, IBM Plex only through a v3 shell, Crimson Text titles) is gone; hard-coded family names are the thing to avoid.
  • The residue is the hard-coded colours. A page that mixes bg-mes-paper with text-blue-600 looks half-migrated, because the token moved and the literal did not. 2,823 such classes remained on 2026-09-11, down from 4,489 when the v3 branch opened; the lint ratchet above means that number can only fall.

Where the site is inconsistent today

ZoneHard-coded paletteNative selectrounded-*Read
apps/web2,345254Carries essentially all of the debt, and all 21 allowlist paths.
apps/lab37441Mostly clean; the 3D surfaces are the exception.
apps/site10030Astro islands, where Radix context is unavailable, take a documented inline disable.
apps/api301No UI bundle; these are incidental strings.
libs/shared/ui27783The library itself breaks its own rules. Fix here first — a primitive that hard-codes a colour re-exports the problem to every consumer.

Adding or changing a component

The order that avoids re-doing it
  1. Look in @messai/ui first, then /design-system. Most of what a surface needs already exists under a name you would not have guessed — Segmented, MetricWithProvenance, AnchoredCard, TrustBadge, Eyebrow.
  2. A new primitive belongs in the library, not the app. That is the rule apps/web/src/components/ui broke twenty times. Shared UI goes in libs/shared/ui; if you need it in a second zone later, it is already there.
  3. Compose from tokens. No hex literals, no stock Tailwind palette classes, no rounded-*. If a token is missing, add it to the preset — one edit that every zone inherits — rather than reaching for #3A6FA0 at the call site.
  4. Tailwind by default; SASS only where Tailwind is ugly. Multi-step keyframes, deep pseudo-element chains, @supports queries, or a component whose class chain would exceed ~6 classes and will not be reused. The module ships beside its component as <Component>.module.scss; shared variables live in libs/shared/ui/src/styles/.
  5. Prove it on a page. /design-system for the library tier, /dev/v3 for v3. A component with no example is a component the next person re-implements.
The three decisions this section cannot make for you
  • Does v3 get adopted or retired? 33 components sitting behind a subpath with one consumer is not a design system, it is a branch. Either migrate surfaces onto the shells or fold the parts worth keeping into the v1 tier.
  • Who empties the colour allowlist? The rule got its linter on 2026-09-12 (a 183-file ratchet). A ratchet only stops the count growing; it shrinks when someone migrates a file and deletes its line. Run pnpm tsx tools/stock-palette-allowlist.ts to list entries that are already clean.
  • Who finishes apps/web/src/components/ui? The 2026-09-12 collapse took it from 22 files to 8. What is left changes rendering when swapped (badge, tabs, dropdown-menu) or belongs in the library (icons.tsx), so each needs a visual before/after, not a find-and-replace.

Conventions in full: docs/ui-conventions.md; the token old→new table in libs/shared/ui/src/v3/tokens/RECONCILIATION.md; the CSS split rule in CLAUDE.md.

11Acquisition tooling · Six generations of scrapers, five stages, ~120 files

Every scraper, downloader, resolver and parser that built the corpus, surveyed 2026-09-07. Two download lineages coexist: lineage A walks a 9-provider chain per paper (download_from_db.py); lineage B runs tiered batch passes built to defeat the ~11k publisher 403s (scripts/acquisition/). Free and official providers only, never Sci-Hub or LibGen. Full per-file table with status and dates: docs/onboarding/acquisition-tooling-inventory.md; visual companion: the Scraper Atlas.

6
eras 2025-07 research agents → 2026-09-03 gap-driven discovery
git first → last commit
5
stages discover · resolve · download · store · parse
all inside pipeline stage 1
~120
files inventoried active · superseded · dead · one-shot
135 paths, all present 2026-09-08
9
providers in download_from_db.py
the mes-paper-retrieval skill still lists 6
EraFromWhatToday
02025-07PubMed / CrossRef / arXiv clients + pdf-parse in the research-agents libdead
12025-08arXiv + PubMed scrapers, seven "final" collection rounds, Nougat OCR, quality-tier PDF storedead
22026-04-25Monorepo consolidation: content-addressed papers/ tree, Snakemake DAG, 9-provider download_from_db.pyactive lineage A
32026-05-09Tiered attack on the 403 wall: paperscraper → curl_cffi → Unpaywall, weekly orchestrator, BioC-PMC XML side-haulactive lineage B
42026-05-18Admin import routes land; BullMQ queue scaffold is mock and never deployedroutes active · queue dead
52026-08-14R2 becomes source of truth, local PDFs a cache; shared DOI / title-match / pdfHash helpersactive
62026-09-03OpenAlex search driven by evidence gaps in effects.json; candidate-only, inserts nothingactive
StageActive todaySuperseded / dead
discoversearch-gaps-openalex.ts · openalex-gap-search.ts + admin/effects/gap-search route · search_openalex_underrepresented.py · check-openalex-mes-coverage.ts · external-search routediscover-papers-via-openalex.ts (superseded) · collect-mes-papers.ts, arxiv/pubmed scrapers, six collection rounds, research-agents external-apis.ts (dead)
resolveresolve.py (4-tier sha256 → DOI) · backfill_identifiers.py · recover_dois_by_title.py · apply_recovered_dois.py · publisher-patterns.ts · scripts/lib/doi.ts + title-match.ts · check-retractions.tsrecover-doi-from-title.ts + 3 siblings (superseded) · repair-paper-* ×7, repair-citations-* (one-shot) · enrich-doi-papers.ts, add-priority-papers* (dead)
downloadA: download_from_db.py · download_playwright.py · export_papers_from_db.sh · download_status.py — B: weekly_pipeline.sh · run-pipeline-for-source.sh · build_local_manifest.py · download_paperscraper.py · download_curl_cffi.py · download_unpaywall.py · xml_to_pmc_pdf.py · download_patents.ts · hydrate-pdfs-from-r2.tsprepare-wastewater-modeling-cohort.py (one-shot) · download-pdfs-smart-storage.ts (dead)
storeSnakefile · discover.py · import_pdfs.py · build_manifest.py · rollup_*.ts · sync-papers-manifest.ts · promote_to_canonical.py · ingest_doi_list.ts · audit_duplicate_papers.py + merge-duplicate-papers.ts · recover-r2-orphans.ts · backfill-r2-source.ts · create-rows-for-orphan-pdfs.ts · pdf-hash-backfill.ts · papers/[id]/pdf routemigrate-local-to-r2.ts + backfill-r2-keys.ts (one-shot, but the only R2 upload path) · insert-openalex-niche-papers.ts, ingest-cheng-logan-2007.ts (one-shot) · pdf-storage-manager.ts, deduplication-service.ts, ~20 import-* scripts, archive/2026-07-prisma-era-scripts (dead)
parsepmc-xml-to-text.ts + run-xml-batch.ts · extract_text.py (marker) · marker_pipeline.py · pdf-triage.ts · extract_tables/figures/charts.py · extract_metadata.py (GROBID) · embed.py · simple_value_extractor.ts · inventory-review-papers.tsorchestrator.ts + retrieval.ts (v1, superseded) · Nougat stack ×5 + run_extraction.sh (superseded) · scientific_paper_extractor.py + 3 (superseded) · abstract-extractor.ts (superseded) · Nougat TS clients ×5, extract-all-mess-papers.ts + siblings, pdf-processor.ts ×2 (dead)
Runtime, and what is not in the repo
  • Retrieval is a batch job by rule, never a page load. The one live acquisition-adjacent route is admin/effects/gap-search, candidate-only. Admin import routes (admin/import, import/manual, import-priority-papers, data/process) are active.
  • The BullMQ layer (apps/api/src/lib/jobs/processors/paper-processing.ts) never ran in production: embeddings write Math.random(), extract returns a hardcoded 1250, no worker host exists. admin/papers/bulk-process front-ends it, dormant.
  • docs/corpus-refresh-architecture.md designed a discover → acquire → extract → sync → embed → refit GitHub Actions DAG on 2026-05-30; the workflow file never landed. An earlier weekly-acquisition.yml (2026-05-11) was archived then deleted.
  • No Zotero importer exists; the two mentions are aspirational. Sci-Hub and LibGen are excluded by rule, so a paper that fails every provider stays unreachable.

Text extraction went through three generations

Nougat (2025) → marker (April 2026) → reading BioC-PMC XML directly (May 2026). The XML path keeps tables intact and yields about 10% more values while skipping OCR entirely. Marker replaced Nougat because it handles tables and multi-column layout better and runs roughly ten times faster, and it pulls from R2 on a cache miss. Value extraction downstream is the separate v2 extractor. The whole 2025 Nougat stack — batch runner, FastAPI queue server, URL extractor, “research-grade” variant, PyMuPDF fallback, plus five TypeScript clients — is superseded or dead but still in the tree, with run_extraction.sh still wired to pnpm nougat:*. Do not restart from it.

Two cautions that cost time. papers/staging/resolution.jsonl is dated 2026-04-29, so DOI matches will miss and stubs named Local PDF <sha8> will be proposed — fix the SHA→DOI map before accepting them into a 23k-row corpus. And marker-quality text is patent-only today: 22 marker text files exist on disk against 1,375 values_v2.json artifacts, so everything else is PyMuPDF output.

12Corpus run · The atomic unit is the experiment record, not the paper

also Part II  This is the Atlas's own summary of a corpus pass. Part II is the runbook itself, stage by stage, with the today-vs-target command pairs: 01 State · 02 Fix first · 03 The run · 04 Ontology duties · 05 Extraction schema · 06 Beyond · 07 Coupling.

Every layer of the corpus pipeline already has a working component. Each one is either not wired to anything, never run, or broken by a small bug. Nothing here needs to be built from scratch; the work is to connect what exists and run it in dependency order, with a measured gate at each step. Grading a paper as complete tells you almost nothing about whether any single experiment inside it can be modeled.

The interactive pipeline map and every stage in detail are in Part II · 00 Pipeline map, under the next tab.

The sequence, and why the order is not optional

  1. Coupling hygiene. Collapse the 31,360 duplicate sets, add @@unique([paperId, label]), then reuse an existing set instead of creating one. About 4 h. Mutates staging, so it needs a stated go. Do not reach for the runId variant — it was considered and rejected, because it makes each re-run legitimately distinct, which is the behaviour being stopped.
  2. Acquisition blockers. The 2.5 h list. Unblocks new papers reaching R2 and staging.
  3. Extraction. Run v2 recency-first on the 1,405 papers from 2025–26, emitting ConditionSet rows at sync time. (The three-pass extractor’s two bugs are unfixed and stay that way: it is not the extractor of record.)
  4. Vectorize. Abstract backfill, then the chunk table.
  5. Gates, then automation. Wire coupling-completeness.ts and the fixture recall check as pass/fail, then land the GitHub Actions DAG.

The order is not optional because a large extraction run on the current backfill reproduces the 24,540 orphaned condition sets of the 807-paper run at ten times the scale.

The extractor of record

Backlog is the 8,743 legacy papers whose rows carry no inline conditions; no backfill can ever couple them, so they must be re-extracted. Decided 2026-09-10 (PR #902): arm B — v2’s flat transport with the v1.2 sub-prompt families ported onto it as separate flat calls — run recency-first so 2025–26 lands before the backlog. Measured on a 20-paper slice: 95% precision [92–98], 62% modelable, $0.05 per paper, and 0 of 109 values sourced from other papers’ reference lists, where v2’s current prompt sourced 99 of 214. The three-pass variant, CMA v3 for bulk and the v1.2 orchestrator are decision history, not options. One caveat could reverse it: the gold set is LLM-adjudicated, not human-verified.

Coupling: the one-line version

41,080 ConditionSet rows describe only 9,720 real (paperId, label) tuples, and 27,419 sets have no linked output at all (staging, 2026-09-08). The backfill re-mints a fresh generation on every run because the unique key names a runId it never sets. The approved fix, in order: collapse the duplicates, add @@unique([paperId, label]), then reuse an existing set at sync time instead of creating one — about 4 h, mutates staging, so it needs a stated go. Setting runId was considered and rejected: it makes each re-run legitimately distinct, which is the behaviour being stopped.

The measured table, the coupling diagram, the five-step fix table and both root causes with their file and line numbers are in Part II · 07 Coupling.

Two systems share the word “experiment”

The platform is really about one thing: a complete experiment record — a design, a full set of conditions, and a co-measured outcome — not a paper. But the schema holds two separate systems under that name, and conflating them is the fastest way to misread this section. The corpus-derived half is ConditionSet plus ExtractedParameterData, everything measured above: one row per condition a paper reports, extracted at scale, coupled by conditionSetId. The user-authored half is Experiment, Run, and Measurement — a personal lab notebook for a signed-in researcher, entirely unrelated to paper extraction, with no shared code path. Everything above this point in this section is the first system.

The shape of that coupling — one paper, many condition sets, each carrying zero or more extracted values — is drawn once, in Part II · 07 Coupling.

The user-facing surfaces: real, but a different feature and hard to find

apps/web/src/app/[lang]/experiments/[id]/page.tsx and apps/web/src/app/[lang]/runs/[id]/page.tsx are genuine server components — they call prisma.experiment.findUnique / prisma.run.findUnique directly, with force-dynamic, and render real rows, not mocked data. But they read the user-authored trio above, not ConditionSet. There is no listing page for either route, and the only inbound link anywhere in apps/web/src is one card on /projects/[id], which itself requires a signed-in session scoped to that user's own projects. There is no page anywhere in apps/web that lets a visitor browse the corpus's coupled experiment records — that story lives only in coupling-completeness.ts script output and the onboarding docs.

RouteReadsReachable fromVerdict
/projectsprisma.project.findMany, scoped to ownerIdauth-gated, no discovery surface foundreal, personal
/experiments/[id]prisma.experiment.findUnique + Run, Project, MethodologyPresetone card on /projects/[id]real, no listing
/runs/[id]prisma.run.findUnique + Experiment; dumps inputs/outputs as raw JSONcards on /experiments/[id]real, no schema-aware rendering

The corpus-wide coupling story on this page — 1,760 complete records, the two bugs above, the class-median result at 14 — has no dedicated UI at all today.

13Priors & predictions · Hierarchical, log-scale, routed through a physics family

The priors turn a pile of comparable measurements into a defensible expectation with honest uncertainty. They are hierarchical, fitted on the log scale, and sampled with NUTS. A prediction is not a prior look-up: it routes to a physics family, draws the base estimate from the hierarchical prior, then passes through a calibration layer.

Power density, current density and resistance are log-distributed across orders of magnitude, so their arithmetic mean is dominated by a few large values and is not a sensible estimate. The earlier v0 approach — analytic method-of-moments — was retired in May 2026 as mathematically wrong for these metrics, not merely imprecise. Everything is now fitted on the log scale.

Partial pooling: sparse classes borrow strength

A shared hyper-prior sits above every system-type estimate. Classes with abundant data stay close to their own evidence; sparse classes are shrunk toward the global expectation rather than over-fitting three noisy points, and the amount of shrinkage is learned from the data, never assumed. Full pooling would erase real between-class differences — an MFC and an MEC are not the same device. No pooling would hand each sparse class a model built on too little data to trust. The hierarchy interpolates by the evidence, so priors degrade gracefully on rare classes instead of returning confident nonsense.

The three prediction steps

Step 1
Route
Dispatch by system class into one of the 5 physics families (02). Subtypes are features, not branches; combinations are coupling rules, not new predictors.
Step 2
Draw
Take the base estimate from the hierarchical log-normal prior for that metric in that class.
Step 3
Calibrate
A split-conformal layer widens or tightens the interval until stated coverage matches observed coverage on held-out data. It makes no distributional assumption — it measures its own errors and adjusts.

What a prior record contains

FieldMeaning
mu_log, tau_logFit-scale location and precision on the log scale
muBack-transformed central estimate
ci95_low / ci95_highBack-transformed 95% credible interval
scalelog or linear — how the metric was modelled
R-hat, ESS, divergencesMCMC health diagnostics, per stratum

Read via GET /api/parameters/[slug]/hierarchical-prior. Always check the diagnostics on sparse strata: a wide interval with R-hat ≈ 1.0 is trustworthy uncertainty; a tight interval built on three data points is not. A condition fitted as an outcome produces R-hat ≈ 3 and ESS 2 — the failure mode the July modeler agent found (15).

The output contract

Every numeric output ships this shape {
  "value": 512,
  "unit": "mW/m^2",
  "ci_low": 180,
  "ci_high": 1450,
  "confidence": 0.62,
  "source": "hierarchical_prior_v2",
  "data_status": "ok"
} Including where it came from. Two refusal states are deliberate: insufficient_inputs means the routed class does not have enough evidence to answer; awaiting_artifact means a computed input is missing. A fabricated confident number is worse than an honest gap, so the engine returns the gap.
Two live priors artifacts coexist — refit them together
  • v2, a Student-t per-class stratified fit, is the served production model behind the main prediction API.
  • v1, a pooled log-normal fit, still backs the parameter-detail pages, the wastewater overlay and the DB-mirrored prior path.
  • Refitting only one leaves the other stale — the fifth harmonization hard rule in 08. The 2026-07-21 batched refit produced 102 params, 202 strata, 58 converged, 89 LOO-reliable; 8 large-N slugs timed out at 180 s each and are omitted from served v2. The v1 parameter-detail path still has no serving gate.
The published coverage figure was computed on a split that does not hold

The public methodology page reports 97.98% out-of-sample coverage across seven strata (n = 940), MFC + MEC only. That number was computed on the stored fit_validation_role split, which 14 Held-out error shows a classifier can separate from design covariates alone — so it is not a held-out result. Treat it as unmeasured until the benchmark is re-run. The audit itself, its adversarial AUCs, and what every model scores on a paper-disjoint split are in 14.

Treat a prediction as a prior to be updated by your own data, not as ground truth. For a novel design the interval will often span an order of magnitude — and given a corpus CoV near 1,285%, anything narrower would be dishonest. Stated limitations from the public page, worth carrying here: extraction success is not currently measured; the Gemini fallback over-assigns OTHER (about 52% versus about 14% on Haiku); holdout validation covers MFC + MEC only; canonical mapping sits near 51.6% on that page’s count against 45% measured on staging 2026-09-07 (08); the physics layer is a 0-D mean function and the calibration is plain split-conformal.

14Held-out error · What every predictor scores on paper-disjoint data

Measured 2026-09-02 on the local corpus, read-only, branch feat/bes-benchmark-v1 (not merged into development or main as of 2026-09-08; the package lives at services/ml-engine/training/benchmark/ on that branch only). The first paper-disjoint, leakage-audited held-out set for BES performance prediction. The result is a clean negative. Markdown port: docs/onboarding/held-out-benchmark.md; visual companion: the bes-benchmark-v1 artifact.

×6
typical held-out error on power density, every model and baseline alike
0.80 dex median |error|
0 of 4
targets where any model passes the ≥10% RMSE-reduction gate vs the class median
M1 −3.2% · M2 −6.0% · M4 −21% on power
96.7%
rows the served physics predictor cannot score — it needs an HRT the paper never reported
3,463 of 3,593 · HRT alone missing on 959
0%
coverage of the served ±25% band on power density where it predicts at all
19% on COD removal
89–91%
coverage of the served priors' 90% band — the one honest interval
4.1 decades wide on power
TargetTransformRowsPapersWithin-paper SDBetween-paper SDB1 class×domain median (the bar)Best model
power_density_areallog10 W/m²1,0373050.61 dex1.10 dex0.80 dexM1 0.81 · M2 0.81 (fail)
current_density_areallog10 A/m²9112420.61 dex1.16 dex0.78 dexB2 0.72 · M2 0.82 (fail)
coulombic_efficiencylogit6001900.962.041.19M2 1.10 (fail)
cod_removallogit1,0452910.861.351.06B0 0.94 · B2 0.99 (fail)

3,593 rows from 590 papers after eight recorded filters (7,336 → 6,884 vintage v2/curated → 6,466 verifier not failed → 5,338 not cited from another paper → 5,281 no physics/dedupe flags → 5,023 BES not review → 4,965 physical bounds → 3,593 exact duplicates removed). Three splits pass the paper-level adversarial audit (hash 0.46–0.53, group 5-fold 0.41–0.62, temporal ≥2024 0.43–0.58); the stored fit_validation_role split leaks on every target (0.77–0.92). Between-paper spread is roughly twice within-paper, so there is explainable variance; the covariates the corpus holds (temperature 49%, pH 44%, anode material 74%, HRT 6%, inoculum 0%) do not explain it.

The claim versus the file

ArtifactWhat it saysWhat it actually is
calibration.json
(live, 2026-07-22)
517 rows · 95.9% coverage · ECE 0.036A random row-level 80/20 split of a 7-feature RandomForest. Papers straddle train and test. Its own power-density test R² is −3.7 × 10⁶. σ is re-estimated from the same held-out residuals, then pooled across volts, ohms and W/m², so the coverage is close to tautological.
“372 held-out, ECE 1.96%”quoted on 6 surfacesNo such file. The lab app bundled a third vintage (377 rows) at build time, so lab and web showed different numbers for the same concept.
fit_validation_role split97.98% OOS coverage · 940 obsNot exchangeable. A classifier on design covariates alone tells fit papers from validation papers at AUC 0.77–0.92 against a permutation null ≈ 0.50. Every May coverage, conformal and PSIS-LOO number was computed on it.
Served 95% interval
/api/ml/predict
“calibrated”A hard-coded ±25% band multiplied by a conformal q̂ fitted to a different model’s residuals. The version-drift warning fires on 100% of responses, correctly.

The split audit — adversarial AUC per target

A LightGBM classifier on design covariates, one vector per paper, 5-fold. 0.5 is exchangeable; above 0.65 the split leaks. The stored split leaks on every target; the three benchmark splits sit at the null, which is what makes them usable.

SplitpowercurrentCECODVerdict
stored fit_validation_role0.8990.9200.7700.845leaks on all four
hash split0.5190.4700.4640.531at the null
group 5-fold0.6000.6180.4120.602passes
temporal ≥ 20240.4290.5240.5800.491passes

The scorecard: every model ties the class median

Gate: a ≥10% RMSE cut against B1 (the median of training papers in the same system class and application domain), with disjoint CIs, on both paper-disjoint splits, direction-correct on 2024+. Median absolute error, group 5-fold out-of-fold, paper-grouped bootstrap CIs.

ModelpowercurrentCECOD
M1 · LightGBM, 35 design covariates−3.2%−1.5%−4.0%+1.5%
M2 · class median + residual GBM−6.0%−4.0%−2.8%−0.9%
M4 · RidgeCV, one-hot + imputed−21%−33%−24%−8%
Same gate, one peak value per paperfailfailfailfail
B1 · class×domain median (the bar)0.795 dex0.784 dex1.1941.056
B2 · served priors v20.8310.7191.1990.989
B3 · served physics, where it can predict0.462 n=15—1.953 n=221.306 n=58

Top GBM gain features are substrate concentration, publication year, reactor volume and electrode area at 9–18% each — no physical driver dominates, and publication year appearing at all is a tell. The covariates the corpus holds for these rows are thin: temperature 49%, pH 44%, anode material 74%, HRT 6%, inoculum 0%.

Intervals: only the priors’ band is honest, and it is four decades wide

Intervalcoverage: powercurrentCECODmean width (power)
B2 · served priors v20.8940.9010.9150.8224.10 decades
B1 · residual band0.8820.8460.8670.8564.02
M3 · conformal quantile regression0.9130.9190.8950.8734.43
served ±25% band0.000—0.0000.1900.22

Target coverage is 0.90. Coverage is bought with width: a 4-decade band on power density spans 0.0001 to 10 W/m². M3 is never narrower than the class-median band. The served band is narrow and wrong.

The served predictor abstains on 3,463 of 3,593 rows

predictForSystem requires temperature, pH, HRT and COD. It predicted on 118 rows (3.3%); 3,463 returned insufficient inputs and 12 were an unroutable class or unit. The top missing-input combinations: HRT alone on 959 rows, all four on 482, pH + HRT on 453, temperature + HRT on 412, temperature + pH + HRT on 344, HRT + COD on 247 — 15 combinations in all. HRT is the binding one, and batch reactors do not have one. The old skill harness filled the gaps with 30 °C, pH 7, 12 h and 1000 mg/L; the benchmark does not, which is why the abstention is visible for the first time.

What “state of the art” means from here
  • Published R² of 0.95 to 0.997 for MFC power density come from within-lab random splits of one dataset — the same artefact this platform diagnosed in its own trainers, where train R² was 1.000 and out-of-fold below zero.
  • On a paper-disjoint benchmark the field’s number, measured here for the first time, is a ×6 typical error and R² ≈ 0 for any covariate model.
  • Accuracy will not move with another model. It moves when complete design → conditions → outcome tuples exist for enough papers — a targeted re-extraction, not a fit. That is the whole argument for the coupling work in 12.
  • The benchmark itself is the contribution; the honest product claim is calibration, not accuracy. 34 offline tests and the leakage audit ship in manifest.json; the handoff is docs/handoffs/2026-09-02-bes-benchmark-v1.md.

Next, as of 2026-09-02 (none landed by 2026-09-08)

  1. Rewire the MFC interval: serve the priors' predictive band instead of σ = 25% of the point (conformal-apply.ts:417), and return insufficient_inputs instead of the 46-paper heuristic when design covariates are absent. The 2026-09-08 MFC commits (bf69e8a49, 4aa9cb2e3, 6a81c5840) fixed material-slug resolution and the null contract, not the interval.
  2. Re-run run_benchmark_v1; the served band should then cover ~90% instead of 0%.
  3. Score the four expert datasets tagged holdout_set as an external test.
  4. Smoke a targeted re-extraction of full tuples on 20 papers before spending on the corpus (decision taken in streamlined-corpus-run.md stage 7; not yet run).

Since 2026-09-02 on the same branch: a within-study contrast census (2026-09-07, staging, read-only) found 18 factor × outcome pairs clearing 10 papers; continuous vs batch operation raises current density ~0.47 dex and power ~0.40 dex in 80–83% of papers, and raising R_ext lowers current density in 8 of 9 papers — the physics direction cross-paper models never recovered. Do not conflate with the separate cohort benchmark export (scripts/derived/build-benchmark-export.ts, regenerated as schema 2.0 from staging in PR #862 on 2026-09-08), which feeds /methodology#schema.

15The week the diagnosis moved upstream

Three Claude Managed Agents ran for the first time, then at corpus scale. Each proved from its own direction that the platform's problems sat upstream of everything being tuned. Click any dot for the detail; bold-ringed dots are the ones to know.

Observations coupled to a condition set
13.6%100%
1,343 of 1,343 · overnight sweep d4fa68922
pH values physically impossible
80.8%0
4,630 of 5,732 → 0 of 60 papers
Outcome priors converged, real gate
12 / 120 / 71
the modeler corrected its own brief
Strata served as "95% CI"
13527
75 withheld · serving gate #736
Product-source type errors
862395
gate was false-green before #735
Cost of the overnight sweep
$70.68
60 deep-extracted · 1,400 screened
Corpus agents Corpus data Serving & priors Gates & hygiene know this one

The causal chain (CORPUS-DIAGNOSIS.md §4)

regex extraction (86%) + no layout model (99.5%)
→ conditions indistinguishable from outcomes (63% of "observations" are conditions)
→ condition_set_label 13.6%, uncertainty 2.4%, replicates 1.1%
→ no within-paper contrasts → 9/10 effect targets non-viable
→ conditions modeled as outcomes → 11 degenerate priors (R-hat ~3, ESS 2)
→ /meta-analysis renders nearly every forest row "not converged"
→ /collab/korth ships 135 priors, 1 converged, to an external collaborator

The schema was never the problem. conditionSetId, uncertaintyPlus/Minus/Type, derivationMethod, bbox all existed. They were unpopulated.

The three agents

AgentJobGate
screenertopic (core / peripheral / off_topic) × doc_type, verbatim evidence, 0–1 confidence, printable-ratio triage before reading8 pass/fail criteria; any write fails the run
extractorevery outcome tied to the condition set it was measured under; chases reference-electrode convention and substrate basisorphan outcomes declared, never guessed; writes drafts, never the DB
modeleroutcome vs condition vs context; merge slugs onto the 687-row vocabulary; serve / withhold"do not soften the gate"; found omnibus_pass rewards vacuity

Sessions run server-side; IDs in EXTRACTION-LEDGER.csv; outputs under docs/corpus-screening-agent/run-0N-*/. Agents hold no credential — presigned URLs only, because vault placeholders cannot sign SigV4.

16Roadmap · Twelve months, six pillars, one brewery

8 Sept 2026 to 27 Aug 2027, in three phases and eighteen workstreams, each with dated deliverables. The roadmap is organised around one first-of-a-kind brewery installation; every other pillar either feeds it or funds it.

Open the full page for the interactive version and every detail.

P0
Prove & prepare · 8 Sept 2026 – 19 Dec 2026
Gate: none — start now
P1
Fund & partner · 5 Jan 2027 – 30 Apr 2027
Gate: funding event OR 2 paid feasibility studies signed · solo → first hire
P2
Build & validate · 3 May 2027 – 27 Aug 2027
Gate: wet-lab partner signed · 1–3 hires

Two dates are hard. The Villano meeting, with the ISMET president, sits at roughly 15 Oct 2026. The P1 gate, a funding event or two paid studies, falls on 19 Dec 2026.

Six pillars, eighteen workstreams

PillarWorkstreams
A Platform & observabilityA1 Product analytics + observability · A2 Agentic ops loop · A3 Reliability + tech-debt burndown
B Science & dataB1 Data acquisition + harmonization + analysis · B2 Model quality + validation · B3 Publications + reporting standard
C Brewery FOAKC1 Brewery FOAK design dossier · C2 Wet-lab partner + de-risking experiments · C3 Pilot instrumentation / DAQ
D Community & partnershipsD1 ISMET partnership · D2 Researcher outreach (tester + collaborator pipeline) · D3 Open source + community
E CommercialE1 Go-to-market + first customer · E2 DAC track · E3 Opportunity portfolio
F CompanyF1 Funding + runway · F2 Legal / IP / data licensing · F3 Team + hiring
The crunch
  • The crunch is the five weeks to 10 Oct 2026: A1-M1, A1-M2, A1-M3, A3-M1, B1-M1, B2-M1, B3-M1, C1-M1, C1-M2, C2-M1, D1-M1, D2-M1, D2-M2, E1-M1, F1-M1, F2-M1. A1-M3 and B1-M2/M3 slip first if something has to.

Where this connects: 12 Corpus run above is workstream B1's ground truth. Its measured coupling and acquisition state is what B1's milestones are scheduled against.

17Open gaps · Ranked, dated, and what unblocks each

Two views of the same backlog. The first is by area, drawn from the onboarding docs; the second is the platform’s own P0 to P2 ranking — P0 blocks the corpus machine, P1 bounds what can be modeled or served, P2 is hygiene. Both are dated. Nothing here is a wish list: every row names what unblocks it.

The backlog, ranked

One register. P0 blocks the corpus machine, P1 bounds what can be modelled or served, P2 is hygiene; rows marked — are known gaps that have not been ranked against the others yet.

SevAreaGapCurrent stateWhat unblocks it
P0acquisitionAcquisition stalled since 2026-05-09Five blockers keep weekly_pipeline.sh from completing; of the 1,405 papers from 2025–26 only 322 are PDF-backed (23%). 2026-09-08Fixes 5–7 landed 2026-09-10 (R2 push + hydrate, PDF triage, --input-format pmc-xml). Remaining: a real recency-first run with R2 and DB credentials.
P0couplingCoupling backfill mints duplicate ConditionSets41,080 rows exist; only 9,720 are real. The backfill has no runId, so it re-mints sets it cannot recognise as present. 2026-09-08Code landed (#893: collapse-duplicate-condition-sets.ts, unique-key migration, reuse at sync). Remaining: run the collapse and migration on staging, then prod. Do this before any large run.
P0pipelineExtraction and acquisition are not scheduledweekly_pipeline.sh and simple_value_extractor.ts run by hand; corpus-refresh.yml is a 2026-05-30 design with no code. Only refits run weekly.Land the GitHub job-DAG, or a scheduled routine step for discover → acquire → extract with the budget cap.
P1ontologyCanonicalization coverage, not extraction, bounds modelability55% of EPD rows have a NULL canonical slug on staging (80.7% when first measured on local). Modelable 31,924 against a ceiling near 40–50k.Extend the alias maps and canonicalize-name.ts; re-run refresh-all steps 2–3.
P1pipelineSection-chunk vectors live on disk onlyembed.py writes chunk vectors to disk; no DB table holds them, so nothing served can search full text. Abstracts cover 13,788 of 23,579 (58%). 2026-09-08The chunk table landed (migration 20260910120000_add_paper_chunk_vectors). Remaining: the embedding backfill.
P1schemaTwo geometry tables not unifiedReactorGeometry (~47 typed columns) feeds the resource-recovery demo and quality scripts; the paper 3D route reads ExperimentalContext + ConditionSet. 2026-09-08Pick one as the source of truth and make the other a view over it.
P1schemaPapers ↔ Materials / Microbes orphanedanodeMaterials and cathodeMaterials are strings; the PaperMaterial and PaperMicrobe junctions are mostly empty.Backfill the junctions from extraction. The read routes exist since 2026-10-02: /api/materials/[id]/papers and /api/microbes/[id]/papers (match on canonicalId; MaterialPaperCrossref is not merged in).
P1pipelineUnread corpus1,444 BioC-PMC XML files parsed but never re-extracted (~$70–150); 2,075 no-DOI papers unreachable; MDPI, Wiley, ACS and Elsevier 403s.A cost-gated v2 run over the XML; the curl_cffi path for no-DOI URLs.
P1uiConfidence and integrity caveats are not surfacedThe parameter detail page now carries an “Honest framing” block citing the CoV and SCIENTIFIC_INTEGRITY.md, and v1 prior intervals are gated (see below). The parameter API responses still carry no aggregate confidence.Add confidence to the parameter responses; a collapsible callout on the parameter pages.
P1priorsGP-SCM fits are data-limited, and can be silently dark in prod12 of 27 MFC nodes fitted; new fits are interpolation artifacts (energy_efficiency flat at 43.02%), toc_removal has one joint observation. A blank GP_SCM_SERVICE_URL disabled the pillar for about seven weeks in July 2026 while it looked wired.Targeted paid re-extraction for the biofilm and biomass nodes — do not lower --min-samples. Check the Vercel env on messai-api.
P1pipelineWeekly routine output never landedNo npe-health.json exists anywhere in the tree, so either the routine has never run with --commit or its commit step has not fired.Run weekly-ml-audit.sh --commit once by hand, then confirm the scheduled session exists.
P1schemaProd FK under-populationAbout 33k ExtractedParameterData rows in prod lack parameterDefinitionId, which affects FK and DAG read paths, not the modelable count.A ref-gated backfill, staging first.
P2mlNo feedback or retraining loopOutcome capture landed (PredictionOutcome, RecordResultsForm, /experiments/predictions, shared computeDrift). No drift-triggered refit runs yet.A feedback widget posting to the predictions route; weekly regeneration when drift exceeds 5%.
P2chatChat tools missing for four packagesqueryDatasets (catalog.json) and queryMethods (the hunter physics-violation flags) landed 2026-10-02. mess-hypotheses holds only synthetic generators and mess-learning holds unsourced figures, so neither is wrapped.One tool per package in @messai/ai-chat.
—infraNo rotation runbook for R2 keys or the Supabase service-role key. Rotating R2 revokes every presigned URL at once.not yet rankeda decision on cadence and owner
—pipeline~1,354 screened-in papers still unextracted; full backlog ≈ $929 Opus / $368 Sonnet vs a ~$100 envelope.not yet rankedan explicit go
—pipelineAgent drafts sit as PENDING; nothing promotes them. The review queue now reads and writes ExtractedParameterData.validationStatus (approve/reject, 2026-10-02); validation/[itemId], export, jobs, metrics and user-activity under admin/extraction still read the missing paper columns.not yet rankeda review-queue UI
—extractorThe v2 system_class prompt taught the legacy 30 buckets while the enum was the canonical 17; 14 of 18 types got the literal string "undefined" as guidance. Fixed 3d89f2901, 2026-09-08, 5-paper smoke passed.not yet rankedre-extract the classes affected before the fix; the MEC smoke paper a5ccb99a still fails on both code paths
—taxonomyScreener verdicts (topic / doc_type / cull) never reach the DB; Tier 3 human review has columns but no writer; 688 papers flagged taxonomyNeedsReview.not yet rankedthe loader exists (scripts/db/load-cull-manifest.ts, dry run by default); smoke 5 rows on staging, then apply
—taxonomyThree parallel system-type columns (primarySystemType, systemType with 8 indexes, v1_1_primary_system_type with no writer); benchmark export reads the frozen one first.not yet rankeda migration that collapses to primarySystemType
—ontologyThe pinned parameter-definitions-rich.json (825) is 10 definitions behind staging (835), so the per-slug fixtures, the npm package and the 826-node KG snapshot are stale. 29 slugs with no definition (membranematerial, cathodematerial, anodematerial, maxpowerdensity, mode, design …) carry over 10,000 rows with no unit, range or FK.not yet rankedre-export rich.json, regenerate fixtures + snapshot; add or alias the 29 orphan slugs
—agentsPresigned-URL expiry blocks a scheduled v3 deployment.not yet rankedrun-time minting or per-run attachment
—acquisitionweekly_pipeline.sh now pushes new PDFs to R2 (2026-09-10) and the skill lists all 9 providers (2026-10-02); neither has run end to end on a fresh batch.not yet rankedfix 5 of the 2.5 h list; a one-line skill edit
—priorsThe v1 parameter-detail route now runs the same serving gate as /api/ml/predict (2026-10-02), but its validation never pairs with the served fit, so it gates on fit quality only. 8 large-N slugs omitted from served v2 at the 180 s timeout.not yet rankednamed follow-ups from #736 / 1e77b1b3d
Resolved since the docs were written — verify before re-flagging
  • Type gate false-green: check-cross-package-imports.sh now fails on dangling node_modules symlinks, on a tsc that crashes or is missing, and on a run that loads too few app source files; each case was shown to fail (2026-10-02).
  • BullMQ scaffold archived to archive/2026-10-bullmq-scaffold/; bullmq/ioredis dropped from apps/api (2026-10-02).
  • Methodology pages describe the v2 extractor (203549b7, 70991c3c; the last lab metadata strings 2026-10-02).
  • NextAuth: only @next-auth/prisma-adapter is installed; @tanstack/react-query is only used by a test util, so devDependencies is correct.
  • fra1 region pinned in apps/api/vercel.json; the session-artifact backup script imports the apps/api R2 client (2026-10-02).
  • EXTRACTION-LEDGER.csv reconciled from disk by refresh-ledger.py: 178 extracted, 2 missing_output (2026-10-02).
  • bes-benchmark-v1 merged with the calibrated MFC interval (0ed50246, #891).
  • ProcessFlowDiagram.tsx is marked as deliberate deck artwork, not a P&ID; engineering drawings stay in @messai/pid-schematic.
  • Chat bifurcation: one canonical /api/chat (2026-05-12).
  • Priors JSON validity: v1 and v2 both parse cleanly (715b63e96).
  • Per-class ML routing: six non-MFC classes route to the analytical predictor (4778f8589, bef7acedd). The lab UI still passes proxy inputs.
  • Harmonization quick wins merged (PR #501): +434 modelable, propagated to staging and prod 2026-07-06/07.
  • GP-SCM all-parents dropna: present-parent fit measured 9 → 12 nodes (2026-07-31, branch unmerged).
  • Legacy extractors already read ANTHROPIC_MODEL with a sane default — that 2026-05 item is closed.
Source docs: docs/onboarding/README.md · infra-access-r2-staging-prod.md · paper-pipeline.md · docs/handoffs/2026-07-19-24-corpus-agent-work.md · docs/corpus-screening-agent/CORPUS-DIAGNOSIS.md. Generated 2026-09-07; section 17 re-verified against the code 2026-10-02.
18Glossary · Every term this page uses, in one place

Science terms are as defined on messai.io/learn/science; platform terms are as used in the repo. Where a term means two things, both senses are given — that ambiguity has cost time before.

BES
Bioelectrochemical system — the umbrella term for any device where microbes (or enzymes) catalyse redox reactions at electrodes wired to an external circuit.
MES
Two senses. As an umbrella: microbial electrochemical system, the whole field. In the extraction schema and primarySystemType: Microbial Electrosynthesis, cathodic CO₂ reduction — one of the 14 devices.
EET
Extracellular electron transfer — how bacteria move respiratory electrons onto an electrode, by direct contact, conductive nanowires, or soluble mediators.
DIET
Direct interspecies electron transfer — electrons passed directly between species through pili or conductive minerals; the basis of syntrophy and electromethanogenesis.
Anode / cathode
The oxidising electrode, where the biofilm gives up electrons, and the reducing electrode, where they are consumed.
Overpotential (η)
Voltage beyond the thermodynamic minimum needed to drive an electrode reaction at a useful rate. Split into activation, ohmic and concentration terms.
Exchange current density (j₀)
How intrinsically fast an electrode reaction runs at equilibrium — higher is better.
Tafel slope
The slope of overpotential against log-current in the kinetically-limited region; encodes the transfer coefficient.
Limiting current
The maximum current a diffusion-limited reaction can sustain, reached when surface reactant concentration falls to zero.
Coulombic efficiency
Fraction of the electrons available in the substrate that are actually recovered as current. Bounded 0–100% and diagnostic of electron leakage.
Power density
Power normalised to electrode area (areal, mW/m²) or reactor volume (volumetric, W/m³). Keep the two apart — they are separate canonical kinds.
Monod kinetics
Saturating relationship between substrate concentration and microbial growth or uptake rate, set by μ_max and the half-saturation constant K_s.
EIS / Randles circuit
Electrochemical impedance spectroscopy, and the equivalent-circuit fit that separates solution resistance, charge-transfer resistance, capacitance and diffusion.
Physics family
One of 5 groupings — anodic oxidation, cathodic reduction, ion transport, selective reduction, sensor/photo — that predictions route through. A calibration and out-of-distribution stratum, not a predictor.
Canonical slug
The normalised identity a free-text parameter name is mapped onto so values from different papers can be compared and pooled. Layer 2 of the five-layer value ontology (09).
Kind
Layer 3: the SI quantity a slug belongs to (SLUG_TO_KIND). A slug outside it is never SI-normalised and never modelable. This, not the category, is the modelable gate.
ConditionSet
The within-paper coupling record that binds an outcome to the operating point it was measured under. Without it a number has no context.
Experiment record
The atomic unit for modeling — not the paper. Grading a paper as complete says almost nothing about whether any single experiment inside it can be modeled.
Hierarchical prior
A Bayesian prior that pools information across system types so sparse classes borrow strength from data-rich ones, with the shrinkage learned rather than assumed.
Conformal calibration
A distribution-free method — here, plain split-conformal — that adjusts prediction intervals so stated coverage matches observed coverage on held-out data.
ECE
Expected calibration error — how far stated interval coverage drifts from observed coverage. Lower is better.
CoV
Coefficient of variation, standard deviation over mean. For power density it is ≈ 1,285%, which is why point estimates are meaningless without an interval.
dex
A decade on a log₁₀ scale. A median absolute error of 0.80 dex is a typical miss of about ×6.
Adversarial AUC
The score of a classifier trained to tell a train split from a test split using covariates alone. 0.5 means exchangeable; above 0.65 the split leaks and every metric computed on it is optimistic.
resolveDbTarget()
The only sanctioned way to resolve a remote database, gated on the Supabase project ref rather than the hostname — staging and production share a pooler host, so a hostname check proves nothing.
Physics families vs. system types
17 primary system types (14 devices + 3 meta classes) exist in the taxonomy; predictions route through 5 physics families. Branching on all 17 would train rare classes on a handful of papers each.