For most of the living world, the compound-by-species toxicology matrix is almost empty. The question is whether historical organism-level outcomes must always be the starting point for prediction.
Dr Alexander D. Kalian posed the problem sharply in a recent thread on biological data and AI: how can we predict the effects of chemicals across millions of undiscovered and understudied organisms when most species have little or no toxicology data?
His case in point was deliberately difficult: a rare symbiotic fungus found on a single tree in the Amazon, represented by perhaps five papers in twenty years and never cultivated in a laboratory. If we introduce a pesticide, pollutant, or microplastic, how could any predictive system determine what it might do to that organism?
If a reliable prediction requires a large toxicology dataset for that particular fungus — or several sufficiently similar organisms — then we are effectively stuck. The necessary data could take decades to produce, may never be funded, and in many cases may never exist at all.
Across millions of known and undiscovered species and an effectively unbounded set of chemical exposures, exhaustive measurement is not a scalable answer.
We can begin without species-specific toxicology data.
Sequence, structure, and physical interaction can support a ranked molecular hypothesis even when the organism has never appeared in a toxicology training set.
We cannot skip biology.
A molecular interaction is not yet an organism-level outcome, and neither is an ecosystem prediction. Exposure, metabolism, compensation, physiology, and ecology still have to be resolved.
The almost-empty compound-by-species matrix
Imagine ecotoxicology as an enormous matrix. Every row is a compound. Every column is an organism. Every cell asks what that compound does to that organism. For a rare or newly discovered species, the entire column may be blank.
Data-driven models can generalize when they have enough populated cells, sufficiently similar compounds, related organisms, or transferable biological outcomes elsewhere in the matrix. When the organism is poorly characterized, evolutionarily distant, or absent from the training distribution, that basis weakens. More compute does not populate the missing cells.
The absence of organism-level data is real. The open question is whether the first useful calculation must depend on organism-level outcomes at all.
Begin with what the compound physically encounters
A chemical does not initially encounter an ecosystem-level outcome. It encounters matter: a protein, ion channel, enzyme, membrane, receptor, nucleic acid, or metabolic process. This first physical disruption is often called the molecular initiating event.
Each step adds context and uncertainty. A molecular interaction depends on geometry, energetics, accessibility, and local environment. A whole-organism outcome also depends on exposure, absorption, metabolism, repair, redundancy, development, and behavior. Population and ecosystem outcomes add reproduction, selection, competition, migration, food webs, symbiosis, transport, and indirect effects.
Matter Computing does not collapse that chain into one number. It changes where an investigation can begin.
From computable matter to Matter Computing
Matter Computing is the deterministic calculation and search of properties, interactions, mechanisms, and physically admissible states directly from structure. FluxMateria is the platform implementing this approach across chemistry, materials, pharmacology, and an emerging Genome Physics program.
At the molecular-interaction layer, a compound structure and a biological target structure can be evaluated without first training on species-specific toxicology outcomes. This does not make every downstream conclusion data-free. It means something more precise:
Historical organism-level outcomes do not have to be the source of the initial compound-target interaction calculation.
The species must still be found, sampled, sequenced, annotated, and represented credibly. What changes is the requirement for thousands of prior compound-species measurements before useful investigation can begin.
Why targets may transfer when outcomes do not
Whole-organism toxicology often transfers poorly between species. The same molecular disruption can produce different outcomes because metabolism, exposure routes, repair systems, physiological redundancy, life cycle, and environment differ.
The earlier links in the chain can be more transferable. Protein families, binding-site geometries, ion channels, membrane systems, and biochemical pathways recur across organisms separated by substantial evolutionary distance. A rare species may have no toxicology record while still possessing identifiable enzymes, receptors, transporters, and structural proteins.
Conservation is evidence, not equivalence
Family-level homology does not guarantee binding-site conservation. A small number of substitutions can shift affinity substantially or abolish binding. The presence of a related target is therefore a hypothesis to test, not a verdict to trust.
The revised question becomes: does this organism possess a molecular target that the compound is physically likely to bind, inhibit, activate, or disrupt, and how much confidence should be assigned to that target representation?
The present FluxMateria capability boundary
The distinction between what can be evaluated now and what remains a research program is essential.
Structure to interaction
FluxMateria's BioTarget capability evaluates binding affinity, target identification, and mechanism-of-action hypotheses from molecular structures across a broad target catalog.
10,065 targets across five kingdoms; CASF-2016 Pearson r = 0.772; ChEMBL mechanism validation accuracy = 91%.Interaction to effect
ADMET and DILI systems propagate molecular evidence toward safety and disposition endpoints, but their present validation is centered on human and conventional model-organism biology.
Three strict #1 public-comparator endpoints; Caco-2 matches the listed reference SOTA; DILI AUROC = 0.9597 on the comparable TDC binary task.Sequence to structure
Protein structure reconstruction is advancing, but performance remains uneven and context dependent. Membranes, cofactors, oligomeric state, modifications, pH, and solvent can all determine whether a model is usable.
A sequence is a starting point, not an authoritative physical target model.The BioTarget benchmark, ADMET benchmark, and DILI benchmark document the current public evidence and scope. These results support molecular hypothesis generation. They do not establish unchanged transfer across fungi, plants, protists, invertebrates, and every other branch of life.
What would change for the Amazonian fungus?
A conventional data-first program might begin by generating enough exposure and toxicity measurements for the fungus to train or validate a predictive model. For a rare symbiont with little commercial visibility, that program may never happen.
A Matter Computing program would begin differently.
Collect and sequence
The organism must still be encountered and described. Fragmentary assemblies and uncertain annotation should be retained as explicit upstream uncertainty.
Reconstruct molecular machinery
Identify proteins, membranes, pathways, transport systems, receptors, enzymes, and other structures relevant to plausible exposure mechanisms.
Calculate plausible interactions
Screen the compound against reconstructed targets and identify candidate binding, inhibition, activation, displacement, or disruption events.
Rank initiating events
Prioritize the strongest, most structurally credible signals and identify pathways where several independent interactions converge.
Propagate cautiously
Ask whether the target is expressed, reachable, persistent, compensable, and consequential at realistic exposure levels.
Run targeted experiments
Test the targets, pathways, doses, and exposure conditions most strongly indicated by the calculations instead of distributing effort blindly.
Uncertainty must remain attached to the result
A useful system should not compress every stage into one confidence score. It should distinguish the provenance of each claim.
Weak structural confidence, unknown cofactors, uncertain exposure, missing pathways, unusual metabolism, and evolutionary novelty should reduce confidence visibly. The objective is not zero experiments. It is fewer blind experiments and a more defensible reason for choosing each one.
The genome-scale path
The longer-term goal is to apply this reasoning across much larger portions of an organism's molecular system: genome-wide target reconstruction, competing compound interactions, metabolic transformation, pathway propagation, transport, organism-specific vulnerability, and eventually population and ecosystem consequences.
The Genome Physics Initiative describes this as a staged scientific program rather than a completed capability. Genome-scale execution will require stronger sequence-to-structure performance, target annotation, comparative biology, experimental datasets, infrastructure, safety review, and independent validation.
Current scientific boundary
The evidence does not support claiming that an ecosystem can already be computed from a genome. It supports testing whether the absence of historical species-specific toxicology data must remain the immovable first barrier.
This does not replace field biology
Matter Computing does not remove the need for naturalists, taxonomists, genomicists, toxicologists, ecologists, or field researchers. It makes their observations and samples more computationally actionable.
The organism still needs to be found. Habitat, symbiosis, developmental state, exposure, metabolism, population dynamics, and ecological relationships still matter. The proposed workflow is therefore not "predict everything without data." It is:
Collect it once. Sequence it. Reconstruct its molecular machinery. Calculate the most plausible initiating interactions. Test the pathways the physics flags. Build outward toward organism and ecosystem effects.
That reframes an almost empty-data problem as a computable hypothesis-generation and experimental-prioritization problem. For organisms that may never receive conventional toxicology programs, it could provide a practical starting point where none exists today.