Executive abstract
What this paper proposes.
This paper defines GenovaX as a proposed computational workflow for generating, predicting, filtering, ranking, and handing off molecular candidates for experimental review. The core requirement is evidence gating: a computational score remains a model output, not evidence of binding, biological activity, safety, efficacy, manufacturability, or therapeutic value. The architecture begins with a defined context of use, preserves target and dataset provenance, separates generation from prediction, applies developability and liability filters, records uncertainty and model disagreement, requires expert review, and produces an experimental handoff package with predeclared acceptance criteria. The existing Planckchron lysozyme benchmark demonstration is computational only and is not a therapeutic program. No candidate described by Planckchron has been represented here as experimentally validated, clinically tested, approved, safe, or effective.
At a glance
Four operating conclusions.
- 01
Structure and interaction predictions generate hypotheses; they do not establish biological truth.
- 02
The context of use determines the evidence required from an AI model.
- 03
Model disagreement and negative results belong in the ranking and handoff record.
- 04
The workflow is complete only when a reviewer can reproduce why a candidate advanced or stopped.
Reference architecture
A control and evidence flow.
Frame
Define the biological question, target, intended use of model output, prohibited interpretation, and decision owner.
Generate
Create or retrieve candidate sequences or structures with complete provenance and diversity controls.
Predict
Run structure, interaction, and property models while preserving versions, parameters, confidence, and failures.
Filter
Apply transparent developability, liability, novelty, and feasibility rules without hiding tradeoffs.
Review
Use multidisciplinary expert review to assess biological plausibility, uncertainty, and experimental value.
Test
Hand off a bounded set with assay plans, controls, acceptance criteria, and a negative-result record.
Proposed designβnot a representation of deployed capability.
1. Context and objective
Modern structure and interaction models can reduce the cost of forming hypotheses, but their outputs are easy to overinterpret. A high confidence score may reflect confidence within a model's learned distribution, not proof of binding or useful biological effect. A visually plausible complex may still fail expression, stability, specificity, kinetics, function, manufacturability, or safety requirements.
GenovaX is proposed as a research workflow that makes those boundaries explicit. It combines computational generation and prediction with provenance, uncertainty, filtering, expert review, and experimental planning. It is not a medical product and this paper does not establish a therapeutic candidate.
2. Context-of-use contract
The workflow begins by defining how the model output will be used. Examples include prioritizing a small set for an assay, exploring sequence diversity, generating a structural hypothesis, or identifying liabilities for redesign. The same model output requires different credibility when used for exploratory ranking than when used to support a regulated decision.
The contract records the biological question, target identity and version, intended and prohibited interpretations, decision owner, input data rights and provenance, model versions, acceptance criteria, uncertainty plan, experimental follow-up, and stop conditions. If the context changes, the credibility assessment must be repeated.
3. Reference workflow
The target and data layer verifies identifiers, sequences, structures, biological assumptions, and licensing or consent constraints. Candidate generation may use retrieval, mutation, language models, diffusion methods, or other design tools, but generation is separated from scoring to reduce circular evidence. Structure and interaction prediction produces ensembles where feasible rather than one preferred image.
A filtering layer evaluates interpretable liabilities such as problematic motifs, unusual cysteine patterns, charge and hydrophobicity extremes, aggregation risk indicators, and other domain-specific rules. A ranking layer combines model evidence without collapsing it into false certainty. Expert review examines biological plausibility, diversity, novelty, failure modes, and assay value. The output is an experimental handoff package, not a therapeutic claim.
4. Evidence ladder
The workflow uses an explicit evidence ladder. A generated candidate is a design artifact. A predicted structure or interaction is a computational observation. Agreement across models or seeds strengthens confidence in reproducibility of the prediction, not in the underlying biology. A controlled assay can establish a result for that assay under its conditions. Replication, orthogonal methods, functional evidence, and later regulated development add different evidence that cannot be inferred from the computational stage.
Public language should identify the rung. Terms such as designed, predicted, ranked, expressed, bound, functional, safe, and effective must not be treated as synonyms.
5. Evaluation and uncertainty
Evaluation should include holdout and out-of-distribution tasks, target leakage checks, calibration, ranking quality, structural accuracy where reference data exist, model disagreement, sensitivity to input changes, and failure analysis by target class. Benchmark selection and data overlap must be disclosed. A single aggregate result can conceal where the workflow fails.
For candidate ranking, uncertainty can include model confidence, ensemble dispersion, disagreement between structure and property methods, distance from training-like examples, rule-based liabilities, and missing biological knowledge. The ranking should preserve these components so experts can disagree with the weighting.
6. Experimental handoff
A strong handoff includes candidate identity and version, design provenance, predicted structures and confidence, ranking rationale, known liabilities, negative computational results, proposed controls, assay sequence, acceptance and stop criteria, sample and data handling requirements, and a plan for returning results to the model record.
The wet-lab plan should be capable of disproving the computational premise. Selecting only assays likely to confirm the model weakens the evidence. Negative results should remain linked to the design lineage so failed patterns are not repeatedly rediscovered.
7. Benchmark demonstration boundary
Planckchron has separately described a methodology demonstration using hen egg-white lysozyme, a standard benchmark target. The reported output is a computational shortlist produced by an AI-assisted workflow. The molecules were not represented as expressed or assayed, and lysozyme is not represented as a Planckchron therapeutic target.
That demonstration is useful for testing workflow mechanics, provenance, ranking, and disclosure. It cannot establish biological activity, translational value, or performance on a new therapeutic target. Any follow-on report should preserve that distinction and publish experimental methods and negative results if testing occurs.
8. Governance, limitations, and next gates
Access control, data governance, biosafety review, dual-use assessment, reproducible environments, model and dataset documentation, and qualified scientific review are required in proportion to the work. Any regulated use would require a separate context-specific quality and evidence program.
The next evidence gates are reproducible reruns of the computational pipeline; stronger benchmark and data-overlap analysis; independent review of the filtering and ranking logic; and, only with appropriate partners and controls, predefined experimental testing. This paper does not represent completion of those gates.

