Capability requirements

Diagnostic data is necessary but not sufficient. The system must interpret that data in the context of the patient's complete genome and known variants, current epigenetic state, tissue-specific and cell-type-specific norms, medical history and current trajectory, likely causal mechanisms underlying observed pathology, and predicted response to candidate interventions.

This is substantially harder than diagnosis because it requires causal reasoning. Pattern recognition tells you that the patient is sick. Interpretation tells you why and what to do.

Current state of the science

Genome sequencing. Whole-genome sequencing now costs under $200 and can be completed in hours. Long-read sequencing (PacBio, Oxford Nanopore) resolves structural variants and complex regions earlier methods missed. We can effectively read the genome.

Variant interpretation. The gap between reading the genome and understanding it is enormous. The American College of Medical Genetics classifies most variants as "variants of uncertain significance" (VUS) — we can see them but don't know what they do. Of the roughly 3 million variants in an average genome, only a small fraction have well-characterised effects.

Functional genomics. Massively parallel reporter assays, CRISPR screens, and saturation mutagenesis are systematically mapping the function of genomic regions and variants. Projects like ENCODE and the Human Cell Atlas are building reference maps of regulatory elements and cell types.

Pharmacogenomics. Specific gene variants affecting drug metabolism (CYP450 family, TPMT, others) are now part of standard care for some prescriptions. This is the leading edge of personalised treatment based on genome.

Proteomics and metabolomics. Mass spectrometry can identify and quantify thousands of proteins and metabolites from small samples. The proteome is more directly relevant to disease state than the genome but harder to measure comprehensively in vivo.

Foundation models for biology. Models like AlphaFold (protein structure), ESM (protein function), and emerging single-cell foundation models are predicting biological behaviour from sequence at unprecedented accuracy. Predictive models of cellular response to perturbation are an active research frontier.

Required advances

Bucket assessment

Interpretation is harder than diagnosis. The genome can be read, but understanding it remains far from complete. Two specific problems — causal disease modelling and rare variant interpretation — may not yield to expected research trajectories and may require fundamental scientific advances comparable to the protein folding breakthrough achieved by AlphaFold.

Without those advances, AIHS interpretation will be limited to well-characterised conditions and known intervention targets — which still covers a substantial fraction of medical need but falls short of the universal capability implied by the AIHS concept.

The dependency on Bucket A is significant. Real-time epigenomic and cellular state inference is impossible without the diagnostic advances above. This study recommends Bucket A and Bucket B research be tightly coupled rather than pursued in parallel silos.