Roadmap
This page distinguishes what is built and runnable today from what remains scaffolded or planned. It is written to be checked against the repository, not aspirational — see design.md for the assumptions this roadmap is consistent with.
MVP: complete and runnable today
- Schemas.
schemas/study.schema.jsonandschemas/evidence.schema.jsonare complete JSON Schema definitions, enforced byaree validate-studyand bysrc/harmonize. - Controlled vocabularies. Phenotype ontology (15 terms), stressor ontology (11 terms), assay types, tissue types, life stages, mapping confidence, and quality flags all exist under
registry/controlled_vocabularies/and are actively checked/consumed by code (src/common/load_vocab,src/intake/schema_validate.py). - Study registry. Six demo studies are registered (
registry/studies/GIGAS_{HEAT01,OA02,PATH03,SAL04,LARV05,GROW06}.yaml), covering the required demo mix: two RNA-seq raw-reanalysis studies (GIGAS_HEAT01,GIGAS_OA02), one methylation study (GIGAS_PATH03), one proteomics processed-results-only study (GIGAS_SAL04), one metabolomics processed-results-only study (GIGAS_GROW06), one imperfect-identifier example (GIGAS_LARV05, legacy symbol-only annotation), and one conflicting-direction example (sod1/LOC105331241acrossGIGAS_OA02andGIGAS_LARV05).registry/study_registry.csvis the resulting index.aree validate-studyandaree register-studyboth work end to end against these files. - Harmonization for all four assay types.
src/harmonizeconverts processed RNA-seq, methylation, proteomics, and metabolomics result tables into the shared evidence schema;aree harmonizeruns successfully against every demo study and writes toreports/evidence/evidence_table.tsv. - Identifier harmonization.
src/harmonize/identifiers.pyimplements the full precedence hierarchy and all sixmapping_confidencelevels, backed by a synthetic but internally consistent crosswalk (data/mappings/gene_id_crosswalk.tsv) and an ambiguous-symbol exception table (data/mappings/ambiguous_symbol_map.yaml). - Meta-analysis.
src/meta_analysis/pooling.pyimplements DerSimonian-Laird random-effects pooling with Q, I², and tau² directly in Python. Only records with a usable standard error contribute to the pooled estimate or replication count; lower-confidence aliases within one comparison are de-duplicated, and multiple contrasts from one study fail closed until a covariance-aware model or prespecified contrast is supplied.aree meta-analyzeruns against the demo evidence table and produces pooled results, including the conflicting-evidence case documented in interpreting_meta_analysis.md. - Candidate prioritization.
src/prioritize/scoring.pyandsrc/prioritize/rank.pyimplement the full transparent scoring formula and the three-tier hard-gated ranking system. Simulation status and species taxid remain partition keys through ranking and reporting; the top tier additionally requires a BH-adjusted pooled p ≤ 0.05.aree build-evidence-cardsranks every candidate intoreports/evidence_cards/candidates.tsvand renders a markdown card (plusindex.json) for each candidate with a significant signal. - CLI. All six commands documented in the README/design docs (
validate-study,register-study,list-studies,harmonize,meta-analyze,build-evidence-cards) are implemented insrc/aree/cli.pyand run against the demo data on a clean install. - Species and assembly reference data.
registry/controlled_vocabularies/species.yamlmaps scientific names and accepted synonyms (Crassostrea/Magallana gigas) onto NCBI taxids, anddata/reference/genome_assemblies.yamlcarries NCBI-verified accessions for both oyster assemblies in use. Both are enforced byaree validate-study, and evidence records now carryspecies_taxidplusidentifier_annotation_releaseso that the crossing between a study’s assembly and the annotation its identifiers were resolved against is visible — see handling_genome_versions.md. - Automated tests.
tests/covers schema validation, registry (duplicate-ID handling), identifier mapping, meta-analysis, prioritization, evidence-card generation, origin/species collision regression cases, and a full demo-pipeline integration test. - CI, interface, and documentation.
.github/workflows/ci.ymlruns Python, real-study, schema, documentation, and Nextflow smoke jobs;app/main.pyprovides the Streamlit browser;docs/_quarto.ymlrenders this site.
Scaffolded but not yet production-ready
- Nextflow workflows. All four now run
processed_results_harmonizationend to end, andworkflows/rnaseqhas additionally run the fullraw_reanalysispath against real public FASTQ (PRJNA1329250, subsampled; see first_raw_reanalysis.md). None of this was true before 2026-08-28 — three of the four did not compile on current Nextflow. Still outstanding: full-depth runs, the raw paths for methylation, proteomics and metabolomics, and any execution inside the declared containers. CI executes all four processed paths on Nextflow 26.04 and uses tiny paired FASTQ fixtures plus-stub-runto exercise the RNA-seq raw DAG; the stub run verifies wiring and artifact contracts, not tool execution or biological validity. - Containers.
containers/README.mddocuments the intended Docker/Apptainer approach, but container images/definitions themselves are not yet built. - Ortholog mapping. A real 33,356-gene M. gigas identifier crosswalk is committed and used for real studies, but its
orthogroup_idfield remains empty. Cross-species pooling still needs a real OrthoFinder/OrthoDB-derived orthology layer — see identifier_mapping.md.
Next milestones
- First full-depth real raw reanalysis. Run all prespecified
CALLA2026_OSHVcontrasts at full depth, but do not combine its repeated contrasts as independent effects. Select one comparison per study/phenotype or implement a covariance-aware within-study model first. - First real cross-study pool. Curate and reanalyze a compatible independent pathogen-challenge study such as
PRJNA593309, aligning phenotype and contrast definitions before pooling. - Container verification. Execute the CI fixtures with the declared Docker images, pin an AREE utility image, and record introspected rather than declared software versions.
- Additional species. The schema and vocabularies are species-aware and ranking/reporting partitions on
species_taxid; onboarding another species still requires its reference data and identifier crosswalk. - Expanded ontologies. Phenotype, stressor, tissue, and life-stage vocabularies cover the terms specified in the founding proposal but will need extension as more species and study types are added (e.g. finfish tissue types, additional stressor classes).
- Hosted/authenticated deployment. Explicitly out of scope for the first release per design.md — no login or cloud dependency is planned for the near term.