Skip to content

Archive historical research and add canonical training recipes #83

Description

@gonzalobenegas

Outcome

Make the maintenance boundary explicit while preserving research reproducibility.

Acceptance criteria

  • Extract one canonical GPN training recipe and one canonical GPN-Star training recipe, each using an already-prepared dataset, a tiny CPU smoke config, and a realistic example GPU config.
  • Do not maintain exact executable profiles for every historical paper run and do not maintain dataset construction.
  • Create a reviewed manifest of the historical material, then prepare an annotated immutable archive tag and lightweight Release.
  • After archive approval, remove analysis/, top-level dataset-building workflow/, GPN-MSA training code, and all notebooks except the existing GPN/GPN-Star/PhyloGPN quick starts from main.
  • Preserve GPN-MSA inference, Sorghum expression inference, and PhyloGPN inference.
  • Add docs/research.md with non-binding guidance for long-lived off-main exploratory/paper branches, self-contained environments, and pinned GPN versions/commits.

The archive tag/Release is prepared in the modernization stack but is not published until final maintainer approval.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions