SimurgArray going public: what it takes, and a calendar

By | September 30, 2026

Part 4 of 5.

Thank you for reading this post, don't forget to subscribe!

Right now SimurgArray lives in a private repository on my own Git server. It has a license, a changelog, a user guide, wheels, a Docker image and a version number that just became 2.0.0. In other words, it looks public. It is not, yet.

Publishing code is easy. Publishing it responsibly is a checklist. This post is that checklist, followed by a calendar I intend to keep, give or take the usual week.

What “public” means here

By public I mean four things:

  1. The source code on a public host, with issues and pull requests open.
  2. Installable packages: a Python wheel, the optional Rust kernel wheels, a container image and a conda recipe.
  3. Documentation someone can follow without asking me anything.
  4. A citable record: a DOI for each release and a short preprint describing the methods and the validation.

The checklist

Licensing

SimurgArray is MIT licensed. Code borrowed from other projects came only from MIT, BSD and Apache licensed sources, and their license texts sit in THIRD_PARTY_NOTICES.md exactly as the authors wrote them. GPL tools are only ever run as separate programs for comparison, never copied.

Before going public I will do one more pass over every file to confirm that each ported function names its origin, and that nothing slipped in during the late night sessions. Late night sessions are when things slip in.

Data that must never be published

The repository must not contain vendor manifests, cluster files, IDATs or GTCs. The .gitignore has blocked them from the start, and the test fixtures are Bioconductor example files under the Artistic 2.0 license plus synthetic data. Still, “we never committed it” is a claim, and a claim deserves a check. I will scan the complete Git history, not just the current tree, for large binaries and for file types that should not be there.

The same goes for secrets. The cloud scripts use credentials from the local environment, never from files in the repository, but a full history scan for keys is cheap and embarrassment is expensive. The cloud account credentials used during development will be rotated before the switch, as a matter of principle.

Names and trademarks

The software reads file formats defined by Illumina and reproduces the behaviour of Illumina’s algorithms from public documents and expired patents. Illumina, Infinium, GenCall, GenTrain and GenomeStudio are their trademarks. The documentation will say clearly that SimurgArray is independent and not affiliated with or endorsed by the vendor, and it will use those names only to describe compatibility.

The name SimurgArray itself needs a check on the Python package index, conda channels and container registries before anything is uploaded. If someone else has it, I will be sad for a day and then pick a variant.

Third party data sources

The PGx database is built on the user’s machine from public sources:

  • CPIC (CC0),
  • PharmVar (their own terms),
  • ClinPGx (CC BY SA).

SimurgArray ships the builder, not the database, so each source keeps its own terms and attribution. The documentation will list those terms next to the download commands, so nobody has to find them in a footnote.

Scope and safety language

Every README, report header and PGx output will state that the software is for research, not for diagnosis or treatment decisions. The PGx JSON already carries the database version and the evidence behind each call. That transparency is the main safety feature, but it is not a substitute for saying the obvious out loud.

Engineering hygiene

  • Continuous integration on public runners for x86_64 and aarch64: lint, type checks, unit tests.
  • Oracle tests that need large public datasets stay opt in, with a script that downloads everything with checksums.
  • Issue templates that ask for the command, the software version and the run manifest, because “it gave a wrong genotype” is hard to debug without them.
  • A contributing guide, a code of conduct and a security policy.
  • A short description of the release process, so that it is not only in my head.

Documentation

The user guide covers every command. What is missing is a gentle path for newcomers: a tutorial that starts from public IDATs (the Bioconductor fixtures and the GEO dataset used in validation), runs the whole pipeline on a laptop, and explains each output. I also want one page per validation report written for people who do not want to read a JSON file. Most people.

Citation

Each tagged release will get a DOI through Zenodo, and the repository will carry a citation file. A preprint on bioRxiv will describe the design, the bit exact reproduction, ArrayTrain, the CN and PGx modules, and every known limitation. If it is going to be cited, it should be cited for what it actually does.

The calendar

The weeks start on Mondays. Everything below costs nothing except time, apart from an optional domain for the documentation site.

Week of October 5, 2026: audit
 
 Full history scan for vendor files and secrets. Credential rotation. Provenance pass over ported code. Trademark and disclaimer text in README, user guide and report templates.

Week of October 12: public repository
 
 Public mirror on a mainstream code host, continuous integration for both architectures, issue templates, contributing guide, code of conduct, security policy. The private Git server stays as my working copy.

Week of October 19: packages
 
 Check name availability. Upload version 2.0.0 wheels to the Python package index, the image to a public container registry, and submit the conda recipe for review.

Weeks of October 26 and November 2: tutorial and docs site
 
 End to end tutorial on public data, validation summaries rewritten for humans, documentation site generated from the repository.

Week of November 16: soft launch
 
 Announce to a few people who work with arrays and will not be polite about problems. Collect issues. Fix the embarrassing ones quietly and the interesting ones loudly.

Weeks of November 30 and December 14: preprint
 
 Write the methods paper, rerun every validation from a clean checkout so the numbers in the paper come from the public code, deposit data summaries, submit to bioRxiv.

Week of January 11, 2027: version 2.1
 
 First release shaped by other people’s feedback. Likely candidates: a second cohort for ArrayTrain validation, iterative CN references for better rare event sensitivity, and anything the soft launch reveals.

After that, a steady rhythm: a minor release every two or three months, a DOI for each, and a public changelog that keeps the habit of writing down what changed and why.

What could delay it

  • A name clash. Solvable, mildly annoying.
  • An unexpected license question about a source. Solvable, possibly slow.
  • Real life. Unavoidable, frequently scheduled.

Why announce a calendar at all

Because a public date is the best cure for “just one more feature”. The software is already more complete than I planned when I started. What it needs now is other people using it, finding its mistakes and telling me which of its choices were wrong.

Next time, the last post: what using SimurgArray actually looks like, from a folder of IDATs to a report, and what the name and the bird are about.

Leave a Reply

Your email address will not be published. Required fields are marked *