SimurgArray: What good could it do?

By | September 30, 2026

Part 3 of 5.

Thank you for reading this post, don't forget to subscribe!

“Benefit to humanity” is a big phrase for a tool that mostly reads binary files. I will try to keep this post the right size. SimurgArray will not cure anything. What it can do is lower a few barriers that currently decide who gets to do genetics research and how much of it they can trust.

Here is where I think it helps, roughly from the most certain to the most hopeful.

1. Reproducible science

Array genotypes feed genome wide association studies, polygenic scores, ancestry analyses and a great deal of population genetics. A lot of that work rests on a sentence like “calls were produced with the manufacturer’s software”. If someone wants to rerun the analysis five years later, with a different version of that software, the numbers may shift, and nobody can say exactly why.

With SimurgArray every run writes a manifest: software version, normalization version, cluster file checksum, command line, even which backend (Python or Rust) was used. The same inputs give the same bytes. And because the pipeline reproduces the vendor’s GTC files exactly on public data, an old analysis can be rebuilt and compared field by field, instead of “roughly”.

This is the least exciting benefit and probably the most important one.

2. Research where budgets are small

Genotyping arrays are the cheapest way to look at hundreds of thousands of variants in a person. That makes them the natural choice for biobanks and studies in countries where sequencing everyone is not realistic. Ironically, the analysis side is where those projects often struggle: licenses, servers, specialist staff, and software that assumes all three.

SimurgArray runs on a Raspberry Pi. I do not recommend running a national biobank on one (I did find its limits, usually around midnight), but a single ordinary server core handles about 80 GSA samples a minute. A modest machine can process a large study in an afternoon, and the software costs nothing.

The pipeline also produces the things a small team needs to trust its data without a bioinformatics department: control probe QC, sample QC, contamination estimates, relatedness, sex checks, and an HTML report that a lab person can open in a browser and actually understand.

3. Better calls for underrepresented populations

This is the part I find most hopeful.

A cluster file encodes where the genotype clouds sit for each SNP. It is trained on a set of samples, and those samples have ancestries. If your cohort comes from a population that was thin or absent in that training set, rare alleles that are common in your population may never have had a proper cluster. The caller then does its best with an imputed guess.

ArrayTrain lets a group train cluster files on their own samples. On public data, a cluster file trained from 120 arrays already matched the vendor’s accuracy against whole genome sequencing within 0.02 percentage points, while calling more genotypes. I have only tested one cohort so far, so I will not claim more than that. But the direction is clear: the people who own the samples can own the clusters too.

4. Pharmacogenomics research that is open to inspection

A growing number of arrays include pharmacogenomic content, because a handful of genes (CYP2D6, CYP2C19, CYP2C9, SLCO1B1, DPYD and others) influence how people respond to common drugs. Calling star alleles from an array is tricky. CYP2D6 in particular has deletions, duplications and hybrid genes that a SNP array can only see indirectly.

SimurgArray builds its PGx database from open sources (CPIC, PharmVar, ClinPGx), calls targeted copy number, and reports diplotypes together with everything that went into them: which variants were seen, which were missing, which alternative diplotypes the array cannot tell apart. On public reference samples it agreed with the GeT RM consensus in 154 of 168 comparable gene calls, and the misses are listed and explained.

I want to be very clear here. This is research software. It is not for choosing a drug or a dose. But openness matters even for research: when an algorithm says someone is a poor metabolizer, a researcher should be able to see why, and disagree.

5. Copy number for rare disease research

Large deletions and duplications cause many rare diseases, and arrays are still a common first look. SimurgArray calls CNVs, loss of heterozygosity and mosaic events across the genome and writes them with ISCN style annotation.

Its honest numbers: on 360 public arrays compared with 30x sequencing, constitutional deletion calls were right far more often than mosaic ones, and rare deletions were found about 30 % of the time at the size and probe density we tested. Common CNVs are mostly invisible to the method by design, because the reference batch absorbs them. For rare disease research, rare events are the interesting ones, so that is the right trade, but it is a trade and it is written down.

6. Teaching

I learned more about arrays by rebuilding the pipeline than from any course. The code is written to be read. Every algorithm names its source. Every oddity the vendor has, and every place we deliberately behave differently, is documented with a date and a reason.

A student who wants to know how a genotype call happens can follow one sample through normalization, clustering and scoring, print intermediate values, and compare them with the reference tools. That used to require either a job at the company or a lot of guessing.

7. Keeping the ecosystem honest

Open alternatives are good for everyone, including the vendor. When a public tool can reproduce outputs exactly, differences become bug reports instead of mysteries. Several findings in known_unknowns.md are small documentation errors or version quirks that anyone using the official tools would benefit from knowing about.

What it will not do

It will not replace clinical laboratories, their validated pipelines, or the people who sign reports. It will not make arrays as informative as sequencing. It will not replace the excellent open tools it builds on. And it will not make genetic data less sensitive: people who run it are responsible for handling their data properly, which the software supports with privacy options in its reports but cannot enforce.

The small version of a big claim

If SimurgArray helps one underfunded study get cleaner genotypes, one student understand what a cluster file is, or one reviewer ask “which version of the normalization did you use?”, it has done its job. The bird in its name is known for bringing many small travellers to one clear place. I will be happy if it brings a few people to a clearer view of their data.

Next time: what it would take to make SimurgArray public properly, and a realistic calendar for doing it.

Leave a Reply

Your email address will not be published. Required fields are marked *