Are there SimurgArray alternatives? Yes. Why build it anyway?

By | September 30, 2026

Part 2 of 5.

Thank you for reading this post, don't forget to subscribe!

Whenever someone tells me they are writing their own genotype caller, my first instinct is to ask whether they have looked around. So it is only fair that I answer the same question about myself.

I did look around. There is a lot of good work out there. Some of it I admire, some of it I depend on, and some of it is the reason SimurgArray can check its own answers at all.

What the vendor gives you

The array manufacturer offers its own software, and it is good at what it does.

GenomeStudio is the classic. A desktop application for Windows where you load a project, look at cluster plots, fix the loci that look strange and export reports. Generations of lab scientists learned arrays on it. It is free to use after registering, but it is closed, it wants a graphical desktop, and it does not scale nicely to tens of thousands of samples.

The command line tools (the array analysis platform and its predecessors) turn IDATs into GTC files without the graphical part. They are fast and they are the reference: when I say SimurgArray matches the vendor byte for byte, these are the outputs I mean. They are also closed, so when something looks odd you can only guess why.

The vendor’s newer analysis software does genotyping, copy number and pharmacogenomics in one package, locally or in the cloud. I studied its public documentation closely, because it describes a sensible feature set and a clear list of outputs. It is still closed, and its licensing and platform assumptions are aimed at institutions, not at a person with a Raspberry Pi and a free evening.

None of this is a complaint. A company building good closed tools for paying customers is a normal thing. It just leaves a gap.

What the open source world gives you

This is where it gets interesting, because several people have already done a lot of the hard work in the open.

gtc2vcf by Giulio Genovese is a bcftools plugin that converts GTC files (and IDATs, via an open reimplementation of the calling) into VCF. It is careful, fast and MIT licensed. SimurgArray uses it as an oracle all the time, and it helped me understand several details that the public documentation leaves out. If you only need GTC to VCF, you may not need anything else.

MoChA, from the same author, detects mosaic chromosomal alterations using phased BAF. It is the right tool for mosaicism in large cohorts, and it is the reason I know my own mosaic caller has a hard limit below about 20 % cell fraction without phasing.

BeadArrayFiles and GTCtoVCF, both from the vendor’s own open source repositories, read the file formats and write VCF. Old, a bit dusty (GTCtoVCF needed two small patches to run on Python 3), and extremely useful as references.

illuminaio in Bioconductor reads IDAT files. It is GPL licensed, so SimurgArray only runs it as a separate program to compare results and never reuses its code.

crlmm and Illuminus are academic callers from the late 2000s with a real statistical model behind them. Their papers shaped how ArrayTrain thinks about clusters.

PennCNV and QuantiSNP are the long standing tools for copy number from LRR and BAF.

PharmCAT, Stargazer, Aldy and friends call star alleles, mostly from sequencing data or from VCFs you already have.

sesame and minfi handle methylation arrays in R, and SimurgArray’s methylation QC is checked against sesame.

PLINK 2 does everything downstream of a VCF, better than I ever would, and I would not dream of replacing it.

So why build another one?

Because none of these, alone or together, gave me the thing I wanted: one open pipeline from raw IDAT to final report, bit exact with the vendor where that is possible, tested against every one of the tools above, and honest about where it differs.

Here are the specific reasons, in the order they occurred to me.

The chain was broken in the middle. You can read IDATs with one tool, convert GTCs with another, call CNVs with a third and PGx with a fourth. Each hands off to the next with its own conventions for strand, allele order, missing values and sample names. SimurgArray keeps one data model from start to finish, and every hand off is tested.

Nobody promised to agree with the vendor to the last bit. Close is usually fine. But close means you cannot tell a real difference from a rounding difference. Once normalization, calling and GTC writing are identical down to the byte, any remaining difference is information instead of noise. Getting there required reproducing float32 quirks, a particular median that is not quite a median, and the sign of a NaN. I am unreasonably proud of that last one.

Cluster files were a dependency without an exit. Every open caller I found needs a cluster file from the vendor. ArrayTrain removes that requirement. For a cohort whose ancestry is not well represented in the samples the vendor trained on, that might matter a lot. I will come back to this.

I wanted it to run on small hardware. Most pipelines assume a cluster or at least a workstation. SimurgArray was written on a Raspberry Pi and validated there as much as possible, with the cloud used only when a dataset simply did not fit, and even then for minutes at a time.

I wanted a specification, not just code. The functional spec lists 416 behaviours with IDs, from command line options to edge cases like an empty sample name column in the vendor’s own template sheet. Tests cite those IDs. It sounds bureaucratic. It is actually liberating, because “is this a bug?” becomes a question with an answer.

I wanted to learn it properly. Reading code teaches you how something works. Reproducing it teaches you why. Several of the most interesting facts about arrays that I know now came from a test failing by one unit in the last place.

And the stubbornness

I should be honest: there is also a bit of “how hard can it be”. The answer, as usual, was “harder than that, but not impossible”. The first estimate for normalization was a few weeks. The first version that matched the vendor bit for bit took longer, and taught me more about floating point contraction on ARM processors than I had planned to know.

It helped that the expired patents on the core normalization and clustering methods are public, that the technical notes are public, and that people like the authors above wrote their work down in the open. SimurgArray stands on all of that. It does not replace any of it.

Where it fits

If you want a mature graphical tool with support, use the vendor’s software. If you only need GTC to VCF, use gtc2vcf. If you want mosaicism in a biobank, use MoChA. If you want PGx from sequencing, use PharmCAT or Aldy.

If you want to go from a folder of IDAT files to genotypes, QC, CNVs, PGx and your own cluster file, on hardware you own, with every step open and every disagreement written down, SimurgArray is the tool I wish I had found when I started.

Next time: why I think this is useful beyond my own curiosity, and who might actually benefit.

Leave a Reply

Your email address will not be published. Required fields are marked *