In 2003, scientists reached a milestone with the Human Genome Project that changed biology by creating the first reference human genome—essentially the instruction book for building and operating the human body. It was not a single individual’s genome but a carefully assembled reference that gave researchers a common starting point for studying DNA. For the first time, scientists had a reliable guide they could use to compare newly sequenced genomes and find the genetic differences that make each of us unique.
Some of these differences involve a change in just one DNA letter. Called single nucleotide polymorphisms (SNPs), these changes help explain everything from eye color and blood type to why some people are more prone to certain diseases. Because SNPs are used to study disease, evolution, and genetic diversity, their accurate identification is one of the most important steps in modern genomics.
However, finding SNPs is not as simple as comparing two DNA sequences letter by letter. Modern sequencing technologies generate millions of short DNA fragments that must be sequenced together and compared to a reference genome using sophisticated computer software. Different programs can analyze the same sequence data in different ways, sometimes disagreeing about whether a SNP is actually present. In a new study published in Genetics, Jackson Laboratory researchers set out to determine which of these commonly used tools accurately identify SNPs in laboratory mice, helping scientists make more confident genetic discoveries.
Comparing different calling programs presents a unique challenge because there is no easy way to know which one is right. If researchers analyze DNA from a real mouse, they don’t know the exact location of each real SNP ahead of time. Without this answer key, it is impossible to compare one program to another.
To overcome this problem, researchers created synthetic mouse genomes. They started with the genomes of 10 well-studied inbred laboratory mouse strains and inserted known SNPs into them. Because they knew exactly where each genetic variant was placed, they could objectively measure how accurately each software program identified those variants while avoiding false positives.
The team wanted to know whether software designed for genetically diverse populations would perform well in these highly homogeneous genomes. Many of today’s most popular variant calling programs were originally developed and tested using the human genome, where individuals inherit different DNA from each parent. Laboratory mouse strains are different. After generations of breeding, individuals within a strain are nearly genetically identical.
The results showed that no single variant calling program consistently outperformed the others.
Each tool had advantages and disadvantages. Some programs identified true SNPs but also reported more false positives, while others were more conservative and missed true variants. Surprisingly, the most reliable results came from combining multiple programs rather than relying on a single program. By searching for SNPs identified by more than one tool, the researchers were able to improve confidence that the variants were real. They also found that variant calling became increasingly difficult as mouse strains became more genetically divergent from the reference genome, highlighting the importance of computational tools used to optimize and analyze reference genomes.
While this study focused on laboratory mice, its implications extend far beyond a model organism. Every genomic discovery begins with the accurate identification of genetic variants. Before researchers can uncover genes linked to disease, reconstruct evolutionary history, or understand the genetic basis of complex traits, they must first know that the species they are studying are real. As sequencing technologies and reference genomes continue to improve, studies like this help ensure that the computational tools used to interpret genomic data keep pace. By providing practical guidance for benchmarking and combining different calling methods, this work strengthens the foundation upon which future discoveries in genetics can be built.
References
-
Garretson, A., Blanco-Berdugo, L., MG Roberts, A., et al. “Benchmarking Genomic Variant Calling Tools in Inbred Mouse Strains: Recommendations and Considerations,” Genetics(2026): 223–3. https://doi.org/10.1093/genetics/iyag131.






