{"id":4471,"date":"2025-07-12T14:53:40","date_gmt":"2025-07-12T14:53:40","guid":{"rendered":"https:\/\/www.newsbeep.com\/us\/4471\/"},"modified":"2025-07-12T14:53:40","modified_gmt":"2025-07-12T14:53:40","slug":"mkdesigner-and-taseq-a-set-of-tools-for-plant-genotyping-by-targeted-amplicon-sequencing-bmc-bioinformatics","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/us\/4471\/","title":{"rendered":"MKDESIGNER and TASEQ: a set of tools for plant genotyping by targeted amplicon sequencing | BMC Bioinformatics"},"content":{"rendered":"<p>DNA-based genetic markers have been used in a wide range of biological fields, including medicine, genetics, and molecular biology. Genotyping is essential for these research, and various methods have been devised [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 13\" title=\"Davey JW, Hohenlohe PA, Etter PD, Boone JQ, Catchen JM, Blaxter ML. Genome-wide genetic marker discovery and genotyping using next-generation sequencing. Nat Rev Genet. 2011;12(7):499\u2013510. &#010;                  https:\/\/doi.org\/10.1038\/nrg3012&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR13\" id=\"ref-link-section-d245668886e477\" rel=\"nofollow noopener\" target=\"_blank\">13<\/a>, <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 25\" title=\"Liu ZJ, Cordes JF. DNA marker technologies and their applications in aquaculture genetics. Aquaculture. 2004;238(1\u20134):1\u201337. &#010;                  https:\/\/doi.org\/10.1016\/j.aquaculture.2004.05.027&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR25\" id=\"ref-link-section-d245668886e480\" rel=\"nofollow noopener\" target=\"_blank\">25<\/a>].<\/p>\n<p>The first generation of genotyping was using restriction fragment length polymorphism (RFLP) or random amplified polymorphic DNA (RAPD) markers [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 1\" title=\"Agarwal M, Shrivastava N, Padh H. Advances in molecular marker techniques and their applications in plant sciences. In: Plant cell reports, vol 27, no 4, pp 617\u2013631; 2008. &#010;                  https:\/\/doi.org\/10.1007\/s00299-008-0507-z&#010;                  &#010;                \" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR1\" id=\"ref-link-section-d245668886e486\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>]. These markers are not frequently in the genome, and the experiments were laborious and time-consuming. In the next stage, genotyping by PCR and electrophoresis using simple sequence repeat (SSR) markers has been commonly performed since 1990\u2019s [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 18\" title=\"Eujayl I, Sorrells ME, Baum M, Wolters P, Powell W. Isolation of EST-derived microsatellite markers for genotyping the A and B genomes of wheat. Theor Appl Genet. 2002;104(2):399\u2013407. &#010;                  https:\/\/doi.org\/10.1007\/s001220100738&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR18\" id=\"ref-link-section-d245668886e489\" rel=\"nofollow noopener\" target=\"_blank\">18<\/a>, <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 26\" title=\"McCouch SR. Development and mapping of 2240 new SSR markers for rice (Oryza sativa L.). DNA Res. 2002;9(6):199\u2013207. &#010;                  https:\/\/doi.org\/10.1093\/dnares\/9.6.199&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR26\" id=\"ref-link-section-d245668886e492\" rel=\"nofollow noopener\" target=\"_blank\">26<\/a>, <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 31\" title=\"Smith JSC, Chin ECL, Shu H, Smith OS, Wall SJ, Senior ML, Mitchell SE, Kresovich S, Ziegle J. An evaluation of the utility of SSR loci as molecular markers in maize (Zea mays L.): comparisons with data from RFLPS and pedigree. Theor Appl Genet. 1997;95(1\u20132):163\u201373. &#010;                  https:\/\/doi.org\/10.1007\/s001220050544&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR31\" id=\"ref-link-section-d245668886e495\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a>]. Although it has led to many genetic achievements, it is now considered labor intensive in the age of next-generation sequencing (NGS). High-throughput NGS genotyping methods are now available. Because whole-genome sequencing for all genetic analysis populations is expensive, it is common to produce a reduced representation of the genome before sequencing [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 13\" title=\"Davey JW, Hohenlohe PA, Etter PD, Boone JQ, Catchen JM, Blaxter ML. Genome-wide genetic marker discovery and genotyping using next-generation sequencing. Nat Rev Genet. 2011;12(7):499\u2013510. &#010;                  https:\/\/doi.org\/10.1038\/nrg3012&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR13\" id=\"ref-link-section-d245668886e498\" rel=\"nofollow noopener\" target=\"_blank\">13<\/a>]. There are several methods to produce a reduced representation of the genome, such as RAD-seq [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"Baird NA, Etter PD, Atwood TS, Currey MC, Shiver AL, Lewis ZA, Selker EU, Cresko WA, Johnson EA. Rapid SNP discovery and genetic mapping using sequenced RAD markers. PLoS ONE. 2008;3(10): e3376. &#010;                  https:\/\/doi.org\/10.1371\/journal.pone.0003376&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR3\" id=\"ref-link-section-d245668886e502\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>], GBS [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 16\" title=\"Elshire RJ, Glaubitz JC, Sun Q, Poland JA, Kawamoto K, Buckler ES, Mitchell SE. A robust, simple genotyping-by-sequencing (GBS) approach for high diversity species. PLoS ONE. 2011;6(5): e19379. &#010;                  https:\/\/doi.org\/10.1371\/journal.pone.0019379&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR16\" id=\"ref-link-section-d245668886e505\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a>], MIG-seq [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 32\" title=\"Suyama Y, Matsuki Y. MIG-seq: an effective PCR-based method for genome-wide single-nucleotide polymorphism genotyping using the next-generation sequencing platform. Sci Rep. 2015;5(1):16963. &#010;                  https:\/\/doi.org\/10.1038\/srep16963&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR32\" id=\"ref-link-section-d245668886e508\" rel=\"nofollow noopener\" target=\"_blank\">32<\/a>], and GRAS-Di\u00ae [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 17\" title=\"Enoki H, Takeuchi Y. New genotyping technology, GRAS-Di, using next generation sequencer. In: Proceedings of the plant and animal genome conference XXVI, San Diego, CA, USA, 13\u201317 January 2018. Available online: &#010;                  https:\/\/pag.confex.com\/pag\/xxvi\/meetingapp.cgi\/Paper\/29067&#010;                  &#010;                . Accessed on 10 Oct 2024\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR17\" id=\"ref-link-section-d245668886e513\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>]. However, these methods do not allow prior knowledge of marker locations. SNP genotyping assay using microarray technology can solve this problem [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"Gunderson KL, Steemers FJ, Lee G, Mendoza LG, Chee MS. A genome-wide scalable SNP genotyping assay using microarray technology. Nat Genet. 2005;37(5):549\u201354. &#010;                  https:\/\/doi.org\/10.1038\/ng1547&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR21\" id=\"ref-link-section-d245668886e516\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a>]. It is commercially available [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 2\" title=\"Akhunov E, Nicolet C, Dvorak J. Single nucleotide polymorphism genotyping in polyploid wheat with the Illumina GoldenGate assay. Theor Appl Genet. 2009;119(3):507\u201317. &#010;                  https:\/\/doi.org\/10.1007\/s00122-009-1059-5&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR2\" id=\"ref-link-section-d245668886e520\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>], but the cost of design novel SNP array is high compared to the NGS-based methods mentioned above. Targeted amplicon sequencing (TAS) is also a promising method. For genotyping by TAS, the sequences around the targeted polymorphisms are amplified by multiplex PCR, and the amplicons are sequenced via next-generation sequencing (NGS). Therefore, markers can be designed at desired positions. The concept of \u200b\u200bTAS was first reported more than a decade ago [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 6\" title=\"Bybee SM, Bracken-Grissom H, Haynes BD, Hermansen RA, Byers RL, Clement MJ, Udall JA, Wilcox ER, Crandall KA. Targeted amplicon sequencing (TAS): a scalable next-gen approach to Multilocus Multitaxa phylogenetics. Genome Biol Evol. 2011;3(1):1312\u201323. &#010;                  https:\/\/doi.org\/10.1093\/gbe\/evr106&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR6\" id=\"ref-link-section-d245668886e523\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>]. Similar methods have been reported under different names, such as Genotyping-in-Thousands by sequencing (GT-seq) [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 9\" title=\"Campbell NR, Harmon SA, Narum SR. Genotyping-in-Thousands by sequencing (GT-seq): a cost effective SNP genotyping method based on custom amplicon sequencing. Mol Ecol Resour. 2015;15(4):855\u201367. &#010;                  https:\/\/doi.org\/10.1111\/1755-0998.12357&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR9\" id=\"ref-link-section-d245668886e526\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>] and Multiplex PCR Targeted Amplicon Sequencing (MTA-Seq) [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 29\" title=\"Onda Y, Takahagi K, Shimizu M, Inoue K, Mochida K. Multiplex PCR targeted amplicon sequencing (MTA-Seq): simple, flexible, and versatile SNP genotyping by highly multiplexed PCR amplicon sequencing. Front Plant Sci. 2018. &#010;                  https:\/\/doi.org\/10.3389\/fpls.2018.00201&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR29\" id=\"ref-link-section-d245668886e529\" rel=\"nofollow noopener\" target=\"_blank\">29<\/a>]. It is also provided commercially as Ion AmpliSeq technology (Thermo Fisher Scientific, Inc.) [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 30\" title=\"Sato M, Hosoya S, Yoshikawa S, Ohki S, Kobayashi Y, Itou T, Kikuchi K. A highly flexible and repeatable genotyping method for aquaculture studies based on target amplicon sequencing using next-generation sequencing technology. Sci Rep. 2019;9(1):6904. &#010;                  https:\/\/doi.org\/10.1038\/s41598-019-43336-x&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR30\" id=\"ref-link-section-d245668886e532\" rel=\"nofollow noopener\" target=\"_blank\">30<\/a>].<\/p>\n<p>In crop science, bi-parental quantitative trait loci (QTL) analyses are often conducted to identify genomic regions related to agronomically important traits. For this application, it is important to ensure enough markers across the genome while keeping costs as low as possible. TAS has a potential to meet this requirement. However, TAS is currently not frequently used for this application for several reasons. First, a reference genome is required for custom primer design. Second, it is expensive to purchase hundreds or thousands of primers. Third, it requires advanced knowledge and skills in bioinformatics. In recent years, the first problem is not so critical because reference genomes have been established for many crop species [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 34\" title=\"Xie L, Gong X, Yang K, Huang Y, Zhang S, Shen L, Sun Y, Wu D, Ye C, Zhu QH, Fan L. Technology-enabled great leap in deciphering plant genomes. In Nature plants, vol 10, no 4, pp. 551\u2013566. Nature Research; 2024. &#010;                  https:\/\/doi.org\/10.1038\/s41477-024-01655-6&#010;                  &#010;                \" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR34\" id=\"ref-link-section-d245668886e538\" rel=\"nofollow noopener\" target=\"_blank\">34<\/a>]. Natsume et al. [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 28\" title=\"Natsume S, Oikawa K, Nomura C, Ito K, Utsushi H, Shimizu M, Terauchi R, Abe A. V-primer: software for the efficient design of genome-wide InDel and SNP markers from multi-sample variant call format (VCF) genotyping data. Breed Sci. 2023;73(4):23018. &#010;                  https:\/\/doi.org\/10.1270\/jsbbs.23018&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR28\" id=\"ref-link-section-d245668886e541\" rel=\"nofollow noopener\" target=\"_blank\">28<\/a>] showed that the second and third problems could be overcome by using cheaper low-concentration mixed primers and developing a free software named V-primer.<\/p>\n<p>The workflow of TAS is roughly divided into three steps (Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>). First, DNA markers are extracted from the NGS data of the parental lines, and primers are designed. Second, the NGS library is constructed using the designed primers, and sequencing is performed. Third, the data obtained by NGS are analyzed, and the genotypes at each DNA marker of each line are determined. The first and third steps require specialized bioinformatics analysis.<\/p>\n<p>Fig.\u00a01<a class=\"c-article-section__figure-link\" data-test=\"img-link\" data-track=\"click\" data-track-label=\"image\" data-track-action=\"view figure\" href=\"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3\/figures\/1\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" aria-describedby=\"Fig1\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/07\/12859_2025_6211_Fig1_HTML.png\" alt=\"figure 1\" loading=\"lazy\" width=\"685\" height=\"329\"\/><\/a><\/p>\n<p>Schematic diagram of genotyping by targeted amplicon sequencing workflow<\/p>\n<p>V-primer [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 28\" title=\"Natsume S, Oikawa K, Nomura C, Ito K, Utsushi H, Shimizu M, Terauchi R, Abe A. V-primer: software for the efficient design of genome-wide InDel and SNP markers from multi-sample variant call format (VCF) genotyping data. Breed Sci. 2023;73(4):23018. &#010;                  https:\/\/doi.org\/10.1270\/jsbbs.23018&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR28\" id=\"ref-link-section-d245668886e573\" rel=\"nofollow noopener\" target=\"_blank\">28<\/a>] can perform the first step, DNA marker extraction and high-throughput primer design, at once. However, no integrated pipelines have been developed to handle the third step, analysis after sequencing. There is also room for improvement in primer design strategies. To make TAS more generally available, it is necessary to establish user-friendly pipelines that covers whole the experiments.<\/p>\n<p>In this study, we developed a novel genome-wide primer design tool, MKDESIGNER, which implements specialized functions for genotyping by TAS. We also developed the post-sequencing analysis tool TASEQ. This tool is a pipeline that takes FASTQ files as input and outputs genotype files in a format that can be used directly for QTL analysis. By implementing these tools, genotyping by TAS for populations which are derived from cross among two to several fixed parental lines can be done easily.<\/p>\n<p>In this paper, we explain how MKDESIGNER and TASEQ work and provide a practical example of genotyping by TAS in rice using these tools.<\/p>\n<p>ImplementationOverview of the tool<\/p>\n<p>MKDESIGNER and TASEQ are command line interface tools that have been verified to work on Ubuntu 20.04 and later. The source codes are written in Python and are available at GitHub (<a href=\"https:\/\/github.com\/KChigira\/mkdesigner\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/KChigira\/mkdesigner<\/a>, <a href=\"https:\/\/github.com\/KChigira\/taseq\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/KChigira\/taseq<\/a>)<\/p>\n<p>They can be installed via Bioconda, including their dependencies. The workflow is shown in Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>. MKDESIGNER has three commands, and TASEQ has four commands. The role of each command and the external tools used are described in the following sections.<\/p>\n<p>Fig.\u00a02<a class=\"c-article-section__figure-link\" data-test=\"img-link\" data-track=\"click\" data-track-label=\"image\" data-track-action=\"view figure\" href=\"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3\/figures\/2\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" aria-describedby=\"Fig2\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/07\/12859_2025_6211_Fig2_HTML.png\" alt=\"figure 2\" loading=\"lazy\" width=\"685\" height=\"329\"\/><\/a><\/p>\n<p>Workflow of genotyping by TAS with automated data analysis using MKDESIGNER and TASEQ<\/p>\n<p>\u2018mkvcf\u2019<\/p>\n<p>This command is responsible for creating a VCF file from the NGS data of the parent varieties. It requires BAM-formatted NGS data from two or more parental lines and a FASTA-formatted reference genome as input. This command produces a VCF file using GATK HaplotypeCaller [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 27\" title=\"McKenna A, Hanna M, Banks E, Sivachenko A, Cibulskis K, Kernytsky A, Garimella K, Altshuler D, Gabriel S, Daly M, DePristo MA. The genome analysis toolkit: a mapreduce framework for analyzing next-generation DNA sequencing data. Genome Res. 2010;20(9):1297\u2013303. &#010;                  https:\/\/doi.org\/10.1101\/gr.107524.110&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR27\" id=\"ref-link-section-d245668886e638\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>] and BCFtools [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 12\" title=\"Danecek P, Bonfield JK, Liddle J, Marshall J, Ohan V, Pollard MO, Whitwham A, Keane T, McCarthy SA, Davies RM, Li H. Twelve years of SAMtools and BCFtools. GigaScience. 2021. &#010;                  https:\/\/doi.org\/10.1093\/gigascience\/giab008&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR12\" id=\"ref-link-section-d245668886e641\" rel=\"nofollow noopener\" target=\"_blank\">12<\/a>]. Pre-created VCF files can also be used for subsequent analysis, but they may lack the information necessary for primer design (ex. sequence depth of each polymorphism). Therefore, we provide a command to create a VCF in a format suitable for subsequent analysis.<\/p>\n<p>\u2018mkprimer\u2019<\/p>\n<p>This command is responsible for designing primers that amplify around polymorphisms. To prepare markers, polymorphisms suitable for primer design are selected from the input VCF table according to the following criteria:<\/p>\n<p>                      (1)<\/p>\n<p>Genotypes differ between parental lines. For example, the GT fields of parental lines A and B in the input VCF are \u20180\/0\u2019 and \u20181\/1\u2019, respectively.<\/p>\n<p>                      (2)<\/p>\n<p>The reliability of polymorphism calling is high (passes GATK VariantFiltration: QD\u2009&lt;\u200920.0 and FS\u2009&gt;\u2009200.0 and SOR\u2009&gt;\u200910.0, fixed values).<\/p>\n<p>                      (3)<\/p>\n<p>The sequence depth is within the specified range (default: 2\u2009~\u2009200, modifiable).<\/p>\n<p>The primers used were designed using Primer3 software [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 33\" title=\"Untergasser A, Cutcutache I, Koressaar T, Ye J, Faircloth BC, Remm M, Rozen SG. Primer3\u2014new capabilities and interfaces. Nucl Acids Res. 2012;40(15):e115\u2013e115. &#010;                  https:\/\/doi.org\/10.1093\/nar\/gks596&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR33\" id=\"ref-link-section-d245668886e690\" rel=\"nofollow noopener\" target=\"_blank\">33<\/a>]. The primers used are designed so that they do not overlap other DNA mutations and are in accordance with other specified conditions (Supplementary Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>). The designed primers were checked to determine whether their amplicons were specific to the genome via BLAST software [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 8\" title=\"Camacho C, Coulouris G, Avagyan V, Ma N, Papadopoulos J, Bealer K, Madden TL. BLAST+: architecture and applications. BMC Bioinform. 2009;10(1):421. &#010;                  https:\/\/doi.org\/10.1186\/1471-2105-10-421&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR8\" id=\"ref-link-section-d245668886e696\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>]. The BLAST condition is based on the settings \u200b\u200bdescribed in the report of Primer-BLAST software [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 35\" title=\"Ye J, Coulouris G, Zaretskaya I, Cutcutache I, Rozen S, Madden TL. Primer-BLAST: a tool to design target-specific primers for polymerase chain reaction. BMC Bioinform. 2012;13(1):134. &#010;                  https:\/\/doi.org\/10.1186\/1471-2105-13-134&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR35\" id=\"ref-link-section-d245668886e699\" rel=\"nofollow noopener\" target=\"_blank\">35<\/a>]. By default, \u2018mkprimer\u2019 explores as many markers as possible (Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>a). If users want to save time, the number of polymorphisms to search for primers can be reduced by selecting the appropriate options. Finally, a VCF-formatted file containing the sequences of the designed primers added to the \u2018INFO\u2019 column is output.<\/p>\n<p>Fig.\u00a03<a class=\"c-article-section__figure-link\" data-test=\"img-link\" data-track=\"click\" data-track-label=\"image\" data-track-action=\"view figure\" href=\"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3\/figures\/3\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" aria-describedby=\"Fig3\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/07\/12859_2025_6211_Fig3_HTML.png\" alt=\"figure 3\" loading=\"lazy\" width=\"685\" height=\"329\"\/><\/a><\/p>\n<p>Example for the output of MKDESIGNER. a Genome-wide markers made by the \u2018mkprimer\u2019 command. b The strategy selecting markers in the \u2018mkselect\u2019 command. a Normally selected 384 markers using the \u2018mkselect\u2019 command. d 384 markers selected by the \u2018mkselect\u2019 command using the \u2018\u2013density\u2019 option to reduce markers near the centromeres<\/p>\n<p>\u2018mkselect\u2019<\/p>\n<p>This command narrows the markers to a specified number at equal intervals. The \u2018mkselect\u2019 works according to the following algorithm (Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>b):<\/p>\n<p>                      (1)<\/p>\n<p>The physical distance between all the DNA markers is calculated, and the pair of markers with the narrowest spacing is identified.<\/p>\n<p>                      (2)<\/p>\n<p>For that pair of markers, the physical distance to the other adjacent markers is calculated. The marker with the smaller distance is removed.<\/p>\n<p>                      (3)<\/p>\n<p>These steps are repeated until the specified number of markers is reached.<\/p>\n<p>The output files are a VCF file containing only the selected markers, a TSV file containing primer information, and a PNG file illustrating the physical positions of the selected markers (Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>c).<\/p>\n<p>In linkage mapping, physical distances among markers are not proportional to genetic distances, especially near centromeres. \u2018mkselect\u2019 can adjust the marker density of such regions by adding a tab-delimited file of the specified format to the \u2018\u2013density\u2019 option (Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>d).<\/p>\n<p>`taseq_hapcall\u2019<\/p>\n<p>The following commands belong to TASEQ, which is responsible for post-sequencing analysis. The graphical workflow is provided in Supplementary Fig.\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>. The first command is responsible for extracting target polymorphisms from raw sequence data (FASTQ) of multiplex PCR amplicons. First, the sequence reads in the input FASTQ files are trimmed using Trimmomatic [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Bolger AM, Lohse M, Usadel B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics. 2014;30(15):2114\u201320. &#010;                  https:\/\/doi.org\/10.1093\/bioinformatics\/btu170&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR4\" id=\"ref-link-section-d245668886e806\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>] (a). Moreover, a list of sequences before and after the target polymorphism was generated using SAMtools in FASTA format (b). Second, the reads in (a) are mapped to (b) using BWA [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 24\" title=\"Li H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM; 2013. &#010;                  arXiv:1303.3997&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR24\" id=\"ref-link-section-d245668886e809\" rel=\"nofollow noopener\" target=\"_blank\">24<\/a>]. Third, the genotypes of the markers are determined using GATK HaplotypeCaller. Finally, the results for all the lines are combined and output as a VCF file.<\/p>\n<p>\u2018taseq_genotype\u2019<\/p>\n<p>This command determines which genotype of each line, and each marker is homozygous for parent A (A), homozygous for parent B (B), heterozygous (H), or missing (-). The algorithm proceeds as follows:<\/p>\n<p>                      (1)<\/p>\n<p>Markers with a sequence depth less than 10 (default, modifiable) are considered missing (-).<\/p>\n<p>                      (2)<\/p>\n<p>If the proportion of minor alleles is less than 10% (default, modifiable), the marker genotype is homozygous for the major allele (A or B).<\/p>\n<p>                      (3)<\/p>\n<p>If the marker genotype is not determined in (2), the chi-square value is calculated for the ratio of the number of reference alleles and alternative alleles. If the p-value of chi-squared test is greater than 0.05 (modifiable), the marker genotype is missing (-); otherwise, it is heterozygous (H).<\/p>\n<p>The genotypes of each line and marker are output as a TSV file.<\/p>\n<p>\u2018taseq_filter\u2019<\/p>\n<p>This command removes markers with a specified percentage of missing data or less than a specified minor allele frequency. If parental lines are included in genotyping, only markers that are supported by the marker genotypes of the parental lines can be selected. The output CSV file is formatted to be used directly as genotype data in R\/qtl [<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 5\" title=\"Broman KW, Wu H, Sen \u015a, Churchill GA. R\/qtl: QTL mapping in experimental crosses. Bioinformatics. 2003;19(7):889\u201390. &#010;                  https:\/\/doi.org\/10.1093\/bioinformatics\/btg112&#010;                  &#010;                .\" href=\"http:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-025-06211-3#ref-CR5\" id=\"ref-link-section-d245668886e866\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a>].<\/p>\n<p>\u2018taseq_draw\u2019<\/p>\n<p>This command visualizes marker genotypes at the chromosomal level for each line. It outputs PNG files for the number of lines in the output directory.<\/p>\n","protected":false},"excerpt":{"rendered":"DNA-based genetic markers have been used in a wide range of biological fields, including medicine, genetics, and molecular&hellip;\n","protected":false},"author":2,"featured_media":4472,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[50],"tags":[5084,3871,5082,5083,5080,200,5076,5081,5079,5078,79,5077],"class_list":["post-4471","post","type-post","status-publish","format-standard","has-post-thumbnail","category-genetics","tag-algorithms","tag-bioinformatics","tag-computational-biology-bioinformatics","tag-computer-appl-in-life-sciences","tag-dna-markers","tag-genetics","tag-genotyping","tag-microarrays","tag-post-sequencing-analysis","tag-primer-design","tag-science","tag-targeted-amplicon-sequencing"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/4471","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/comments?post=4471"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/4471\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media\/4472"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media?parent=4471"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/categories?post=4471"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/tags?post=4471"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}