{"id":119392,"date":"2025-08-29T22:08:24","date_gmt":"2025-08-29T22:08:24","guid":{"rendered":"https:\/\/www.newsbeep.com\/us\/119392\/"},"modified":"2025-08-29T22:08:24","modified_gmt":"2025-08-29T22:08:24","slug":"new-anvil-data-explorer-makes-valuable-datasets-more-accessible-for-health-research","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/us\/119392\/","title":{"rendered":"New AnVIL Data Explorer makes valuable datasets more accessible for health research"},"content":{"rendered":"<p>Key takeaways<\/p>\n<p>The AnVIL Data Explorer makes it easier for researchers to access and reuse high-value datasets, including those that focus on specific diseases like Alzheimer\u2019s, cancer, and rare genetic disorders<\/p>\n<p>It currently organizes over 280 datasets from major NIH-supported consortia, with new data added quarterly<\/p>\n<p>By helping scientists combine information from multiple studies and build on existing resources, the Data Explorer will accelerate discovery, especially for rare and understudied conditions<\/p>\n<p>Collecting high-quality genomic data is time-consuming, expensive, and often only possible through large-scale national efforts. Thankfully, a new tool developed by the UC Santa Cruz Genomics Institute\u2019s Computational Genomics Lab is making existing datasets easier to find and use, ensuring that more researchers can build on these precious resources rather than starting from scratch.<\/p>\n<p>The AnVIL Data Explorer, which is now live and ready to use on the AnVIL platform, gives researchers fast, easy access to hundreds of human genomic datasets and enables them to construct research-specific groupings to accelerate scientific progress and new discoveries for health conditions ranging from cancer to rare disease.<\/p>\n<p>\u201cThis tool is about amplifying the impact of data and making the most of past public investments,\u201d said Benedict Paten, professor of biomolecular engineering and director for computational genomics for the UC Santa Cruz Genomics Institute. \u201cWe\u2019re making it easier for researchers to find the exact datasets they need, build cohorts, and start analyzing them right away. That means faster insights, more collaboration, and ultimately, better outcomes for human health.\u201d<\/p>\n<p>A human\u2019s genetic sequence is billions of base pairs long, which means that studying even a single genome generates massive amounts of raw data. Multiply that by hundreds or even thousands of individuals in larger genome-wide studies, and the data volume quickly becomes enormous.\u00a0<\/p>\n<p>Traditional genomic research workflows have researchers download these huge datasets to local servers, which is inefficient and costly. To address this problem, the National Human Genome Research Institute (NHGRI) created the Genomic Data Science Analysis, Visualization, and Informatics Lab-space (AnVIL) to allow researchers to access their data and analyze it centrally in the cloud.\u00a0<\/p>\n<p>The Data Explorer is a central component of AnVIL. It connects users with hundreds of datasets contributed by NHGRI-funded consortia and makes it easy to search, request access, and organize data for analysis\u2014all within a secure, scalable environment.<\/p>\n<p>\u201cThe goal is to help scientists spend less time collecting and searching for data and more time making discoveries that could improve human health,\u201d Paten said. \u201cIt also enables researchers to combine data across studies, opening the door to insights that wouldn\u2019t be possible with smaller datasets alone.\u201d<\/p>\n<p>The AnVIL Data Explorer is currently a gateway to over 280 datasets, including contributions from major consortia like 1000 Genomes Project, the <a href=\"https:\/\/news.ucsc.edu\/2023\/05\/pangenome-draft\/\" rel=\"nofollow noopener\" target=\"_blank\">Human Pangenome Reference Consortium<\/a>, <a href=\"https:\/\/news.ucsc.edu\/2022\/03\/t2t-genome\/\" rel=\"nofollow noopener\" target=\"_blank\">Telomere-to-Telomere<\/a>, and the Center for Alzheimer\u2019s and Related Dementias, among others. The number of available datasets is continually growing and supports studies in rare diseases, neurogenomics, aging, cancer, and beyond.\u00a0<\/p>\n<p>Although most datasets require access approval, users can browse core information for all of them before requesting access through a streamlined system tied to NIH credentials. Users can then analyze their data directly in Terra, which is AnVIL\u2019s secure analysis platform built on Google Cloud.<\/p>\n<p>The Explorer offers five views, organized by dataset, donor, biosample, activity, or file name, and includes a search function for quick navigation. In keeping with AnVIL\u2019s commitment to open science and interoperability, data from the Explorer can also be accessed through other platforms on Google Cloud, such as the National Heart, Lung, and Blood Institute\u2019s BioData Catalyst.\u00a0<\/p>\n<p>New users can create a free account by visiting<a href=\"https:\/\/anvilproject.org\" rel=\"nofollow noopener\" target=\"_blank\"> anvilproject.org<\/a> and clicking on \u201cLaunch Terra.\u201d Once signed in, they can access the AnVIL Data Explorer directly to browse datasets, and follow simple prompts to request access through dbGaP or DUOS. The platform\u2019s built-in help guides and documentation provide step-by-step guidance.\u00a0<\/p>\n<p>While AnVIL already supports a wide range of high-impact research, the platform is still in its early days. As with all collaborative science efforts, its potential for empowering discoveries will increase as more researchers jump on board. The Data Explorer team encourages users to submit feedback and suggestions via the \u201chelp\u201d section of the platform.<\/p>\n<p>\u201cThe platform is designed to grow and improve as more groups contribute data and provide feedback,\u201d Paten said. \u201cWe\u2019re eager to see the research community engage with it and help shape what comes next.\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"Key takeaways The AnVIL Data Explorer makes it easier for researchers to access and reuse high-value datasets, including&hellip;\n","protected":false},"author":2,"featured_media":119393,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[50],"tags":[200,79],"class_list":["post-119392","post","type-post","status-publish","format-standard","has-post-thumbnail","category-genetics","tag-genetics","tag-science"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/119392","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/comments?post=119392"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/119392\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media\/119393"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media?parent=119392"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/categories?post=119392"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/tags?post=119392"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}