The Semantic Precision for AI Retrieval of Knowledge (SPARK) dbGaP Challenge
The Semantic Precision for AI Retrieval of Knowledge (SPARK) dbGaP Challenge
Build AI that enables cross study biomedical discovery.
This challenge aims to build a modern, AI-powered search engine to identify relevant biomedical research studies that explore genomic and phenotypic associations stored within the NLM Database of Genotypes and Phenotypes (dbGaP).
Code Upload Open until 01/15/2027
Total prizes: $250,000
Description
Subject of the Challenge
The Semantic Precision for AI Retrieval of Knowledge (SPARK) dbGaP Challenge aims to advance the biomedical data and tools of the National Institutes of Health (NIH) and National Library of Medicine (NLM), making these resources high-quality, trusted, and widely used for public benefit. NIH/NLM encourages broad use of the resources it generates and treats data shared through its mechanisms as a common foundation for open discovery, freely available for everyone to build upon. Specifically, this challenge is to build a modern, AI-powered search engine to identify relevant biomedical research studies that explore genomic and phenotypic associations stored within the NLM Database of Genotypes and Phenotypes (dbGaP). Right now, public dbGaP study metadata and variable descriptions are heterogeneous across studies making study discovery haphazard and cross-study comparison difficult.
To solve this problem, the challenge consists of two distinct engineering tasks:
Task 1 (Semantic Variable Normalization and Ontological Alignment): Task 1 participants are tasked with developing an AI system that maps raw, unstructured dbGaP study variables to standardized ontology concepts in preferred source vocabularies LOINC, RxNorm, and SNOMED CT, thereby enabling precise, interoperable study discovery and cohort evaluation and other downstream use cases and applications. If no mapping is found in these sources, the Unified Medical Language System® (UMLS®) Metathesaurus, should be used.
The following provides an example mapping from dbGaP study variable to its corresponding expected target node in one of the task designated source vocabularies for this challenge task.
- Here, OCEA37L is a variable occurring in the StudyAccession phs000810.v2.p2 (Hispanic Community Health Study /Study of Latinos (HCHS/SOL)). It has a corresponding Variable Description “Occ Exp (Long Occ) - tobacco smoke (OCEA37L)”.
- The following provides a table of the information about this variable, available via the variables page associated with this study; for more details, see https://ncbi.nlm.nih.gov/projects/gap/cgi-bin/GetListOfAllObjects.cgi?study_id=phs000810.v2.p2&object_type=variable.
|
Study Accession |
Variable Accession |
Variable Name |
Variable Description |
|
phs000810.v2.p2 |
phv00527205.v1.p2 |
OCEA37L |
Occ Exp (Long Occ) - tobacco smoke (OCEA37L) |
Expected Response for OCEA37L:
- For the example variable OCEA37L given above, the expected response should identify the most appropriate source vocabulary target concept(s). This corresponds to SNOMED CT: 16090771000119104.
- SNOMED CT: 16090771000119104 provides the best semantic match for the OCEA37L variable and is considered a high-confidence exact mapping.
- LOINC: 20222-11: this can be considered a related match; however, it is considered a less precise mapping because it does not fully capture the meaning and context of the dbGaP variable (exposure at the primary occupation site).
- In this situation, the correct response is SNOMED CT: 16090771000119104 as it is the best fit. LOINC: 20222-11 is incorrect.
Task 2 (Conversational Cohort Feasibility & Discovery): Task 2 participants are tasked with developing an intuitive, privacy-preserving, reasoning-enabled, and conversational AI interface for cohort feasibility assessment and study discovery. Systems should enable investigators to formulate complex research information needs in natural language and reason over the available evidence in the challenge dataset -including publicly available dbGaP study documentation, data dictionaries, genotype and genomic metadata, and aggregated or summary statistics - to identify studies relevant to those needs. Task 2 participants are highly encouraged, but not required, to leverage the newly structured metadata generated by the ontological mapping pipelines in Task 1.
The following provides an example information need and the corresponding expected response for this Challenge Task:
- Information need: Find dbGaP studies involving adult participants with lung adenocarcinoma as the primary disease, where participant-level data include a history of smoking or smoking exposure and whole genomic sequences.
- Expected Response: List of Relevant Study Accessions: "phs001954.v2.p1" other study accessions considered relevant to this information need.
Challenge Tasks
Participants (whether an individual, team, or entity) must build a system that addresses one or both of the following distinct engineering tasks:
Task 1: Semantic Variable Normalization and Ontological Alignment
Background & Operational Bottleneck
A primary bottleneck in multi-study biomedical data aggregation is the inability to effectively discover and link relevant datasets when participant characteristics, clinical endpoints, and laboratory assays are described using disparate shorthand, local naming conventions, and non-standardized protocol descriptions. True semantic interoperability requires moving beyond simple string-based keyword matching through traditional search interfaces; it demands mapping highly variable, unstructured data dictionary definitions to precise, context-appropriate reference vocabulary concepts while preserving their embedded semantic relationships and protocol qualifiers.
Challenge Objective
Task 1 participants are tasked with developing an AI system that maps raw, unstructured dbGaP study variables to standardized ontology concepts in preferred source vocabularies LOINC, RxNorm, and SNOMED CT, thereby enabling precise, interoperable study discovery and cohort evaluation and other downstream use cases and applications. If no mapping is found in these sources, the Unified Medical Language System® (UMLS®) Metathesaurus, should be used.
The core objective is to ingest raw, unstructured, publicly available dbGaP study data, metadata, data dictionaries, and related public study information, and algorithmically map study variables to standardized reference concepts most appropriate in the context of the study. To support downstream interoperability, mappings should resolve to the established, high-quality reference vocabularies listed above. Systems should select the most appropriate target vocabulary concepts or combination of concepts based on the underlying data type and context, with the goal of maximizing semantic precision, consistency, and classification performance.
Task 2: Conversational Cohort Feasibility & Discovery
Background & Operational Bottleneck
Securing approval for controlled-access clinical data through the formal Data Access Request (DAR) process represents a massive investment of time and administrative overhead for both independent researchers and the National Institutes of Health (NIH). Far too often, investigators navigate this months-long pipeline only to discover post-approval that a specific study cohort's shared patient attributes or intersectional criteria do not actually match their specific analytical requirements, resulting in a severe slowing of research velocity.
Challenge Objective
Task 2 participants are tasked with developing an intuitive, privacy-preserving, reasoning-enabled, and conversational AI interface for cohort feasibility assessment and study discovery. Systems should enable investigators to formulate complex research information needs in natural language and reason over the available evidence in the challenge dataset -including publicly available dbGaP study documentation, data dictionaries, genotype and genomic metadata, and aggregated or summary statistics - to identify studies relevant to those needs.
Relevance should be determined by the required combination of study characteristics specified in the information need, rather than by the presence of individual terms or concepts in the study documentation. Systems should therefore demonstrate the ability to perform multi-hop reasoning across study characteristics, integrating multiple pieces of evidence, resolving relationships among characteristics, distinguishing required characteristics from incidental ones, and providing a clear, reasoned basis for why a study meets the investigator’s stated criteria.
Partners
This Challenge is brought to you by the National Institutes of Health (NIH) National Library of Medicine (NLM) and is co-sponsored by the NIH Chief AI Officer (CAIO) and the Office of Data Science Strategy (ODSS) within the NIH Office of the Director.
Statutory Authority to Conduct the Challenge
NLM is conducting this Challenge under the America Creating Opportunities to Meaningfully Promote Excellence in Technology, Education, and Science (COMPETES) Reauthorization Act of 2010, as amended (15 U.S.C. § 3719). NLM's authority to conduct the Challenge is further grounded in Section 465 of the Public Health Service Act (42 U.S.C. § 286), which establishes NLM's mission to acquire, organize, and disseminate biomedical information. The SPARK Challenge furthers this mission by directly incentivizing AI-powered semantic search tools that improve access to the data that NLM stewards.
Timeline
- Registrations Open: September 15, 2026
- Code Upload Close: January 15, 2027
- Submissions Open: January 18, 2027
- Submissions Close: January 31, 2027
- Judging Period: February 1, 2027-March 29, 2027
- Winner Announced: March 30, 2027
Prizes
Amount of the Prize
A total prize pool of $250,000 will be available to top-performing participants (whether an individual, team, or entity) in this Challenge. Task One and Task Two will each be assigned $125,000 in prize money, with eligibility contingent upon meeting all reproducibility, documentation, and compliance requirements established for the Challenge. Participants must provide sufficient materials to enable independent verification and reproduction of their submitted results and must comply with all applicable Challenge rules, data-use requirements, and submission procedures. The same team is eligible to win prize awards in both Task One and Task Two, provided its submissions independently satisfy the evaluation criteria and all reproducibility and compliance requirements for each Task.
Prize Money Distribution
The total allocated prize purse across both tasks is $250,000. For each individual task, a total pool of $125,000 will be split among the top two performing participants (whether an individual, team, or entity) using fixed percentage proportions:
|
Finish Position |
Prize Money Proportion |
Prize Amount per Task |
|
First Place (Winner) |
80% |
$100,000 |
|
Second Place |
20% |
$25,000 |
|
Total Per Task |
100% |
$125,000 |
Award Approving Official
The Award Approving Official will be Richard Scheuermann, PhD, Scientific Director, National Library of Medicine.
Payment of the Prize
Prizes awarded under this Challenge will be paid by electronic funds transfer and may be subject to federal income taxes. HHS/NIH will comply with Internal Revenue Service withholding and reporting requirements, where applicable.
Entities participating in this Challenge are encouraged, but not required, to obtain a free Unique Entity ID (UEI) via SAM.gov, as this will expedite prize payment. Additional information is available at sam.gov/content/entity-registration.
NIH/NLM reserves the right, in its sole discretion, to (a) cancel, suspend, or modify the Challenge, or any part of it, for any reason, and/or (b) not award any prizes if no submissions are deemed worthy.
NIH/NLM also reserves the right to validate submissions based on the Docker Images, repositories, and/or documentation provided by participants. Evidence of gaming the system will result in disqualification.
Use of Prize Funds
Participants are reminded that under this challenge announcement, NIH/NLM is awarding unrestricted cash prizes, not grants. The purpose of this Challenge is to reward innovation, not provide financial assistance, and NIH/NLM does not limit how winners may use cash prizes awarded to them.
Rules
Eligibility Rules
To be eligible to win a prize under this Challenge, a Participant (whether an individual, team, or entity):
- Shall have registered to participate in the Challenge under the rules promulgated by the National Institutes of Health (NIH) as published in this announcement;
- Shall have complied with all the requirements set forth in this announcement;
- In the case of a private entity, shall be incorporated in and maintain a primary place of business in the United States, and in the case of an individual, whether participating singly or in a group, shall be a citizen or permanent resident of the United States. However, non-U.S. citizens and non-permanent residents can participate as a member of a team that otherwise satisfies the eligibility criteria. Non-U.S. citizens and non-permanent residents are not eligible to win a monetary prize (in whole or in part). Their participation as part of a winning team, if applicable, may be recognized when the results are announced.
- Shall not be a federal entity or federal employee acting within the scope of their employment;
- Shall not be an employee of the Department of Health and Human Services (HHS, or any other component of HHS) acting in their personal capacity;
- Who is employed by a federal agency or entity other than HHS (or any component of HHS), or federally funded research and development center (FFRDC), should consult with an agency ethics official to determine whether the federal ethics rules will limit or prohibit the acceptance of a prize under this Challenge;
- Shall not be on the Excluded Parties List (i.e., not currently suspended or debarred from doing business with the federal government);
- Shall not be a judge of the Challenge, or any other party involved with the design, production, execution, or distribution of the Challenge or the immediate family of such a party (i.e., spouse, parent, stepparent, child, or stepchild); and
- Shall be 18 years of age or older at the time of submission.
Participation Rules
- Participants (whether individuals, team, or entities) may not use federal funds from a cooperative agreement, grant award or other transaction (OT) awards to develop their Challenge submissions or to fund efforts in support of their Challenge submissions unless use of such funds is consistent with the purpose, terms, and conditions of the grant, cooperative agreement, or OT award. Participants intending to use federal grant, cooperative agreement, or OT award funds must register for and participate in the challenge on behalf of the awardee institution, organization, or entity. If a Participant uses federal grant, cooperative agreement, or OT award funds and wins the Challenge, the prize must be treated as program income for purposes of the original grant, cooperative agreement, or OT award in accordance with applicable Uniform Administrative Requirements, Cost Principles, and Audit Requirements for Federal Awards (2 CFR § 200).
- Federal contractors may not use federal funds from a contract to develop their Challenge submissions or to fund efforts in support of their Challenge submissions.
- By participating in this Challenge, each Participant (whether an individual, team, or entity) agrees to assume any and all risks and waive claims against the federal government and its related entities, except in the case of willful misconduct, for any injury, death, damage, or loss of property, revenue, or profits, whether direct, indirect, or consequential, arising from participation in this Challenge, whether the injury, death, damage, or loss arises through negligence or otherwise.
- Based on the subject matter of the Challenge, the type of work that it will possibly require, as well as an analysis of the likelihood of any claims for death, bodily injury, property damage, or loss potentially resulting from Challenge participation, no Participant (whether an individual, team, or entity) participating in the Challenge is required to obtain liability insurance, or demonstrate financial responsibility, or agree to indemnify the federal government against third party claims for damages arising from or related to Challenge activities in order to participate in this Challenge.
- A Participant (whether an individual, team, or entity) shall not be deemed ineligible because the Participant used federal facilities or consulted with federal employees during the Challenge if the facilities and employees are made available to all Participants participating in the Challenge on an equitable basis.
- By participating in this Challenge, each Participant (whether an individual, team, or entity) warrants that they are sole author or owner of, or has the right to use, any copyrightable works that the submission comprises, that the works are wholly original with the Participant (or is an improved version of an existing work that the Participant has sufficient rights to use and improve), and that the submission does not infringe any copyright or any other rights of any third party of which the Participant is aware.
- By participating in this Challenge, each Participant (whether an individual, team, or entity) grants to the NIH an irrevocable, paid-up, royalty-free nonexclusive worldwide license to reproduce, publish, post, link to, share, and display publicly the submission on the web or elsewhere, and a nonexclusive, nontransferable, irrevocable, paid-up license to practice, or have practiced for or on its behalf, the solution throughout the world. Each Participant will retain all other intellectual property rights in their submissions, as applicable. To participate in the Challenge, each Participant must warrant that there are no legal obstacles to providing the above-referenced nonexclusive licenses of the Participant's rights to the federal government. To receive an award, Participants will not be required to transfer their intellectual property rights to NIH, but Participants must grant the federal government the nonexclusive licenses recited herein.
- Each Participant (whether an individual, group of individuals, or entity) agrees to follow all applicable federal, state, and local laws, regulations, and policies.
- Each Participant (whether an individual, team, or entity) participating in this Challenge must comply with all terms and conditions of these rules, and participation in this Challenge constitutes each such Participant's full and unconditional agreement to abide by these rules. Winning is contingent upon fulfilling all requirements herein.
- Each Participant (whether an individual, team, or entity) is permitted to enter a maximum of three (3) submissions per task. Failure to adhere to this limit—including any attempts to bypass this restriction via the creation of multiple accounts, ghost profiles, or collaborative collusion across distinct team registrations – will result in the immediate disqualification of the Participant, their entire team (if applicable), and all associated submissions from the challenge. All evaluations and potential prize eligibility for the offending parties will be rendered null and void.
- As a condition for winning a cash prize in this Challenge, each Participant (whether an individual, team, or entity) that has been selected as a winner must complete and submit all requested winner verification and payment documents to NIH within 14 business days of formal notification. Failure to return all required verification documents by the date specified in the notification may be a basis for disqualification of a cash prize winning submission.
- Each Participant in the Challenge (whether an individual, team, or entity) must successfully complete the official registration process. As a mandatory condition of entry, all Participants shall formally attest to and agree to strictly abide by the Participation and Eligibility Rules outlined herein. Failure to complete onboarding or adhere to these Rules at any point during the competition will result in immediate disqualification and forfeiture of any potential prize eligibility.
How to Enter
Participant Dataset
Participants will be provided with a fixed snapshot of dbGaP documents, associated metadata, and data dictionaries, including relevant version information where applicable. Responses to Task 2 information needs must be limited to and based on this snapshot. The snapshot is derived from official NCBI public repositories and includes the document content and associated metadata required to support the challenge tasks. Participants will therefore work from a common, fixed corpus, rather than independently downloading or accessing documents from dbGaP. The challenge dataset does not require access to controlled-access individual-level data. Use of controlled-access individual-level data is strictly prohibited. Any submission found to have used controlled-access individual-level data, either directly or indirectly through an external data source, service, or system, will be disqualified from the challenge.
The data for this challenge can be downloaded via the following link: https://ftp.ncbi.nlm.nih.gov/dbgap/NLM_Challenge/dbgap.tar.gz
This will result in one compressed TAR file being downloaded and titled as “dbgap.tar.gz”. Participants should unpack the archive to create a base directory named dbgap. The directory contains a README.txt file introducing the snapshot of dbGaP (the database of Genotypes and Phenotypes) studies and their associated objects, together with supporting reference material.
The dataset is organized so that each study and its associated objects can be inspected independently while preserving the relationships between objects:
dbgap/
├── studies/ one directory per dbGaP study, containing its objects
├── collections/ dbGaP collections (logical groupings of studies)
├── collections.json aggregate summary of collections and their members
└── web/ supporting dbGaP reference pages (HTML/PDF)
At a glance, the Challenge Dataset covers:
|
Object |
dbGaP Prefix |
Approx. Count |
|
Studies |
phs |
3,578 |
|
Phenotype datasets |
pht |
16,070 |
|
Variables |
phv |
465,853 |
|
Genotype datasets |
phg |
2,122 |
|
Documents |
phd |
8,079 |
|
Analyses |
pha |
17,277 |
|
Collections |
phs (collection hosts) |
23 |
The total size of the snapshot is approximately 6.3 GB unpacked (1.6 GB Compressed), dominated by approximately 466,000 variable JSON files and per-object HTML captures.
For each track/task, a small validation dataset will be released to participants at Challenge launch. For Track/Task 1, the dataset will consist of sample mappings. For Track/Task 2, it will consist of sample queries and associated studies labeled as relevant or not relevant.
Participants may also use the following resources to better understand the topic or problem:
- dbGaP Website: https://dbgap.ncbi.nlm.nih.gov/home/
- dbGaP Advanced Search: https://www.ncbi.nlm.nih.gov/gap/advanced_search
- https://www.ncbi.nlm.nih.gov/gap/advanced_search/?OBJ=variable&TERM=phv00054125.v1.p1
- Variable details and ontology mappings: https://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/variable.cgi?study_id=phs000280.v9.p3&phv=297987#ontology=LOINC
- Relevant Past TREC Biomedical Challenges: https://trec-biogen.github.io
- NIH Policy: https://osp.od.nih.gov/policies/artificial-intelligence/
- NLM Strategic Futures Catalyst Workshop: Artificial Intelligence and the Future of NLM (Day 1) https://videocast.nih.gov/watch/38915b89-5ddc-11f1-82c0-124f0a52e769
- NLM Strategic Futures Catalyst Workshop: Artificial Intelligence and the Future of NLM (Day 2) https://videocast.nih.gov/watch/38143122-5ddd-11f1-82c0-124f0a52e769
- Regarding access to mentors, advisors, or other support for crafting effective submissions: NLM will host 2–3 public information sessions, with recordings made available for each session.
Registration Process
To participate in this challenge, you will need to first register for the challenge by completing the registration with your participation information. Pay close attention to the registration dates posted for the challenge to ensure you submit your application on time.
Before you register, you will need to determine if you are registering as an Individual, Team Lead, Team Member, or a participant on behalf of an entity:
- Individual: Registering on behalf of themselves. To be eligible to receive a cash prize, the Individual must be a citizen or permanent resident of the United States.
- Team Lead: Registering on behalf of a team. To be eligible to receive a cash prize, the Team Lead must be a citizen or permanent resident of the United States. Non-U.S. citizens and non-permanent residents can participate as a member of a team that otherwise satisfies the eligibility criteria; however, citizens and non-permanent residents are not eligible to win a monetary prize (in whole or in part). If a dispute regarding the identity of the Team Lead who submitted the entry cannot be resolved to the satisfaction of the NIH, the affected submission will be deemed ineligible.
- Entity: Registering on behalf of an organization. To be eligible to receive a cash prize, the Entity must be incorporated in and maintain a primary place of business in the United States. The Entity must not be a federal entity, federal employee acting within the scope of their employment, or federally funded research and development center (FFRDC), and must not be on the Excluded Parties List (i.e., not currently suspended or debarred from doing business with the federal government). An officer or employee of the Entity must be designated as the point of contact for registration and communication purposes; the prize, if won, is awarded to the Entity rather than to the individual point of contact.
To enter this Challenge, participants must complete the following steps:
- Create an account on the Challenge Management Platform: https://repository.niddk.nih.gov/data-challenge Select the “Sign In / Register” button in the top right corner of the Challenge Management Platform homepage to get started. You will be directed to the RAS sign-in page.
- Select the eRA Commons, or Login.gov option and enter login credentials. The system will redirect you to the homepage after correctly entering your credentials.
- Click on the Challenge Title of the challenge that you wish to apply for. This will bring you to the challenge overview page.
- Click on the “Start Application” button to begin your application. The system will display a modal that will prompt you to specify whether you are registering as an individual, team lead, or entity. Select the individual, team lead, or entity option.
- You will be directed to the Data Challenge Application form where you will provide your applicant information, specify team members (if applicable), and challenge project information.
- Follow the on-screen prompts and enter information about your participation challenge.
- After completing each section, click on the “Next” button at the bottom right of the page. You may also click on the “Save” button at the bottom of the page to save your application.
- Once you have completed all sections, click on the “Submit” button at the bottom right of the page.
To ensure a rigorous and reproducible evaluation, all participating systems will operate under a standardized testing environment:
- Challenge Data: Participants will draw from the same snapshot of dbGaP data described above, ensuring a level evaluation playing field.
- Benchmark Dataset: Systems will be evaluated against expert-curated biomedical evaluation datasets, one set for each task, developed by NLM to rigorously assess mapping accuracy, and retrieval accuracy using task-specific evaluation metrics. More details are provided below.
- Deployment Architecture: Participants will develop systems to address one or both task tasks and must submit the following deliverables (see later section for more details):
- Their system responses to the task specific evaluation datasets in JSON-format,
- It is a reproducible set of Docker images containing required operating system, software, dependencies, and configurations. Submissions must be delivered fully packaged as a Docker image that can reproduce the responses provided in Deliverable 1 for each task. The Challenge Organizers reserve the right to seamlessly execute submitted systems against the private test set in a secure, isolated environment, as needed.
- More specific details and requirements for each task submission is available via the Challenge Management Platform.
- It is a GitHub repository containing source code, execution instructions, architecture documentation, dependencies, external resources, licenses, and reproduction instructions. Noncompliant or incomplete submissions may be deemed ineligible.
Submission Process
To be eligible to win a prize under a challenge, participants will need to submit their challenge solution for evaluation. Pay close attention to the submission dates posted for the challenge to ensure you submit your solution on time.
- Log into the Challenge Management Platform using your eRA Commons, or Login.gov credentials.
- Select Submit Solution from the upper navigation bar. This will bring you to the Submit Challenge Solution page.
- Under the Upload Solution Package section, click on the “Select File to Upload” button to upload your solution package (i.e., ZIP file containing the required submission materials, as specified by the challenge instructions).
- Under the Confirm Upload to GitHub section, select the checkbox to confirm that you have uploaded your solution to GitHub.
- Click on the “Submit” button at the bottom of the page to submit your solution.
- Note: If participating as a team lead or entity, the Point of Contact should submit the solution package on behalf of the team or entity.
Task 1: Submission Requirements
For Task 1, the Semantic Variable Normalization and Ontological Alignment Task, an evaluation dataset consisting of 1,000 dbGaP study variables will be made available to registered participants prior to the close of the Challenge. Participants must apply their systems to the evaluation dataset and generate the required system outputs. The resulting participant submission file must be provided in JSONL format and conform to the schema, file structure, naming conventions, and other technical specifications established by the Challenge organizers. Detailed submission specifications and instructions is available through the Challenge Management Platform.
To be considered complete and eligible for evaluation, each submission must include the following components:
- Submission File: Participants must upload a submission file containing their system-generated responses to the evaluation dataset. The output file must conform to the required JSONL schema, and other formatting and validation requirements specified by the Challenge organizers. Submissions that do not conform to the required technical specifications may be deemed incomplete or ineligible for evaluation.
- Containerized Application Submission: By January 15, 2027, participants must submit a containerized application the solution, a description of their system and its components, together with sufficient technical documentation and instructions to enable NIH to reproduce, deploy, and evaluate the system within an NIH-managed AWS environment. All software, models, indexes, data resources, documentation, and other materials that comprise or support the submitted system must have been created or last modified on or before January 15, 2027. A link to the participant’s GitHub/Docker repository for the solution will be accepted.
- System Architecture Description: Submissions should describe the system architecture and components, including retrieval, ranking, embeddings, indexes, models, LLM inference, and any other software or data resources required to execute the system. Participants may design their systems as a single containerized application or as multiple containerized services, including distributed architectures in which different components run on separate compute nodes. Participants should also provide deployment and execution instructions, along with an estimate of the computational and storage resources required (e.g., CPU, memory, GPU, and storage). Submissions must explicitly disclose all training data sources, base model selections, architectural diagrams, post-processing heuristics, and any supplemental third-party datasets utilized. Furthermore, any specific adaptations applied to the pipeline or its components—such as fine-tuning, Reinforcement Learning from Human Feedback (RLHF), post-training enhancements, or other algorithmic optimizations—must be clearly documented.
- Initial Submission Scope: Participants will not be required to submit executable container images, source code, or large data artifacts (e.g., embeddings, indexes, or models) as part of this initial submission. These materials will be requested during the evaluation phase in accordance with the Challenge procedures and technical guidance.
- AWS Infrastructure Responsibility: For evaluation, NIH will provision and manage the underlying AWS infrastructure, including compute, storage, networking, and orchestration. Participants will not be required to provision or manage AWS resources, Kubernetes clusters, or Terraform infrastructure.
- Deadline Verification: To support verification of the January 15, 2027, submission deadline, participants may be required to provide cryptographic hashes or other verifiable records for software, models, indexes, data artifacts, and other system components identified in their submission. When requested during the evaluation phase, the corresponding executable or data artifacts must match the versions documented and verified at the submission deadline.
- Model and Resource Restrictions: Submitted systems may use only publicly available, open-weight Large Language Models (LLMs), open-source software, and/or other publicly available computing resources. Systems may not invoke external APIs to proprietary frontier models (e.g., commercial hosted LLM endpoints) or rely on non-public, restricted-access software, models, datasets, or other resources during execution. Any pretrained/fine-tuned models used by the system must be included within the submitted Docker images.
- Docker Image Execution and Network Restrictions:
- The challenge organizers reserve the right to run the submitted Docker Images and verify system reproducibility, resource utilization, and compliance with challenge rules. Specifically:
- Submitted Docker Images will be executed in a completely isolated environment with all external network calls blocked. Limited access will be permitted to NLM-hosted services, such as the UMLS Terminology Services (UTS), and other challenge specific services. Details of these approved services and endpoints will be made available when the challenge is published. This zero-egress policy (excluding the permitted services) prevents both data leaks and runtime reliance on external computing resources, tools, or services.
- Disclosure Responsibility: Participants must clearly identify all external data, models, software, and other third-party resources used in connection with their submission, including applicable licenses, terms of use, access restrictions, and attribution requirements. Participants are responsible for ensuring that they possess all rights, permissions, licenses, and authorizations necessary to use such resources in connection with the Challenge and to permit evaluation of their submissions by the Challenge organizers.
- Proof of Execution: Prior to final submission, participants must execute a mandatory validation script directly on their Docker Image, which will be deployed on the specific AWS hardware corresponding to the challenge-designated AWS base image. This step is required to self-certify that the Docker Image compiles, runs smoothly, and correctly outputs data in the required format within the target environment. Submissions that fail this baseline execution test will not be evaluated. More details on the exact requirements are available on the Challenge Management Platform.
- Disclaimer: Participants are solely responsible for ensuring their submissions and activities comply with all applicable laws, regulations, and policies. NIH cannot provide legal advice to third parties. Participants should consult their own legal counsel as necessary and appropriate.
Task 2: Submission Requirements
The submission requirements for Task 2 are the same as those specified for Task 1 above and are therefore not repeated here for brevity. Task 2 comes with a separatee set of evaluation information needs.
Judging Criteria
Basis Upon Which a Winner Will be Selected. Submissions in both tasks of the Challenge will be evaluated by a Technical Review Panel composed of subject matter experts in artificial intelligence, data science, information technology, and biomedical research. The Evaluation Panel will assess eligible submissions using the published evaluation criteria and provide feedback to the Judging Panel composed of federal employees.
The Judging Panel, with representation from NIH and other operating divisions, will review the Technical Review Panel's assessments alongside other relevant Challenge inputs, including compliance with applicable laws, regulations, and Challenge Rules.
Based on these factors, the Judging Panel will recommend final prize selections to the Award Approving Official. Final selections are made by the Award Approving Official, in their sole discretion, consistent with the published evaluation framework and applicable federal procedures. The Award Approving Official will review and approve final selections in accordance with applicable NIH requirements. A total of 4 winners is anticipated (the top 2 participants within each task).
Evaluation Criteria
Evaluation Metrics
Task 1: Semantic Variable Normalization and Ontological Alignment
Task 1 submissions will be evaluated using the following criteria:
- Mapping Alignment Accuracy (MAA) score: Measures the Percentage of dbGaP study variables for which the system identifies the expert-verified target concept(s)
- Ease of Use: Assesses the effort required to deploy, configure, operate, and use the system in the standardized NIH cloud environment.
- Compute Time: Measures the computational time required to generate responses.
- Compute Effort: Measures the computational effort required to generate responses, including the resources and processing required by the system.
- Memory Footprint: Measures the memory and other computational resources required to operate the system and generate responses.
MAA will serve as the primary metric for determining final system rankings and identifying the overall winning systems. In the event of a tie in this metric, the additional evaluation criteria listed above may be used to differentiate systems and determine the final ranking.
Task 2: Conversational Cohort Feasibility & Discovery
Task 2 evaluates an information retrieval (IR) engine's ability to process unstructured natural-language queries and returns a ranked list of relevant dbGaP studies.
Task 2 will use the following evaluation criteria to select winning submissions:
- NDCG: Measures the relevance and ranking quality of retrieved studies against the expert-verified reference set.
- Ease of Use: Assesses the effort required to deploy, configure, operate, and use the system in the standardized NIH cloud environment.
- Compute Time: Measures the computational time required to generate responses.
- Compute Effort: Measures the computational effort required to generate responses, including the resources and processing required by the system.
- Memory Footprint: Measures the memory and other computational resources required to operate the system and generate responses.
NDCG will serve as the primary metric for determining final system rankings and identifying the overall winning systems. In the event of a tie in this metric, the additional evaluation criteria listed above may be used to differentiate systems and determine the final ranking.
Submission Packet Validation
To efficiently manage computing resources and ensure the scientific integrity of the results, submissions will undergo a multi-stage validation, filtering, and dual-layer evaluation process.
1. Timeline and Initial Validation Phase
- Validation Set Release: For each task, a small validation dataset will be released to participants at challenge launch time. This will consist of sample mappings in the case of Task 1, and of sample queries and relevant/not-relevant studies in the case of Task 2. Mandatory Local Sanity Check: Before submitting to their containerized systems, participants must execute a mandatory baseline validation script locally over the provided validation data to ensure their Docker Image compiles and outputs data in the correct format (e.g., JSON).
2. Anti-Gaming Probes and Adversarial Security
- Anti-gaming Checks: To prevent "memorization gaming"—where a Docker Image is hardcoded with pre-computed answers mapping directly to known validation or public study data—the NLM will inject entirely unseen, private adversarial probes into the final test phase.
- Behavioral Auditing & Adversarial Security Post-competition: Challenge organizers reserve the right to execute participant Docker Images over private queries and synthetic variables. Systems that exhibit a drastic drop in accuracy or display hardcoded outputs tailored specifically to public metadata schemas will be flagged for anomalous "Trojan horse" behavior. Upon judicial review by the panel, any system found to be using pre-computed or non-generalized outputs will be disqualified.
3. Scope Boundaries
To maintain a fair, standardized, and legally compliant evaluation environment, participants must strictly adhere to the following operational boundaries:
- In-Scope (Permitted Data):
- Publicly Available Metadata: Participants are explicitly permitted to utilize any publicly available metadata associated with each target dbGaP study, including study descriptions, data dictionaries, variable text, and trial design protocols available in the challenge dataset.
- Documented Supplementary Data: If any external or supplementary data sources (such as open-access clinical vocabularies, pre-trained biomedical language models, or public medical ontologies) are integrated into the pipeline, they must be fully disclosed and meticulously detailed within the participant's formal submission documentation.
- Out-of-Scope (Strictly Prohibited Data):
- Controlled-Access Genotypic/Phenotypic Data: Participants are strictly prohibited from attempting to access, download, or reverse-engineer individual-level patient data, genomic sequence data, or any other or controlled-access from dbGaP. Any integration of unauthorized or non-public controlled-access clinical trial data will result in immediate disqualification.
- Use study data beyond challenge dataset (such as scientific articles) is not permitted.
Additional Information
Participants are encouraged to join one of the following information sessions to learn more and ask questions:
- October 8 at 3:00 PM EST Register Here
- November 5 at 12:00 PM EST Register Here
- January 13 at 2:00 PM EST Register Here
Point of Contact
Questions can be directed to the following email mailbox for this challenge at [email protected]
This page last reviewed on
