The Semantic Precision for AI Retrieval of Knowledge (SPARK) PubMed/PMC Challenge

The Semantic Precision for AI Retrieval of Knowledge (SPARK) PubMed/PMC Challenge

Build AI that transforms biomedical discovery

Image
The Semantic Precision for AI Retrieval of Knowledge (SPARK) PubMed/PMC Challenge

This challenge aims to catalyze the development of AI-powered systems capable of generating synthesized, citation-backed responses that support complex literature exploration and help biomedical researchers identify and take appropriate next steps, substantially reducing the time and cognitive burden associated with navigating and interpreting the scientific literature.

Registration open until 1/15/2026

Total prizes: $250,000

Description

Subject of the Challenge

The Underlying Problem

The Semantic Precision for AI Retrieval of Knowledge (SPARK) PubMed/PMC Challenge aims to advance the biomedical data and tools of the National Institutes of Health (NIH) and National Library of Medicine (NLM), making these resources high-quality, trusted, and widely used for public benefit. NIH/NLM encourages broad use of the resources it generates and treats data shared through its mechanisms as a common foundation for open discovery, freely available for everyone to build upon. Specifically, the volume of biomedical literature reporting the results and conclusions from research endeavors continues to grow exponentially. NLM resources, such as PubMed and PubMed Central (PMC), now index and disseminate tens of millions of abstracts and full-text articles. While this expanding body of literature represents an unparalleled scientific asset, it frequently leaves researchers overwhelmed by the sheer volume and structural complexity of search results, even using the advanced keyword search options.

Currently, investigators querying PubMed and PMC must manually sift through, evaluate, and synthesize findings from lengthy lists of keyword-ranked abstracts and full-text articles—a time-consuming, labor-intensive process that offers limited support for identifying the most relevant, credible, and high-impact evidence that meet the investigators' objectives. This cognitive bottleneck slows down systematic reviews, hypothesis generation, and rapid evidence of discovery and validation.

Challenge Objective and Intended Effect

This exponential growth presents a critical opportunity to modernize literature discovery at the NLM. By leveraging artificial intelligence (AI) to enable semantic search, contextual relevance, and intent-aware retrieval, we can empower researchers to efficiently navigate, prioritize, and synthesize relevant evidence across massive literature corpuses.

Accordingly, the NIH invites AI-driven solutions capable of intelligent, context-aware, and impact-oriented discovery of biomedical literature within PubMed and PMC. The overarching objective is to move beyond rigid, keyword-based retrieval toward systems that understand scientific meaning, prioritize robust evidence, and accelerate cross-disciplinary scientific discovery.

Ultimately, this challenge aims to catalyze the development of AI-powered systems capable of generating synthesized, citation-backed responses that support complex literature exploration and help biomedical researchers identify and take appropriate next steps, substantially reducing the time and cognitive burden associated with navigating and interpreting the scientific literature.

The Desired Solutions: Technical Requirements

Participants (whether an individual, team, or entity) are tasked with building an end-to-end system that accepts a complex biomedical information need as input (framed as a natural language question or instruction) and returns an intelligent, grounded output. Systems must fulfill three core capabilities:

  1. Grounding: Every response must be strictly grounded in valid PubMed/PMC citations and explicit, relevant text passages.
  2. Synthesis: Information must be compiled into a coherent, objective, and contextually accurate response.

Challenge Tasks

Participants (whether an individual, team, or entity) must build a system that addresses one or both of the following distinct engineering tasks:

Task 1: Known-Item & Provenance Retrieval

  • Focus: Given a challenge-specific corpus of documents comprised of documents selected from PubMed and PMC (details will provided below) retrieve the specific, relevant document or set of documents that satisfy an information need expressed in natural language.
  • Example:
    • Input Query: “What was the landmark paper that demonstrated the mechanistic discovery and potential to exploit CRISPR-Cas9 for RNA-programmable genome editing in vitro?”
    • Expected Response: PMID: 22745249

Task 2: Open-Ended Exploratory Information Needs

  • Focus: Designed for broad, exploratory biomedical research inquiries requiring thematic synthesis across multiple sources. The system must generate a natural language synthesis that is systematically backed by valid PubMed/PMC citations and supporting passages drawn from a challenge-specific corpus of documents comprised of documents selected from PubMed and PMC (details will provided below). The ideal response is highly objective—covering all essential facts without extraneous information—and must explicitly address any controversial findings or scientific disagreements.
  • Example:
    • Input Query: What genetic and other molecular changes are associated with the acquisition of treatment resistance leading to disease progression in chronic lymphocytic leukemia (CLL) patients treated with ibrutinib?
    • Expected Response: Ibrutinib (aka PCI-32765) inhibits a number of protein tyrosine kinases, including Bruton’s Tyrosine Kinase (BTK), through covalent bonding to cysteine residues in their active sites, and has been approved for the treatment of chronic lymphocytic leukemia (CLL), where signaling through the B cell antigen receptor mediated by BTK is required for cell activation and survival [1]. Using targeted and whole-exome sequencing and clonal evolution analysis, mutations in the BTK, PLCG2, and ITPKB genes were associated with clonal expansion in different patients with relapsing disease, suggesting the acquisition of treatment resistance due to mutations in these genes [2].
    • References:
      • [1] Ahn IE, Brown JR. Targeting Bruton's Tyrosine Kinase in CLL. Frontiers in Immunology. 2021. PMID: 34248972, PMCID: PMC8261291.
      • [2] Landau DA, Sun C, Rosebrock D, et al. The evolutionary landscape of chronic lymphocytic leukemia treated with ibrutinib targeted therapy. Nature Communications. 2017. PMID: 29259203, PMCID: PMC5736707.

Partners

This Challenge is brought to you by the National Institutes of Health (NIH) National Library of Medicine (NLM) and is co-sponsored by the Chief AI Officer (CAIO) and the Office of Data Science Strategy (ODSS) within the NIH Office of the Director.

Statutory Authority to Conduct the Challenge

NLM is conducting this Challenge under the America Creating Opportunities to Meaningfully Promote Excellence in Technology, Education, and Science (COMPETES) Reauthorization Act of 2010, as amended (15 U.S.C. § 3719). NLM's authority to conduct the Challenge is further grounded in Section 465 of the Public Health Service Act (42 U.S.C. § 286), which establishes NLM's mission to acquire, organize, and disseminate biomedical information. The SPARK Challenge furthers this mission directly by incentivizing the development of AI-powered semantic search tools that improve access to the biomedical literature and the data that NLM stewards.

Timeline

  • Registrations Open: September 15, 2026
  • Code Upload Close: January 15, 2027
  • Submissions Open: January 18, 2027
  • Submissions Close: January 31, 2027
  • Judging Period: February 1, 2027-March 29, 2027
  • Winner Announced: March 30, 2027

Prizes

Amount of the Prize

A total prize pool of $250,000 will be available to top-performing participants (whether an individual, team, or entity) in this Challenge. Task One and Task Two will each be assigned $125,000 in prize money, with eligibility contingent upon meeting all reproducibility, documentation, and compliance requirements established for the Challenge. Participants must provide sufficient materials to enable independent verification and reproduction of their submitted results and must comply with all applicable Challenge rules, data-use requirements, and submission procedures. The same team is eligible to win prize awards in both Task One and Task Two, provided its submissions independently satisfy the evaluation criteria and all reproducibility and compliance requirements for each Task.

Prize Purse Distribution

The total allocated prize purse across both tasks is $250,000. For each individual task, a total pool of $125,000 will be split among the top two performing participants (whether an individual, team, or entity) using fixed percentage proportions:

Finish Position

Prize Money Proportion

Prize Amount per Task

First Place (Winner)

80%

$100,000

Second Place

20%

$25,000

Total Per Task

100%

$125,000

Award Approving Official

The Award Approving Official will be Richard Scheuermann, PhD, Scientific Director, National Library of Medicine.

Payment of the Prize

Prizes awarded under this Challenge will be paid by electronic funds transfer and may be subject to federal income taxes. HHS/NIH will comply with Internal Revenue Service withholding and reporting requirements, where applicable.

Entities participating in this Challenge are encouraged, but not required, to obtain a free Unique Entity ID (UEI) via SAM.gov, as this will expedite prize payment. Additional information is available at sam.gov/content/entity-registration.

NIH/NLM reserves the right, in its sole discretion, to (a) cancel, suspend, or modify the Challenge, or any part of it, for any reason, and/or (b) not award any prizes if no submissions are deemed worthy.

NIH/NLM also reserves the right to validate submissions based on the Docker Images, repositories, and/or documentation provided by participants. Attempts to manipulate or circumvent the evaluation process will result in disqualification.

Use of Prize Funds

Participants are reminded that under this challenge announcement, NIH/NLM is awarding unrestricted cash prizes, not grants. The purpose of this Challenge is to reward innovation, not provide financial assistance, and NIH/NLM does not limit how winners may use cash prizes awarded to them.

Rules

Eligibility Rules

To be eligible to win a prize under this Challenge, a Participant (whether an individual, team, or entity):

  1. Shall have registered to participate in the Challenge under the rules promulgated by the National Institutes of Health (NIH) as published in this announcement;
  2. Shall have complied with all the requirements set forth in this announcement;
  3. In the case of a private entity, shall be incorporated in and maintain a primary place of business in the United States, and in the case of an individual, whether participating singly or in a group, shall be a citizen or permanent resident of the United States. However, non-U.S. citizens and non-permanent residents can participate as a member of a team that otherwise satisfies the eligibility criteria. Non-U.S. citizens and non-permanent residents are not eligible to win a monetary prize (in whole or in part). Their participation as part of a winning team, if applicable, may be recognized when the results are announced.
  4. Shall not be a federal entity or federal employee acting within the scope of their employment;
  5. Shall not be an employee of the Department of Health and Human Services (HHS, or any other component of HHS) acting in their personal capacity;
  6. Who is employed by a federal agency, entity other than HHS (or any component of HHS), or federally funded research and development center (FFRDC), should consult with an agency ethics official to determine whether the federal ethics rules will limit or prohibit the acceptance of a prize under this Challenge;
  7. Shall not be on the Excluded Parties List (i.e., not currently suspended or debarred from doing business with the federal government);
  8. Shall not be a judge of the Challenge, or any other party involved with the design, production, execution, or distribution of the Challenge or the immediate family of such a party (i.e., spouse, parent, stepparent, child, or stepchild); and
  9. Shall be 18 years of age or older at the time of submission.

Participation Rules

  1. Participants (whether an individual, team, or entity) may not use federal funds from a cooperative agreement, grant award or other transaction (OT) awards to develop their Challenge submissions or to fund efforts in support of their Challenge submissions unless use of such funds is consistent with the purpose, terms, and conditions of the grant, cooperative agreement, or OT award. Participants intending to use federal grant, cooperative agreement, or OT award funds must register for and participate in the challenge on behalf of the awardee institution, organization, or entity. If a Participant uses federal grant, cooperative agreement, or OT award funds and wins the Challenge, the prize must be treated as program income for purposes of the original grant, cooperative agreement, or OT award in accordance with applicable Uniform Administrative Requirements, Cost Principles, and Audit Requirements for Federal Awards (2 CFR § 200).
  2. Federal contractors may not use federal funds from a contract to develop their Challenge submissions or to fund efforts in support of their Challenge submissions.
  3. By participating in this Challenge, each Participant (whether an individual, team, or entity) agrees to assume any and all risks and waive claims against the federal government and its related entities, except in the case of willful misconduct, for any injury, death, damage, or loss of property, revenue, or profits, whether direct, indirect, or consequential, arising from participation in this Challenge, whether the injury, death, damage, or loss arises through negligence or otherwise.
  4. Based on the subject matter of the Challenge, the type of work that it will possibly require, as well as an analysis of the likelihood of any claims for death, bodily injury, property damage, or loss potentially resulting from Challenge participation, no Participant (whether an individual, team, or entity) participating in the Challenge is required to obtain liability insurance, or demonstrate financial responsibility, or agree to indemnify the federal government against third party claims for damages arising from or related to Challenge activities in order to participate in this Challenge.
  5. A Participant (whether an individual, team, or entity) shall not be deemed ineligible because the Participant used federal facilities or consulted with federal employees during the Challenge if the facilities and employees are made available to all Participants participating in the Challenge on an equitable basis.
  6. By participating in this Challenge, each Participant (whether an individual, team, or entity) warrants that they are sole author or owner of, or has the right to use, any copyrightable works that the submission comprises, that the works are wholly original with the Participant (or is an improved version of an existing work that the Participant has sufficient rights to use and improve), and that the submission does not infringe any copyright or any other rights of any third party of which the Participant is aware.
  7. By participating in this Challenge, each Participant (whether an individual, team, or entity) grants to the NIH an irrevocable, paid-up, royalty-free nonexclusive worldwide license to reproduce, publish, post, link to, share, and display publicly the submission on the web or elsewhere, and a nonexclusive, nontransferable, irrevocable, paid-up license to practice, or have practiced for or on its behalf, the solution throughout the world. Each Participant will retain all other intellectual property rights in their submissions, as applicable. To participate in the Challenge, each Participant must warrant that there are no legal obstacles to providing the above-referenced nonexclusive licenses of the Participant's rights to the federal government. To receive an award, Participants will not be required to transfer their intellectual property rights to NIH, but Participants must grant the federal government the nonexclusive licenses recited herein.
  8. Each Participant (whether an individual, team, or entity) agrees to follow all applicable federal, state, and local laws, regulations, and policies.
  9. Each Participant (whether an individual, team, or entity) participating in this Challenge must comply with all terms and conditions of these rules, and participation in this Challenge constitutes each such Participant's full and unconditional agreement to abide by these rules. Winning is contingent upon fulfilling all requirements herein.
  10. Each Participant (whether an individual, team, or entity) is permitted to enter with a maximum of three (3) submissions per task. Failure to adhere to this limit—including any attempts to bypass this restriction via the creation of multiple accounts, ghost profiles, or collaborative collusion across distinct team registrations – will result in the immediate disqualification of the Participant, their entire team (if applicable), and all associated submissions from the challenge. All evaluations and potential prize eligibility for the offending parties will be rendered null and void.
  11. As a condition for winning a cash prize in this Challenge, each Participant (whether an individual, team, or entity) that has been selected as a winner must complete and submit all requested winner verification and payment documents to NIH within 14 business days of formal notification. Failure to return all required verification documents by the date specified in the notification may be a basis for disqualification of a cash prize winning submission.
  12. Each Participant in this Challenge (whether an individual, team, or entity) must successfully complete the official registration process. As a mandatory condition of entry, all Participants shall formally attest to and agree to strictly abide by the Participation and Eligibility Rules outlined herein. Failure to complete onboarding or adhere to these Rules at any point during the competition will result in immediate disqualification and forfeiture of any potential prize eligibility.

How to Enter

Participant Dataset

This challenge comes with its own Challenge Dataset (Corpus) that can be used in responding to the questions in this task. It consists of two constituent corpora:

PubMed

PubMed Central (PMC)

  • A subset of PMC shall be used as a source of full-text documents for responding to questions in this task. The resulting subset includes PMC publications that satisfy the specified Creative Commons licensing, NIH funding or intramural status, publication-date, and publication-status criteria.
  • A list of PMCIDs defining the PMC component of the Challenge Corpus is available in the file named “Challenge_PMCIDs.txt”
  • The corresponding PMC articles may be downloaded or accessed programmatically using NCBI/PMC-supported services, including the PMC Cloud Service, PMC OAI-PMH API, BioC API, NCBI E-Utilities, and PMC ID Converter API. The PMC Cloud Service supports bulk access to machine-readable article files, while the OAI-PMH and BioC APIs support programmatic retrieval of article content. See https://pmc.ncbi.nlm.nih.gov/tools/developers/ Participants are responsible for complying with applicable article-level license and usage requirements. PMC identifies these services as the supported mechanisms for automated retrieval of PMC content.
  • PMC Cloud Service / AWS — Bulk Download
    • The PMC Cloud Service provides a convenient mechanism for downloading the Challenge Corpus in bulk.
    • Available article files include JATS XML, plain text, metadata, and, where applicable, PDFs and supplementary materials.
    • The PMC datasets are accessible through AWS without requiring an AWS account or login.
    • Participants should use the provided PMCID list to identify and retrieve the articles comprising the PMC component of the Challenge Corpus.
  • For more information on the PMC Cloud Service and available datasets, please see: https://pmc.ncbi.nlm.nih.gov/tools/cloud/
  • For general information about PubMed Central, please see: https://pmc.ncbi.nlm.nih.gov/

Participants may also use the following resources to better understand the topic or problem:

Registration Process

To participate in this challenge, you will need to first register for the challenge by completing the registration with your participation information. Pay close attention to the registration dates posted for the challenge to ensure you submit your application on time.

Before you register, you will need to determine if you are registering as an Individual, Team Lead, or Entity:

Individual: Registering on behalf of themselves

  • To be eligible to receive a cash prize, the Individual must be a citizen or permanent resident of the United States.

Team Lead: Registering on behalf of a group of individuals

  • To be eligible to receive a cash prize, the Team Lead must be a citizen or permanent resident of the United States. Non-U.S. citizens and non-permanent residents can participate as a member of a team that otherwise satisfies the eligibility criteria; however, citizens and non-permanent residents are not eligible to win a monetary prize (in whole or in part). If a dispute regarding the identity of the Team Lead who submitted the entry cannot be resolved to the satisfaction of the NIH, the affected submission will be deemed ineligible.

Entity: Registering on behalf of an organization

  • To be eligible to receive a cash prize, the Entity must be incorporated in and maintain a primary place of business in the United States. The Entity must not be a federal entity, federal employee acting within the scope of their employment, or federally funded research and development center (FFRDC), and must not be on the Excluded Parties List (i.e., not currently suspended or debarred from doing business with the federal government). An officer or employee of the Entity must be designated as the point of contact for registration and communication purposes; the prize, if won, is awarded to the Entity rather than to the individual point of contact.

To enter this Challenge, participants must complete the following steps:

  1. Create an account on the Challenge Management Platform: https://repository.niddk.nih.gov/data-challenge. Select the “Sign In / Register” button in the top right corner of the Challenge Management Platform homepage to get started. You will be directed to the RAS sign-in page.
  2. Select the eRA Commons, or Login.gov option and enter login credentials. The system will redirect you to the homepage after correctly entering your credentials.
  3. Click on the Challenge Title of the challenge that you wish to apply for (e.g., SPARK PubMed/PMC Challenge). This will bring you to the challenge overview page.
  4. Click on the “Start Application” button to begin your application. The system will display a modal that will prompt you to specify whether you are registering as an Individual, Team Lead, or Entity. Select either the Individual, Team Lead, or Entity option.
  5. You will be directed to the Data Challenge Application form where you will provide your applicant information, specify team members (if applicable), and challenge project information.
  6. Follow the on-screen prompts and enter information about your challenge participation.
  7. After completing each section, click on the “Next” button at the bottom right of the page. You may also click on the “Save” button at the bottom of the page to save your application.
  8. Once you have completed all sections, click on the “Submit” button at the bottom right of the page.

To ensure a rigorous and reproducible evaluation, all participating systems will operate under a standardized testing environment:

  • Literature Corpus: Participants will draw from the same frozen subset of PubMed/PMC articles ensuring a level evaluation playing field.
  • Benchmark Dataset: Systems will be evaluated against expert-curated biomedical benchmark question sets, one set for each task, developed by NLM to rigorously assess retrieval accuracy, synthesis fidelity, and scientific reasoning using task-specific evaluation metrics.
  • Deployment Architecture: Participants will develop systems to address one or both tasks and must submit the following deliverables (see later section for more details):
    • Their system responses to the task specific evaluation datasets in JSON-format,
    • Participants must submit their systems as reproducible containerized applications, packaged as Docker image(s) containing all software, dependencies, models, and configurations necessary to reproduce the responses provided in Deliverable 1 for each track/task. Submissions must be deployable within the Challenge’s AWS evaluation environment and may use either single-node or distributed architectures. Participants are responsible for providing all components required by their solutions, including retrieval, ranking, orchestration, and inference. They may use appropriate software frameworks, models, retrieval algorithms, vector databases, search technologies, and other tools, subject to the Challenge’s technical, security, licensing, and data-use requirements. Each submission must include deployment and execution instructions, along with estimated CPU, memory, GPU, and storage requirements. The Challenge Organizers reserve the right to execute submitted systems against the private test set in a secure, isolated environment to verify reproducibility and evaluate performance. More specific details and requirements for each task submission are available via the challenge management platform.
    • Links to a Docker repository and a GitHub repository containing source code, execution instructions, architecture documentation, dependencies, external resources, licenses, and reproduction instructions. Noncompliant or incomplete submissions may be deemed ineligible.

Submission Process

To be eligible to win a prize under a challenge, participants will need to submit their challenge solution for evaluation. Pay close attention to the submission dates posted for the challenge to ensure you submit your solution on time.

  1. Log into the Challenge Management Platform using your eRA Commons, or Login.gov credentials.
  2. Select Submit Solution from the upper navigation bar. This will bring you to the Submit Challenge Solution page.
  3. Under the Upload Solution Package section, click on the “Select File to Upload” button to upload your solution package (i.e., ZIP file containing the required submission materials, as specified by the challenge instructions).
  4. Under the Confirm Upload to GitHub section, select the checkbox to confirm that you have uploaded your solution to GitHub.
  5. Click on the “Submit” button at the bottom of the page to submit your solution.
  6. Note: If participating as a team lead or entity, the Point of Contact should submit the solution package on behalf of the team or entity.

Submission Requirements

System Requirements: For each Task, an evaluation dataset will be made available to registered participants prior to the Challenge closing. Participants must apply their systems to the evaluation dataset and generate the required system outputs and store in a submission file. This submission file must be provided in JSONL format and conform to the schema, file structure, naming conventions, and other technical specifications established by the Challenge organizers. Detailed submission specifications and instructions are available through the Challenge Management Platform.

To be considered complete and eligible for evaluation, each submission must include the following components:

  • Submission File: Participants must upload a challenge submission file containing their system-generated responses to the task evaluation datasets. Each task output file must conform to the required JSONL schema and other formatting, and validation requirements specified by the Challenge organizers. Submissions that do not conform to the required technical specifications may be deemed incomplete or ineligible for evaluation.
  • Containerized Application Submission: By January 15, 2027, participants must submit a Docker image (s) containing the solution, a description of their system and its components, together with sufficient technical documentation and instructions to enable NIH to reproduce, deploy, and evaluate the system within an NIH-managed AWS environment. All software, models, indexes, data resources, documentation, and other materials that comprise or support the submitted system must have been created or last modified on or before January 15, 2027. A link to the participant’s GitHub/Docker repository for the solution will be accepted. 
  • System Architecture Description: Submissions should describe the system architecture and components, including retrieval, ranking, embeddings, indexes, models, LLM inference, and any other software or data resources required to execute the system. Participants may design their systems as a single containerized application or as multiple containerized services, including distributed architectures in which different components run on separate compute nodes. Participants should also provide deployment and execution instructions, along with an estimate of the computational and storage resources required (e.g., CPU, memory, GPU, and storage). Submissions must explicitly disclose all training data sources, base model selections, architectural diagrams, post-processing heuristics, and any supplemental third-party datasets utilized. Furthermore, any specific adaptations applied to the pipeline or its components—such as fine-tuning, Reinforcement Learning from Human Feedback (RLHF), post-training enhancements, or other algorithmic optimizations—must be clearly documented.
  • Initial Submission Scope: Participants will not be required to submit executable container images, source code, or large data artifacts (e.g., embeddings, indexes, or models) as part of this initial submission. These materials will be requested during the evaluation phase in accordance with the Challenge procedures and technical guidance. 
  • AWS Infrastructure Responsibility: For evaluation, NIH will provision and manage the underlying AWS infrastructure, including compute, storage, networking, and orchestration. Participants will not be required to provision or manage AWS resources, Kubernetes clusters, or Terraform infrastructure. 
  • Deadline Verification: To support verification of the January 15, 2027, submission deadline, participants may be required to provide cryptographic hashes or other verifiable records for software, models, indexes, data artifacts, and other system components identified in their submission. When requested during the evaluation phase, the corresponding executable or data artifacts must match the versions documented and verified at the submission deadline. 
  • Model and Resource Restrictions: Submitted systems may use only publicly available, open-weight Large Language Models (LLMs), open-source software, and/or other publicly available computing resources. Systems may not invoke external APIs to proprietary frontier models (e.g., commercial hosted LLM endpoints) or rely on non-public, restricted-access software, models, datasets, or other resources during execution. Any pretrained/fine-tuned models used by the system must be included within the submitted Docker images.
  • Docker Image Execution and Network Restrictions
    • The challenge organizers reserve the right to run the submitted containerized application and verify system reproducibility, resource utilization, and compliance with challenge rules. Specifically:
    • External Calls: Submitted Docker Images will be executed in a completely isolated environment with all external network calls blocked. Limited access will be permitted to NLM-hosted services, such as the UMLS Terminology Services (UTS), and other challenge specific services. Details of these approved services and endpoints will be made available when the challenge is published. This zero-egress policy (excluding the permitted services) prevents both data leaks and runtime reliance on external computing resources, tools, or services.
  • Disclosure Responsibility: Participants must clearly identify all external data, models, software, and other third-party resources used in connection with their submission, including applicable licenses, terms of use, access restrictions, and attribution requirements. Participants are responsible for ensuring that they possess all rights, permissions, licenses, and authorizations necessary to use such resources in connection with the Challenge and to permit evaluation of their submissions by the Challenge organizers.
  • Proof of Execution: Prior to final submission, participants must execute a mandatory validation script directly on their Docker Image, which will be deployed on the specific AWS hardware corresponding to the challenge-designated AWS base image. This step is required to self-certify that the Docker Image compiles, runs smoothly, and correctly outputs data in the required format within the target environment. Submissions that fail this baseline execution test will not be evaluated. More details on the exact requirements are available on the Challenge Management Platform.
  • Disclaimer: Participants are solely responsible for ensuring their submissions and activities comply with all applicable laws, regulations, and policies. NIH cannot provide legal advice to third parties. Participants should consult their own legal counsel as necessary and appropriate.
  • Additional technical guidance, including submission formats, deployment specifications, resource limits, verification requirements, and evaluation procedures, will be provided to participants sufficiently in advance of the January 15, 2027, submission deadline.

Judging Criteria

Submissions in both tasks of the Challenge will be evaluated by a Technical Review Panel composed of subject matter experts in artificial intelligence, data science, information technology, and biomedical research. The Evaluation Panel will assess eligible submissions using the published evaluation criteria and provide feedback to the Judging Panel composed of federal employees.

The Judging Panel, with representation from NIH and other operating divisions, will review the Technical Review Panel's assessments alongside other relevant Challenge inputs, including compliance with applicable laws, regulations, and Challenge Rules.

Based on these factors, the Judging Panel will recommend final prize selections to the Award Approving Official. Final selections are made by the Award Approving Official, in their sole discretion, consistent with the published evaluation framework and applicable federal procedures. The Award Approving Official will review and approve final selections in accordance with applicable NIH requirements. A total of 4 winners is anticipated (the top 2 participants within each task).

Evaluation Criteria

To select winners for each Challenge Task, a challenge-specific evaluation dataset comprising hundreds of distinct information needs and their corresponding reference responses will be used.

Task 1 Ranking

Task 1 will use the following evaluation criteria to select winning submissions:

  • MRR (Mean Reciprocal Rank): Measures how highly a system ranks the relevant item for each question, rewarding systems that place the relevant item near the top of their ranked results. Higher MRR scores reflect a system's superior ability to correctly retrieve relevant information. 
  • Ease of Use: Assesses the effort required to deploy, configure, operate, and use the system in the standardized NIH cloud environment.
  • Compute Time: Measures the computational time required to generate responses.
  • Compute Effort: Measures the computational effort required to generate responses, including the resources and processing required by the system.
  • Memory Footprint: Measures the memory and other computational resources required to operate the system and generate responses.

MRR will serve as the primary metric for determining final system rankings and identifying the overall winning systems. In the event of a tie in MRR, the additional evaluation criteria listed above may be used to differentiate systems and determine the final ranking.

Task 2 Ranking

Task 2 will use the following evaluation criteria to select winning submissions:

  • Task 2 Quality Score (T2QS) composed of:
    • Grounded Nugget F1: Measures the precision and recall of correct, evidence-grounded nuggets.
    • Conflict / Controversy Coverage: Measures whether systems appropriately identify and characterize conflicting findings, scientific uncertainty, competing hypotheses, or other areas of disagreement when present.
  • Ease of Use: Assesses the effort required to deploy, configure, operate, and use the system in the standardized NIH cloud environment.
  • Compute Time: Measures the computational time required to generate responses.
  • Compute Effort: Measures the computational effort required to generate responses, including the resources and processing required by the system.
  • Memory Footprint: Measures the memory and other computational resources required to operate the system and generate responses.

T2QS will serve as the primary metric for determining final system rankings and identifying the overall winning systems. In the event of a tie, the additional evaluation criteria listed above may be used to differentiate systems and determine the final ranking.

Submission Packet Validation

Filtering & Evaluation Process

To efficiently manage computing resources and ensure the scientific integrity of the results, submissions will undergo a multi-stage validation, filtering, and dual-layer evaluation process.

1. Timeline and Initial Validation Phase

  • Validation Set: A small validation dataset is provided for each track to help participants develop and test their pipeline architectures:
    • Task 1: Consists of narrow, objective questions along with their corresponding reference responses.
    • Task 2: Consists of open-ended questions, evidence-backed reference responses with supporting PM/PMC items.
  • Mandatory Local Sanity Check: Prior to submitting their containerized systems, participants must execute a mandatory baseline validation script locally against the validation data. This ensures the Docker Image(s) compiles properly and outputs data in the strict, required format (e.g., JSON).

2. Anti-Gaming Probes and Adversarial Security

  • Trojan Horse Detection: To prevent “memorization gaming”—where a submitted Docker Image is hardcoded with pre-computed answers mapping directly to known validation or public study data—NLM Challenge Team will inject entirely unseen, private adversarial probes and information needs into the final test phase.
  • Behavioral Auditing & Adversarial Security Post-competition: The Challenge organizers reserve the right to execute participant Docker Images over private queries and synthetic variables. Systems that exhibit a drastic drop in evaluation of metrics or display hardcoded outputs tailored specifically to PubMed/PMC data will be flagged for anomalous “Trojan horse” behavior. Upon judicial review by the panel, any system found to be using pre-computed or non-generalized outputs will be disqualified.

Additional Information

Participants are encouraged to join one of the following information sessions to learn more and ask questions:

Point of Contact

Questions can be directed to the following email mailbox for this challenge at [email protected].

This page last reviewed on