Discover NAIRR: A New National Compute Resource for AI Research
Transcript
Hi everyone and welcome to today's Lunch and Learn on Discover Nair, a new national compute resource for AI research. My name is Elizabeth Kwon and I am an Embedded Research Computing Specialist on the Research Computing Services team. The goal of today's Lunch and Learn is to increase awareness of available resources like Nair and to help everyone better understand how to use them to advance research and educational initiatives at Columbia.
So, you know, for those who are not familiar with our group, Research Computing Services, we are a team within CUIT's Research Services Department and we provide technology resources to support Columbia's researchers and faculty with their computing needs. We help researchers identify and access resources that might fit their research efforts such as Nair and we also collaborate closely with our partner team, High Performance Computing, to offer shared HPC resources and related services across the university. Looking at our, you know, just our general service portfolio, our Research Services team offers various research compute resources for data transfer, storage, analysis, data collaboration, high performance computing, scientific software like Lab Archives and GraphPad Prism, access to national compute resources, and our Embedded Research Computing Support Service which provides tailored support to the needs of a research group or a center.
So, you know, if you are from a school, department, IT, or research-facing group and you are interested in learning more about our resources, feel free to get in touch with us at rcs.columbia.edu. Taking a quick glance at the agenda that I have planned for today, we're going to be talking about, you know, what is Nair, what does this mean for universities and researchers, what are some existing projects that are on Nair, Nair resources, and I'll go into a couple of them, for example, eligibility requirements, what are some current opportunities, and additional resources that you can leverage, and then Q&A. And then also, you know, as part, as we go through this, you know, we'll be throwing in some links in the chat, so, you know, if you, you know, don't feel free, feel free that, you know, you will be sharing all of this information afterwards as well. So, what is Nair? So, Nair stands for the National Artificial Intelligence Research Resource, and it's a shared national research infrastructure program that's designed to connect U.S. researchers and educators to essential AI resources, including computation, data, software, various models, and training materials.
Its primary goal is to bridge the gap in access to these tools to advance research, discovery, and innovation in AI. So, you can think of it kind of like as a pool of resources and tools that scientists and engineers across the U.S. can dip into and try a resource out. Nair is led by the U.S. National Science Foundation, NSF, in partnership with 13 other federal agencies and 28 non-governmental partners.
This collective provides the necessary AI resources to be shared on a national level for the research and education community, and there are currently over 500 research projects using Nair across 49 U.S. states. So, you know, Nair represents a unique opportunity for universities to democratizing access to advanced computing, data, and AI tools that were previously only available to those in the tech industry. So, universities can conduct research and compute on a level comparable to some leading technology companies, and this will also enable broader participation in cutting-edge AI research across disciplines.
Some key objectives of the Nair program is that it aims to accelerate AI and AI-powered discovery and innovation to help researchers across all disciplines use AI as a catalyst for new insights and research breakthroughs. It seeks to expand the AI workforce also by training the next generation of AI researchers and educators to ensure that students and professionals alike can gain hands-on experience with advanced AI tools and resources. Nair also works to increase capacity, integration, and use of these world-class public and private sector AI resources to create a more connected and accessible AI research ecosystem.
And then finally, Nair's focus on advancing AI interpretability. So, security, trust, these are all kind of critical for building responsible and transparent AI systems that serve the public good. Note that Nair resources are available at no charge, lowering the barrier to experiment with larger models, but at the same time the application process as a result can be a bit competitive, particularly for projects that require significant compute power.
So, there are different types of opportunities researchers and educators alike can take advantage of depending on their specific use cases and computational needs. Note that the Nair pilot currently provides infrastructure resources for research engagements, so these are the resources themselves, and they do not provide funding for the research work. So, looking at some of the existing use cases, researchers across the U.S. are currently using Nair in the following research fields.
They're building foundation models for aquatic sciences, training vision language models for agricultural resilience, generative models of neural images for assessing Alzheimer's disease, and so many more. What used to be only utilized in traditional STEM disciplines like physics, chemistry, engineering, it's now being adopted across, you know, the humanities, social sciences, and other non-STEM focused fields. So, researchers are really using high-performance computing to analyze large-scale social data, simulate economic systems, and support creative work in digital design.
We're also seeing a strong convergence between AI and HPC with the growing demand for large-scale data analysis and model training. Many research groups across the U.S. are integrating machine learning and generative AI techniques into their workflows. This shift is really highlighting how important it is to have accessible, scalable, and well-supported compute infrastructure, which is exactly what programs like Nair are aiming to provide.
Nair includes over 40 different resources across different providers, so this includes, you know, pre-trained models, data sets, and relevant data and compute platforms. Some of the existing resources include AWS, Databricks, Google Cloud, Hugging Face Spaces, Indiana University's JetStream 2, NVIDIA, and a bunch more. If you have possibly attended a previous Lunch and Learn, I held on Access, which is another NSF-led program.
Many of Nair's resources are actually part of the Access program as well, and our team, Research Computing Services, we can help provide guidance to researchers who are interested in navigating access to these national resources. Resources on Nair have varying compute systems, hardware, they can have different GPU accelerators, different data storage systems, software libraries, etc., so there's a lot to leverage. So, you know, let's take a look at some of these resources a bit deeper.
So, the first one, Hugging Face Spaces. So, if you're not familiar, Hugging Face is a leading open-source AI platform that supports collaboration, model sharing, and deployment of machine learning applications. Through Hugging Face Spaces, researchers can host machine learning applications directly on their Hugging Face profiles to showcase their projects and collaborate with others.
Nair provides support in upgrading hardware resources through Hugging Face Spaces, such as, GPUs or TPUs, and this will help run more demanding models and compute-intensive experiments. Hugging Face also offers inference endpoints, which are a managed service to deploy your AI models, like transformers, sentence transformers, diffusion models, etc., to production. So, these are especially useful for researchers who want to annotate, augment, generate data sets, or create interactive model demos.
And in addition, Hugging Face provides a no-code tool called Autotrain, which is a platform designed to make training, cutting-edge models more simple and more accessible. So, whether you're working with, you know, natural language processing, computer vision, or even, like, with tabular data, what makes Autotrain kind of stand out is that it builds on the advanced frameworks that already created by Hugging Face's own team. So, this allows anyone to leverage, you know, state-of-the-art AI and machine learning without needing that deep technical expertise.
Overall, Hugging Face Spaces offer an excellent opportunity for researchers to prototype, deploy, and share AI models in a collaborative, scalable environment. And then, another thing to note, Hugging Face is GDPR compliant and also SOC2 Type 2 certified. JetStream 2 is a flexible, user-friendly cloud computing environment that was designed and housed in Indiana University.
JetStream 2 is one of the many resources part of NAIR and other national programs, and it's primarily designed for researchers with minimal high-performance compute experience to those interested in exploring cloud native resources. JetStream 2 provides an on-demand virtual cloud environment where users can create and manage their own virtual machines without queues or runtime limits. It currently supports over 2,500 active users on their resource with a 8 petaflops cloud computing system and 17.2 petabytes in storage.
It has 384 compute nodes of AMD Milan EPYC CPUs with 128 cores and 512 gigabytes of RAM and 90 GPU nodes with 4 NVIDIA A100s each and 32 large memory nodes. In partnership with NAIR, JetStream 2 also has an AI Fellows Program. It's a professional development program for early career and established researchers to learn how to use JetStream 2 for AI-driven research.
And this fellowship has four main components. First, project consulting. Fellows will receive one-on-one support to help overcome technical or logistical challenges that they have, and also onboarding to, you know, optimizing their use of JetStream 2 GPU capabilities.
Second, training. Participants are going to engage in practical sessions using JetStream 2's AI Workbench and other training modules focused on AI methods and tools. Priority access to high-end GPUs.
This means that AI fellows are going to get priority access to JetStream 2's powerful GPU nodes, allowing them to accelerate their research and achieve results faster. And then finally, community contribution. Each fellow will work with the JetStream 2 team to share lessons learned, you know, provide feedback, sort of creating, you know, not only helping them with their research, but this will also create future resources and examples that will benefit other researchers in the broader community.
So, you know, overall, this program not only advances participants' own research, but also strengthens the entire AI research ecosystem supported by JetStream 2. And in terms of like security, JetStream 2 security controls include, you know, defense-in-depth approach that relies on using user-managed security groups, firewalls, you know, host-based firewalls, in addition to Indiana University's policies and infrastructure security controls. So now, OpenAI. Through Nair, OpenAI is providing a million dollars in OpenAI API credits for research focused on safe, secure, and trustworthy AI for the public benefit.
Researchers are encouraged to review the OpenAI pricing information to, you know, determine what's the estimated scale of their project needs. But essentially, OpenAI API credits are going to be allocated in 5, 25, 50, and 100k units per PI proposal. OpenAI is providing researchers access to powerful AI models for building applications and conducting studies in areas like language, vision, reasoning, and, you know, depending on the researcher's goals and computing needs, they can leverage OpenAI's frontier models.
So, these are OpenAI's most advanced and capable models, you know, including the GPT-5 family. There's specialized models. So, these are purpose-built models for specific domains.
You can think of like Sora for video generation, GPT image, and Dolly for image creation. There's deep research models for complex analysis, and open weight models to allow researchers to freely modify and deploy for experimentation and scalable applications. So, you know, with OpenAI, researchers can generate and analyze text using powerful language models.
They can process and interpret images or even create new ones with built-in vision and image generation tools, work with audio by transcribing, analyzing, generating speech, build intelligent agents that can reason, use tools, perform complex tasks, extract structured data from model outputs, and, you know, customize models through fine-tuning and evaluations to fit specific applications or domains. So, you know, now that we've gone through some of the resources, who is eligible to use these resources? So, NARE is open to qualified U.S.-based researchers and educators from academic institutions such as Columbia University, in addition to nonprofits, government agencies, federally funded R&D centers, eligible startups, or small businesses that have federal support. So, note that, you know, certain specific resources that are offered through NARE may have additional eligibility requirements and proposal requirements that will align with those constraints.
But, you know, as you're going through the process of submitting an application, it will become clear what those constraints might possibly be. Additionally, NARE resources on the default opportunities are for non-sensitive data use only. So, you know, research projects must be handling non-sensitive data as this is an open science program, fundamentally for fostering, you know, open science, open public research, unless they're using what's called NARE Secure.
And this is something I'll talk about shortly after, but if your research does involve sensitive data, I also highly recommend taking a look at alternative resources that our RCS team supports, such as the Secure Data Enclave or the Columbia Data Platform. So, the first opportunity I'll talk about is the research resources call. So, the NARE open call for allocation proposals invites projects that are aligned with a broad set of cross-cutting and domain-specific focus areas, which I've listed here in this slide.
You know, so some cross-cutting priorities include advancing AI methods for scientific discovery, developing open-source foundation models, integrating simulations, AI, and like other kinds of like experimental data. So, this call also emphasizes, you know, responsible AI education, public engagement with scientific data, and privacy-preserving methods for sensitive data. Domain-specific areas target high-impact challenges, such as like AI safety, healthcare innovation, infrastructure resilience, advanced manufacturing, climate, and environmental solutions.
And then, all approved NARE projects are awarded for a 12-month duration. The second call is a classroom educator resources call. So, this call is open to proposals by U.S.-based educators and researchers who are teaching undergraduate or graduate courses, or shorter-duration training sessions to U.S.-based students that include subject matter in artificial intelligence, and, you know, require those students to use advanced computational resources as part of their coursework.
So, courses from any discipline are eligible for this program. However, courses and trainings must, you know, not allow participants who are not U.S.-based. So, you know, to request access to NARE, you know, classroom resources, individuals must prepare a description that's no longer than three pages. You know, you have to describe your course and requirements for computational resources available to this program.
So, it's a fairly straightforward process. And then, the NARE startup project call, which is a new resource call that was recently launched. So, this was created to help researchers get started with NARE resources and to ensure that they're fully prepared to use them effectively.
So, startup requests is generally a smaller, lighter-weight proposal that provides short-term access to resources. Decisions are made quickly, usually within about two weeks, making it an ideal entry point for researchers who are either new to the NARE pilot or, you know, need to estimate the requirements before submitting a full-scale research proposal for NARE. So, you know, startup projects can be used for several purposes.
For example, researchers might want to learn how federally-coordinated computing and data resources are allocated, if they're just doing some, like, small-scale building and testing for a software environment or workflow, kind of like as a proof of concept, if they are running a scaling study to determine the right amount of compute or data needed for a larger project. So, you know, based off of, you know, these examples, you know, each startup project is awarded a smaller duration. It's a three-month duration, giving researchers time to experiment, gather information they need, and then from there, they can work and either fill out a researcher resources call or the classroom educator resources call.
So, yeah. So, after starting a startup, they're encouraged to submit longer-term research requests. And, you know, these startup projects kind of serve as, like, a stepping stone, helping researchers transition from exploration to full-scale AI research.
And then Nair Secure. So, this is one of the other pilot programs within the Nair Initiative that is co-led by the Department of Energy and the National Institutes of Health and NIH, and it's designed to enable AI research involving sensitive data. So, any kind of data that requires extra layers of protection and secure computing environments.
This working group had, you know, three main goals when they create the Nair Secure resource, and it's, you know, to refine the requirements and infrastructure design patterns needed for future Nair Secure resources, including ways to integrate diverse and sensitive datasets safely. They want to explore new methods for combining data that still preserves privacy and security to ensure that research can still move forward responsibly, and to, you know, identify key research and training use cases that will help shape development of the next generation of Nair Secure systems and resources. So, you know, note that, you know, a limited number of resource providers that are on the Nair resource list can accommodate sensitive data.
So, always make sure to double check when applying or submitting a request that the resource that you're requesting for is part of the Nair Secure program if you are handling sensitive data. You know, and then looking at two demonstrations of what, you know, Nair Secure projects already have been doing in terms of research, we have Althena, which uses the latest open LLM models to power biomedical research, discovery, and understanding on scientific literature for secure scalable discovery. And there's another project that focuses on democratizing AI for cancer research through privacy preserving synthetic data generation.
And, you know, they're helping identify cancer cases while protecting patient privacy in this process. So, you know, these two projects are examples that demonstrate how Nair Secure is paving the way for responsible privacy conscious AI research that can still leverage the power of real world data. And to apply, you know, you can easily just navigate to the Nair pilot website, nairpilot.org, go to the available submission opportunities page that's at the following link from the current available opportunity.
So, you're going to be seeing the research resources, classroom educator, startup, Nair Secure. From there, select the opportunity of interest, you can then fill out the request form with the respective information that's requested, and then talk about like your project, your resource needs. You know, if there are other people working on the project with you, you would add them here as well and just submit your project proposal requests.
So note that, you know, proposals, they are reviewed on a ongoing monthly cycle. So typically requests that are submitted by the 15th of the month will be reviewed and have the outcome decided by the end of the following month. And then again, just to reiterate, projects are awarded for 12 month durations, unless you're going with a startup request, which is a three month duration.
Going over just, you know, some other general expectations of Nair projects. Again, you know, because Nair is built on the principle of openness, all project results should be publishable and publicly available whenever possible. You know, going back to Nair's goals, they want to generate open, pre-competitive research outcomes that can benefit the broader scientific community.
So, you know, also some things to consider are if your project might lead to commercializable intellectual property, you might need to discuss this early on with your assigned resource provider and confirm that it fits in with their Nair guidelines. Projects are also, you know, again, intended to be completed within 12 months. So if you do require more time, you'll need to request additional support either through Nair or through other allocation programs.
And, you know, transparency. Each project's PI, affiliation, title, and also like their project abstract can be listed on the Nair website. And this is to promote transparency and visibility on the Nair effort.
And then community and reporting. So, PIs will also join the Nair researcher community mailing list, and they will also be expected to provide short updates at one in six months, and then a brief final report at the end of the project. And finally, Nair is going to collect, you know, usage on data and on allocated resources to help evaluate their impact and the impact of their program.
So again, you know, due to the competitive nature of requests, projects that don't use their allocations efficiently may see those resources reduced or might get a, you know, a touch base from the reviewers and technical team from Nair about, you know, just what people are doing with their research and whether, you know, there's some additional accommodations or updates that can be made. So, you know, in short, the Nair program emphasizes openness, accountability, effective use of shared computing resources to support high-impact AI research across the U.S. And, you know, in addition to Nair, there's also Access, which is the other resource that I mentioned earlier. You know, if you're thinking about what else is out there, I do highly recommend also checking out Access or asking our team about it.
So Access stands for the Advanced Cyber Infrastructure Coordination and Ecosystem Services and Support. It's an optional, I know. It formally was known by its predecessing name XSEED, and it's a program that's funded by the National Science Foundation, NSF, to help researchers and educators utilize a nationwide collection of supercomputing systems.
So it's, you know, another library of resources. And, you know, there are also some overlap in terms of the resource providers and some new ones as well on Access. But, you know, they provide a wide range of resources and services such as, you know, compute resources, so high-performance computing clusters that are CPU and GPU-based, storage resources, so, you know, data storage systems for storing and managing large amounts of data, specialized support, so additional tools, resources to help streamline resources, and, you know, Access resources are also available at no cost.
And, you know, as a result, you know, their application process can also be competitive depending on the compute resources needed, and they have a kind of like a sliding scale of allocation sizes to address small and large projects. And, you know, Access also has researcher and educator-focused resource allocations. The main difference, I will say, is that Access is mainly, is just for a non-sensitive data use, while NAIR has the NAIR Secure resource allocation option.
So, you know, if you're interested in either, you know, feel free to touch base with us. And with that, you know, I am at the end of the presentation. If you have any additional, you know, research computing questions, feel free to contact us at the following information.
You can also check out our site, which has more information about different services that we provide and, you know, goes through our service portfolio. But other than that, you know, I hope, you know, if there are any researchers or educators on call that you reach out and take advantage of this great opportunity. Thank you.
Thanks, Liz. There's been a little bit of discussion in the chat about the eligibility, and I know the language in both NAIR and I think Access is like U.S.-based research. Does that mean that someone like an international postdoc that is at Columbia would be eligible if their PI signed off? Right, you know, so that's a really good question.
So, you know, if you know, if this does this mean that only like Americans or, you know, U.S. citizens can participate and use the resources? No. So, you know, citizenship does not impact general eligibility. So the requirement is just that you must be affiliated with a U.S.-based institution and be working in the U.S. or conducting your research in the U.S. I hope that answers that question.
Let me check the chat. Yeah, I think that definitely is good. And of course, feel free to speak up if anybody has any questions.
Sorry, I have a follow-up on that. Go ahead. Yeah.
So I was wondering because researchers means that it can be a student or a postdoc as well, right? So they are also eligible to write a proposal or it has to be a faculty or a research scientist. And which particular opportunity are you talking about, the researcher resources or the classroom educator resources? Researcher resources, just for my research project. So for that, you would have to be at a minimum a grad student or higher.
So the undergrad opportunities are through the classroom resources opportunity, which need to be done in parallel with either a professor or some educator that's leading it and then allowing their students to then use the resources, if that makes sense. Yeah, that's very helpful. Thank you.
So, you know, for those who are not familiar with our group, Research Computing Services, we are a team within CUIT's Research Services Department and we provide technology resources to support Columbia's researchers and faculty with their computing needs. We help researchers identify and access resources that might fit their research efforts such as Nair and we also collaborate closely with our partner team, High Performance Computing, to offer shared HPC resources and related services across the university. Looking at our, you know, just our general service portfolio, our Research Services team offers various research compute resources for data transfer, storage, analysis, data collaboration, high performance computing, scientific software like Lab Archives and GraphPad Prism, access to national compute resources, and our Embedded Research Computing Support Service which provides tailored support to the needs of a research group or a center.
So, you know, if you are from a school, department, IT, or research-facing group and you are interested in learning more about our resources, feel free to get in touch with us at rcs.columbia.edu. Taking a quick glance at the agenda that I have planned for today, we're going to be talking about, you know, what is Nair, what does this mean for universities and researchers, what are some existing projects that are on Nair, Nair resources, and I'll go into a couple of them, for example, eligibility requirements, what are some current opportunities, and additional resources that you can leverage, and then Q&A. And then also, you know, as part, as we go through this, you know, we'll be throwing in some links in the chat, so, you know, if you, you know, don't feel free, feel free that, you know, you will be sharing all of this information afterwards as well. So, what is Nair? So, Nair stands for the National Artificial Intelligence Research Resource, and it's a shared national research infrastructure program that's designed to connect U.S. researchers and educators to essential AI resources, including computation, data, software, various models, and training materials.
Its primary goal is to bridge the gap in access to these tools to advance research, discovery, and innovation in AI. So, you can think of it kind of like as a pool of resources and tools that scientists and engineers across the U.S. can dip into and try a resource out. Nair is led by the U.S. National Science Foundation, NSF, in partnership with 13 other federal agencies and 28 non-governmental partners.
This collective provides the necessary AI resources to be shared on a national level for the research and education community, and there are currently over 500 research projects using Nair across 49 U.S. states. So, you know, Nair represents a unique opportunity for universities to democratizing access to advanced computing, data, and AI tools that were previously only available to those in the tech industry. So, universities can conduct research and compute on a level comparable to some leading technology companies, and this will also enable broader participation in cutting-edge AI research across disciplines.
Some key objectives of the Nair program is that it aims to accelerate AI and AI-powered discovery and innovation to help researchers across all disciplines use AI as a catalyst for new insights and research breakthroughs. It seeks to expand the AI workforce also by training the next generation of AI researchers and educators to ensure that students and professionals alike can gain hands-on experience with advanced AI tools and resources. Nair also works to increase capacity, integration, and use of these world-class public and private sector AI resources to create a more connected and accessible AI research ecosystem.
And then finally, Nair's focus on advancing AI interpretability. So, security, trust, these are all kind of critical for building responsible and transparent AI systems that serve the public good. Note that Nair resources are available at no charge, lowering the barrier to experiment with larger models, but at the same time the application process as a result can be a bit competitive, particularly for projects that require significant compute power.
So, there are different types of opportunities researchers and educators alike can take advantage of depending on their specific use cases and computational needs. Note that the Nair pilot currently provides infrastructure resources for research engagements, so these are the resources themselves, and they do not provide funding for the research work. So, looking at some of the existing use cases, researchers across the U.S. are currently using Nair in the following research fields.
They're building foundation models for aquatic sciences, training vision language models for agricultural resilience, generative models of neural images for assessing Alzheimer's disease, and so many more. What used to be only utilized in traditional STEM disciplines like physics, chemistry, engineering, it's now being adopted across, you know, the humanities, social sciences, and other non-STEM focused fields. So, researchers are really using high-performance computing to analyze large-scale social data, simulate economic systems, and support creative work in digital design.
We're also seeing a strong convergence between AI and HPC with the growing demand for large-scale data analysis and model training. Many research groups across the U.S. are integrating machine learning and generative AI techniques into their workflows. This shift is really highlighting how important it is to have accessible, scalable, and well-supported compute infrastructure, which is exactly what programs like Nair are aiming to provide.
Nair includes over 40 different resources across different providers, so this includes, you know, pre-trained models, data sets, and relevant data and compute platforms. Some of the existing resources include AWS, Databricks, Google Cloud, Hugging Face Spaces, Indiana University's JetStream 2, NVIDIA, and a bunch more. If you have possibly attended a previous Lunch and Learn, I held on Access, which is another NSF-led program.
Many of Nair's resources are actually part of the Access program as well, and our team, Research Computing Services, we can help provide guidance to researchers who are interested in navigating access to these national resources. Resources on Nair have varying compute systems, hardware, they can have different GPU accelerators, different data storage systems, software libraries, etc., so there's a lot to leverage. So, you know, let's take a look at some of these resources a bit deeper.
So, the first one, Hugging Face Spaces. So, if you're not familiar, Hugging Face is a leading open-source AI platform that supports collaboration, model sharing, and deployment of machine learning applications. Through Hugging Face Spaces, researchers can host machine learning applications directly on their Hugging Face profiles to showcase their projects and collaborate with others.
Nair provides support in upgrading hardware resources through Hugging Face Spaces, such as, GPUs or TPUs, and this will help run more demanding models and compute-intensive experiments. Hugging Face also offers inference endpoints, which are a managed service to deploy your AI models, like transformers, sentence transformers, diffusion models, etc., to production. So, these are especially useful for researchers who want to annotate, augment, generate data sets, or create interactive model demos.
And in addition, Hugging Face provides a no-code tool called Autotrain, which is a platform designed to make training, cutting-edge models more simple and more accessible. So, whether you're working with, you know, natural language processing, computer vision, or even, like, with tabular data, what makes Autotrain kind of stand out is that it builds on the advanced frameworks that already created by Hugging Face's own team. So, this allows anyone to leverage, you know, state-of-the-art AI and machine learning without needing that deep technical expertise.
Overall, Hugging Face Spaces offer an excellent opportunity for researchers to prototype, deploy, and share AI models in a collaborative, scalable environment. And then, another thing to note, Hugging Face is GDPR compliant and also SOC2 Type 2 certified. JetStream 2 is a flexible, user-friendly cloud computing environment that was designed and housed in Indiana University.
JetStream 2 is one of the many resources part of NAIR and other national programs, and it's primarily designed for researchers with minimal high-performance compute experience to those interested in exploring cloud native resources. JetStream 2 provides an on-demand virtual cloud environment where users can create and manage their own virtual machines without queues or runtime limits. It currently supports over 2,500 active users on their resource with a 8 petaflops cloud computing system and 17.2 petabytes in storage.
It has 384 compute nodes of AMD Milan EPYC CPUs with 128 cores and 512 gigabytes of RAM and 90 GPU nodes with 4 NVIDIA A100s each and 32 large memory nodes. In partnership with NAIR, JetStream 2 also has an AI Fellows Program. It's a professional development program for early career and established researchers to learn how to use JetStream 2 for AI-driven research.
And this fellowship has four main components. First, project consulting. Fellows will receive one-on-one support to help overcome technical or logistical challenges that they have, and also onboarding to, you know, optimizing their use of JetStream 2 GPU capabilities.
Second, training. Participants are going to engage in practical sessions using JetStream 2's AI Workbench and other training modules focused on AI methods and tools. Priority access to high-end GPUs.
This means that AI fellows are going to get priority access to JetStream 2's powerful GPU nodes, allowing them to accelerate their research and achieve results faster. And then finally, community contribution. Each fellow will work with the JetStream 2 team to share lessons learned, you know, provide feedback, sort of creating, you know, not only helping them with their research, but this will also create future resources and examples that will benefit other researchers in the broader community.
So, you know, overall, this program not only advances participants' own research, but also strengthens the entire AI research ecosystem supported by JetStream 2. And in terms of like security, JetStream 2 security controls include, you know, defense-in-depth approach that relies on using user-managed security groups, firewalls, you know, host-based firewalls, in addition to Indiana University's policies and infrastructure security controls. So now, OpenAI. Through Nair, OpenAI is providing a million dollars in OpenAI API credits for research focused on safe, secure, and trustworthy AI for the public benefit.
Researchers are encouraged to review the OpenAI pricing information to, you know, determine what's the estimated scale of their project needs. But essentially, OpenAI API credits are going to be allocated in 5, 25, 50, and 100k units per PI proposal. OpenAI is providing researchers access to powerful AI models for building applications and conducting studies in areas like language, vision, reasoning, and, you know, depending on the researcher's goals and computing needs, they can leverage OpenAI's frontier models.
So, these are OpenAI's most advanced and capable models, you know, including the GPT-5 family. There's specialized models. So, these are purpose-built models for specific domains.
You can think of like Sora for video generation, GPT image, and Dolly for image creation. There's deep research models for complex analysis, and open weight models to allow researchers to freely modify and deploy for experimentation and scalable applications. So, you know, with OpenAI, researchers can generate and analyze text using powerful language models.
They can process and interpret images or even create new ones with built-in vision and image generation tools, work with audio by transcribing, analyzing, generating speech, build intelligent agents that can reason, use tools, perform complex tasks, extract structured data from model outputs, and, you know, customize models through fine-tuning and evaluations to fit specific applications or domains. So, you know, now that we've gone through some of the resources, who is eligible to use these resources? So, NARE is open to qualified U.S.-based researchers and educators from academic institutions such as Columbia University, in addition to nonprofits, government agencies, federally funded R&D centers, eligible startups, or small businesses that have federal support. So, note that, you know, certain specific resources that are offered through NARE may have additional eligibility requirements and proposal requirements that will align with those constraints.
But, you know, as you're going through the process of submitting an application, it will become clear what those constraints might possibly be. Additionally, NARE resources on the default opportunities are for non-sensitive data use only. So, you know, research projects must be handling non-sensitive data as this is an open science program, fundamentally for fostering, you know, open science, open public research, unless they're using what's called NARE Secure.
And this is something I'll talk about shortly after, but if your research does involve sensitive data, I also highly recommend taking a look at alternative resources that our RCS team supports, such as the Secure Data Enclave or the Columbia Data Platform. So, the first opportunity I'll talk about is the research resources call. So, the NARE open call for allocation proposals invites projects that are aligned with a broad set of cross-cutting and domain-specific focus areas, which I've listed here in this slide.
You know, so some cross-cutting priorities include advancing AI methods for scientific discovery, developing open-source foundation models, integrating simulations, AI, and like other kinds of like experimental data. So, this call also emphasizes, you know, responsible AI education, public engagement with scientific data, and privacy-preserving methods for sensitive data. Domain-specific areas target high-impact challenges, such as like AI safety, healthcare innovation, infrastructure resilience, advanced manufacturing, climate, and environmental solutions.
And then, all approved NARE projects are awarded for a 12-month duration. The second call is a classroom educator resources call. So, this call is open to proposals by U.S.-based educators and researchers who are teaching undergraduate or graduate courses, or shorter-duration training sessions to U.S.-based students that include subject matter in artificial intelligence, and, you know, require those students to use advanced computational resources as part of their coursework.
So, courses from any discipline are eligible for this program. However, courses and trainings must, you know, not allow participants who are not U.S.-based. So, you know, to request access to NARE, you know, classroom resources, individuals must prepare a description that's no longer than three pages. You know, you have to describe your course and requirements for computational resources available to this program.
So, it's a fairly straightforward process. And then, the NARE startup project call, which is a new resource call that was recently launched. So, this was created to help researchers get started with NARE resources and to ensure that they're fully prepared to use them effectively.
So, startup requests is generally a smaller, lighter-weight proposal that provides short-term access to resources. Decisions are made quickly, usually within about two weeks, making it an ideal entry point for researchers who are either new to the NARE pilot or, you know, need to estimate the requirements before submitting a full-scale research proposal for NARE. So, you know, startup projects can be used for several purposes.
For example, researchers might want to learn how federally-coordinated computing and data resources are allocated, if they're just doing some, like, small-scale building and testing for a software environment or workflow, kind of like as a proof of concept, if they are running a scaling study to determine the right amount of compute or data needed for a larger project. So, you know, based off of, you know, these examples, you know, each startup project is awarded a smaller duration. It's a three-month duration, giving researchers time to experiment, gather information they need, and then from there, they can work and either fill out a researcher resources call or the classroom educator resources call.
So, yeah. So, after starting a startup, they're encouraged to submit longer-term research requests. And, you know, these startup projects kind of serve as, like, a stepping stone, helping researchers transition from exploration to full-scale AI research.
And then Nair Secure. So, this is one of the other pilot programs within the Nair Initiative that is co-led by the Department of Energy and the National Institutes of Health and NIH, and it's designed to enable AI research involving sensitive data. So, any kind of data that requires extra layers of protection and secure computing environments.
This working group had, you know, three main goals when they create the Nair Secure resource, and it's, you know, to refine the requirements and infrastructure design patterns needed for future Nair Secure resources, including ways to integrate diverse and sensitive datasets safely. They want to explore new methods for combining data that still preserves privacy and security to ensure that research can still move forward responsibly, and to, you know, identify key research and training use cases that will help shape development of the next generation of Nair Secure systems and resources. So, you know, note that, you know, a limited number of resource providers that are on the Nair resource list can accommodate sensitive data.
So, always make sure to double check when applying or submitting a request that the resource that you're requesting for is part of the Nair Secure program if you are handling sensitive data. You know, and then looking at two demonstrations of what, you know, Nair Secure projects already have been doing in terms of research, we have Althena, which uses the latest open LLM models to power biomedical research, discovery, and understanding on scientific literature for secure scalable discovery. And there's another project that focuses on democratizing AI for cancer research through privacy preserving synthetic data generation.
And, you know, they're helping identify cancer cases while protecting patient privacy in this process. So, you know, these two projects are examples that demonstrate how Nair Secure is paving the way for responsible privacy conscious AI research that can still leverage the power of real world data. And to apply, you know, you can easily just navigate to the Nair pilot website, nairpilot.org, go to the available submission opportunities page that's at the following link from the current available opportunity.
So, you're going to be seeing the research resources, classroom educator, startup, Nair Secure. From there, select the opportunity of interest, you can then fill out the request form with the respective information that's requested, and then talk about like your project, your resource needs. You know, if there are other people working on the project with you, you would add them here as well and just submit your project proposal requests.
So note that, you know, proposals, they are reviewed on a ongoing monthly cycle. So typically requests that are submitted by the 15th of the month will be reviewed and have the outcome decided by the end of the following month. And then again, just to reiterate, projects are awarded for 12 month durations, unless you're going with a startup request, which is a three month duration.
Going over just, you know, some other general expectations of Nair projects. Again, you know, because Nair is built on the principle of openness, all project results should be publishable and publicly available whenever possible. You know, going back to Nair's goals, they want to generate open, pre-competitive research outcomes that can benefit the broader scientific community.
So, you know, also some things to consider are if your project might lead to commercializable intellectual property, you might need to discuss this early on with your assigned resource provider and confirm that it fits in with their Nair guidelines. Projects are also, you know, again, intended to be completed within 12 months. So if you do require more time, you'll need to request additional support either through Nair or through other allocation programs.
And, you know, transparency. Each project's PI, affiliation, title, and also like their project abstract can be listed on the Nair website. And this is to promote transparency and visibility on the Nair effort.
And then community and reporting. So, PIs will also join the Nair researcher community mailing list, and they will also be expected to provide short updates at one in six months, and then a brief final report at the end of the project. And finally, Nair is going to collect, you know, usage on data and on allocated resources to help evaluate their impact and the impact of their program.
So again, you know, due to the competitive nature of requests, projects that don't use their allocations efficiently may see those resources reduced or might get a, you know, a touch base from the reviewers and technical team from Nair about, you know, just what people are doing with their research and whether, you know, there's some additional accommodations or updates that can be made. So, you know, in short, the Nair program emphasizes openness, accountability, effective use of shared computing resources to support high-impact AI research across the U.S. And, you know, in addition to Nair, there's also Access, which is the other resource that I mentioned earlier. You know, if you're thinking about what else is out there, I do highly recommend also checking out Access or asking our team about it.
So Access stands for the Advanced Cyber Infrastructure Coordination and Ecosystem Services and Support. It's an optional, I know. It formally was known by its predecessing name XSEED, and it's a program that's funded by the National Science Foundation, NSF, to help researchers and educators utilize a nationwide collection of supercomputing systems.
So it's, you know, another library of resources. And, you know, there are also some overlap in terms of the resource providers and some new ones as well on Access. But, you know, they provide a wide range of resources and services such as, you know, compute resources, so high-performance computing clusters that are CPU and GPU-based, storage resources, so, you know, data storage systems for storing and managing large amounts of data, specialized support, so additional tools, resources to help streamline resources, and, you know, Access resources are also available at no cost.
And, you know, as a result, you know, their application process can also be competitive depending on the compute resources needed, and they have a kind of like a sliding scale of allocation sizes to address small and large projects. And, you know, Access also has researcher and educator-focused resource allocations. The main difference, I will say, is that Access is mainly, is just for a non-sensitive data use, while NAIR has the NAIR Secure resource allocation option.
So, you know, if you're interested in either, you know, feel free to touch base with us. And with that, you know, I am at the end of the presentation. If you have any additional, you know, research computing questions, feel free to contact us at the following information.
You can also check out our site, which has more information about different services that we provide and, you know, goes through our service portfolio. But other than that, you know, I hope, you know, if there are any researchers or educators on call that you reach out and take advantage of this great opportunity. Thank you.
Thanks, Liz. There's been a little bit of discussion in the chat about the eligibility, and I know the language in both NAIR and I think Access is like U.S.-based research. Does that mean that someone like an international postdoc that is at Columbia would be eligible if their PI signed off? Right, you know, so that's a really good question.
So, you know, if you know, if this does this mean that only like Americans or, you know, U.S. citizens can participate and use the resources? No. So, you know, citizenship does not impact general eligibility. So the requirement is just that you must be affiliated with a U.S.-based institution and be working in the U.S. or conducting your research in the U.S. I hope that answers that question.
Let me check the chat. Yeah, I think that definitely is good. And of course, feel free to speak up if anybody has any questions.
Sorry, I have a follow-up on that. Go ahead. Yeah.
So I was wondering because researchers means that it can be a student or a postdoc as well, right? So they are also eligible to write a proposal or it has to be a faculty or a research scientist. And which particular opportunity are you talking about, the researcher resources or the classroom educator resources? Researcher resources, just for my research project. So for that, you would have to be at a minimum a grad student or higher.
So the undergrad opportunities are through the classroom resources opportunity, which need to be done in parallel with either a professor or some educator that's leading it and then allowing their students to then use the resources, if that makes sense. Yeah, that's very helpful. Thank you.
