For this workshop, the AI agent will help you create and submit SLURM jobs, check their status, troubleshoot problems, and adjust computing resources.
Throughout this workshop, we provide example prompts that you can copy and use with the AI agent. These are suggested prompts, not commands with fixed syntax.
You do not need to memorize or reproduce the exact wording. You can describe what you want in your own words, ask follow-up questions, change your request, or ask the agent to explain something you do not understand.
Think of the prompts as examples of what information to give the agent and what questions to ask, rather than syntax that must be followed literally.
As you become more comfortable working with AI agents, you will likely use shorter prompts and develop the analysis through conversation.
If this is your first time using Codex or Claude Code on BioHPC, follow the appropriate setup instructions:
Codex: Codex setup instructions
Claude Code: Claude Code setup instructions
ssh your_user_id@cbsulogin.biohpc.cornell.edu
You can also connect using VS Code Remote SSH.
xxxxxxxxxxmkdir -p /home2/$USER/bioslurm_workshop
xcd /home2/$USER/bioslurm_workshop#ChatGPT usercodex#Claude userclaude
This file provides the information that AI Agent should know about BioSlurm.
xxxxxxxxxxRead /programs/ai_pipelines/slurm/bioslurm.md
For this workshop, you only have access to computing nodes in the training partition of BioSlurm.
xxxxxxxxxxUse only the training partition.
xxxxxxxxxxDo I have access to BioSlurm?
xxxxxxxxxxsacctmgr show assoc | grep $USER
Interpretation:
Output containing your username and bioslurm indicates an account association.
No output may mean that you do not have an association.
xxxxxxxxxxWhat is the current status of BioSlurm?What jobs are currently in the queue?Do I have any jobs in the queue?
xxxxxxxxxxsqueue -M bioslurm -p trainingsqueue -M bioslurm -p training -u $USER
Interpretation:
Displays all active BioSlurm jobs, includes both pending (PD) and running (R) jobs.
xxxxxxxxxxWhat computing resources are available on BioSlurm?What partitions are available, and when should I use each one?
The agent will explain the CPU/GPU nodes and short, long, gpu, and debug partitions from the provided instructions.
xxxxxxxxxx# Estimated start time for pending jobssqueue -M bioslurm -p training -u "$USER" --start# Partitions and available nodessinfo -M bioslurm -p training# Detailed job informationscontrol -M bioslurm show job JOB_ID
xxxxxxxxxxUse R to generate 10 pairs of random numbers and make a scatter plot. Run the analysis as a BioSlurm job.Check the status of my job.Did the job finish successfully? Show me the output files.
The AI agent should create a SLURM batch script, then use the sbatch command to submit the job to BioSlurm.
After the agent submits the job, look in your project directory. Can you identify the files it created? Ask the agent to explain the SLURM batch script and the resources requested for the job.
You can also ask:
xxxxxxxxxxCheck the status of my job.Did the job finish successfully? What output files were created?
In this exercise, we have FASTQ files from three samples. Each sample can be aligned to the reference genome independently, making this analysis suitable for parallel execution.
xxxxxxxxxxCreate a project directory at /home2/$USER/alignments and copy all files from /shared_data/bioslurm_ws/sequences into it.Align the FASTQ files from all samples to the reference genome using BWA-MEM. Run the sample alignments in parallel on BioSlurm.Do not submit the jobs yet. Show me your proposed approach and resource requests first.
Then ask Agent:
xxxxxxxxxxHow are you parallelizing this analysis?
The agent may create multiple independent SLURM jobs or use a SLURM job array. Ask it to explain the approach it selected.
Also examine the requested CPU, memory, scratch space, and runtime. For real analyses, these requirements are not always known precisely in advance. The AI agent can estimate them, but you may need to adjust the requests based on your knowledge of the data and previous jobs.
Then let the agent submit the jobs and monitor their progress.
Some analyses require specific hardware or computing resources. For example, GPU software may require a particular type of GPU.
BioSlurm currently has P100 and A40 GPUs, and a GPU job must explicitly request a GPU resource.
xxxxxxxxxxOn BioSlurm, create and submit a small GPU test job in /home2/$USER/gpu_test.Use the gpu partition and request exactly one P100 GPU, one CPU, 1 GB of memory, and five minutes. Run NVIDIA's CUDA vector addition example.
After the agent submits the job, ask:
xxxxxxxxxxShow me the SLURM resource requests you used. Explain what each one does.
Then perhaps:
xxxxxxxxxxWhich BioSlurm nodes could this job run on?
Nextflow is a workflow management system that can run multi-step analyses and submit individual tasks to a SLURM cluster.
In this exercise, we will use the sequencing data from Exercise 3 for variant calling.
The variant-calling workflow should:
align reads to the reference genome using BWA-MEM;
generate variant likelihoods using bcftools mpileup;
call variants using bcftools call;
use an appropriate nf-core pipeline.
It is OK if you do not remember the technical details of how to perform variant calling. The Bioinformatics Facility is working on curating and developing commonly used data-analysis workflows as AI skills. Once these skills are available, you will be able to load the appropriate skill and ask the AI agent to perform the analysis following the curated workflow.
For this exercise, the variant-calling skill is not provided. Instead, we will give the AI agent the necessary workflow instructions directly in the prompt. In other words, you will provide the "skill" through the prompt..
xxxxxxxxxx- Using the sequencing data in /home2/$USER/alignments, perform variant calling using an appropriate nf-core Nextflow pipeline. - Use BWA-MEM for read alignment, followed by bcftools mpileup and bcftools call for variant calling.- Run the workflow on BioSlurm. Nextflow itself should run in a small BioSlurm job and submit the individual analysis tasks as BioSlurm jobs.- Do not start the analysis yet. First identify the appropriate nf-core pipeline and show me your proposed workflow.
Then ask
xxxxxxxxxxExplain how this nf-core pipeline implements the variant-calling steps I requested.
Review the proposed workflow with the AI agent. Ask questions about any steps, parameters, or resource requests that you do not understand. You can also ask the agent to modify the plan if necessary.
Once you are satisfied with the plan, tell the agent:
xxxxxxxxxxGo ahead and run the workflow.