Monitoring Jobs
This page provides information about monitoring user jobs. Slurm provides different commands to gather information about current, future and past jobs.
squeue- provides information on current and future jobsscontrol- provides detailed information on currently running jobssacct- provides detailed information on past jobs
Note that the output format of most Slurm commands is highly configurable to your needs. Look for the —format or —Format options in the man pages of the command.
squeue
Section titled “squeue”Syntax
squeue [options]Common options
--me Display all your currently pending or running jobs--jobs=<job_id[,job_id[,...]]> Request specific jobs to be displayed--partition=<part[,part[,...]]> Request jobs to be displayed from a comma separated list of partitions--states=<state[,state[,...]]> Display jobs in specific states. Comma separated list or "all". Default: "PD,R,CG"The default output format is as follows:
JOBID PARTITION NAME USER ST TIME NODES NODELIST(REASON)where
JOBID Job or step ID. For array jobs, the job ID format will be of the form <job_id>_<index>PARTITION Partition of the job/stepNAME Name of the job/stepUSER Owner of the job/stepST State of the job/step. See below for a description of the most common statesTIME Time used by the job/step. Format is days-hours:minutes:seconds (days,hours only printed as needed)NODES Number of nodes allocated to the job or the minimum amount of nodes required by a pending jobNODELIST(REASON) For pending jobs: Reason why pending. For failed jobs: Reason why failed. For all other job states: List of allocated nodes. See below for a list of the most common reason codesSee the man page for more information: man squeue
Job States
Section titled “Job States”During its lifetime, a job passes through several states:
PD Pending. Job is waiting for resource allocationR Running. Job has an allocation and is runningS Suspended. Execution has been suspended and resources have been released for other jobsCA Cancelled. Job was explicitly cancelled by the user or the system administratorCG Completing. Job is in the process of completing. Some processes on some nodes may still be activeCD Completed. Job has terminated all processes on all nodes with an exit code of zeroF Failed. Job has terminated with non-zero exit code or other failure conditionWhy is my job still pending?
Section titled “Why is my job still pending?”(Resources)
The job is waiting for resources to become available so that the jobs resource request can be fulfilled.
(Priority)
The job is not allowed to run because at least one higher prioritized job is waiting for resources.
(Dependency)
The job is waiting for another job to finish first (—dependency=… option).
(DependencyNeverSatisfied)
The job is waiting for a dependency that can never be satisfied. Such a job will remain pending forever. Please cancel such jobs.
(QOSMaxCpuPerUserLimit)
The job is not allowed to start because your currently running jobs consume all allowed CPU resources for your user in a specific partition. Wait for jobs to finish.
(AssocGrpCpuLimit)
dito.
(AssocGrpJobsLimit)
The job is not allowed to start because you have reached the maximum of allowed running jobs for your user in a specific partition. Wait for jobs to finish.
(ReqNodeNotAvail, UnavailableNodes:…)
Some node required by the job is currently not available. The node may currently be in use, reserved for another job, in an advanced reservation, DOWN, DRAINED, or not responding. Most probably there is an active reservation for all nodes due to an upcoming maintenance downtime and your job is not able to finish before the start of the downtime. Another reason why you should specify the duration of a job (—time) as accurately as possible. Your job will start after the downtime has finished. You can list all active reservations using scontrol show reservation.
Why can’t I submit further jobs?
Section titled “Why can’t I submit further jobs?”_sbatch: error: Batch job submission failed: Job violates accounting/QOS policy (job submit limit, user's size and/or time limits)… indicates that you have reached the maximum of allowed jobs to be submitted to a specific partition.
Examples
Section titled “Examples”List all your currently running jobs:
squeue --me --states=RList all your currently running jobs in the gpu partition:
squeue --me --partition=gpu --states=Rscontrol
Section titled “scontrol”Use the scontrol command to show more detailed information about a job
Syntax
scontrol [options] [command]Examples
Section titled “Examples”Show detailed information about job with ID 500:
scontrol show jobid 500Show even more detailed information about job with ID 500 (including the jobscript):
scontrol -dd show jobid 500Use the sacct command to query information about past jobs
Syntax
sacct [options]Common options
--endtime=end_time Select jobs in any state before the specified time.--starttime=start_time Select jobs in any state after the specified time.--state=state[,state[,...]] Select jobs based on their state during the time period given. By default, the start and end time will be the current time when the --state option is specified, and hence only currently running jobs will be displayed.