chokepoints.ai
SUBSCRIBE
10 layers580 nodes2,376 dependencies9 chokepoints112 bottlenecks6,500+ companiesnode size = companies identified

HPC batch schedulers (Slurm and successors)

BOTTLENECK

SchedMD's Slurm dominates academic clusters but cannot scale elastically, forcing reliance on rigid queue-based architectures.

Queue-based batch systems like Slurm, inherited from supercomputing, still rule academic and national-lab training clusters. They handle large-node jobs and priority queues well but cannot scale elastically like cloud-native tools. SchedMD and HPC integrators earn licensing and support revenue.

Queue-based workload managers from the HPC tradition, dominant in large-scale training and research clusters.

Why the concentration exists

Slurm is an open-source workload manager designed for Linux-based HPC clusters, handling job scheduling and resource management. Users submit batch jobs with "#SBATCH" directives that specify resource requirements like walltime, memory, and CPU count. The system allocates compute nodes, manages job queues, and enforces priority policies across the cluster. It supports heterogeneous nodes and multifactor fair-share job priority, which allows administrators to balance resource access across user groups.[2][7][10][19]

The scheduler has been tested with up to 120,000 exec nodes and 100,000 submitted jobs. It supports up to 1,000 job submissions per second with 600 executions per second. This scalability makes Slurm suitable for installations ranging from small laboratory clusters to the world's largest supercomputers. The architecture uses multiple controllers and node sharing to maintain performance at extreme scale.[10][12][9]

Concentration around Slurm stems from network effects and the open-source model that eliminates licensing fees. Many HPC centers have coalesced around Slurm as administrators gain experience with its tooling. Alternative schedulers have lost momentum, with SGE's open-source codebase no longer supported as of September 2022. PBS has seen declining adoption as Slurm overtook it in both research and industry applications.[1][15]

What the evidence shows

Many places are coalescing around Slurm, making it the most popular batch scheduler.

reddit.com

Slurm is now the most popular batch scheduler, but 10 years ago it was PBS/TORQUE.

mattermodeling.stackexchange.com

Slurm has overtaken PBS in research and industry, leading to declining adoption of PBS.

vantagecompute.ai
RESCORED JUL 2026near-monopolyscaling3 companies

Who supplies it

SchedMD maintains Slurm and provides commercial support services for enterprise deployments. SchedMD is now part of NVIDIA, which acquired the company to integrate Slurm expertise into its AI infrastructure offerings. Slurm is the scheduler of choice for over half of the top one-hundred systems in the TOP500 ranking of supercomputers. This dominance gives SchedMD significant influence over scheduling standards in HPC environments.[10][13]

IBM Spectrum LSF remains a key commercial vendor in the HPC job scheduling market. IBM designed Spectrum LSF to support diverse distributed workloads including batch jobs, interactive workloads, and containerized applications. The scheduler can make thousands of scheduling decisions per second while efficiently sharing resources among workloads with vastly different runtimes. Most large-scale HPC and AI supercomputers use HPC-oriented workload managers such as Spectrum LSF.[14][18]

Who controls it

No named supplier is publicly confirmed for this node yet.

No independently verified market-size figure is published for this node yet.

What it depends on, and what depends on it

Slurm runs on Linux-based cluster environments with heterogeneous compute nodes. AWS ParallelCluster automates deployment of HPC clusters on AWS infrastructure using Slurm as the native scheduler. Azure CycleCloud supports deploying and managing Slurm alongside OpenPBS, PBS Pro, and LSF. These integrations allow Slurm to operate in hybrid cloud environments while maintaining familiar scheduling semantics.[2][6][17]

Slurm serves researchers, HPC administrators, data scientists, and AI engineers across multiple sectors. Users span academia, government, life sciences, finance, manufacturing, and energy industries. Major installations include Livermore Computing clusters at Lawrence Livermore National Laboratory. CERN processes 330 petabytes across 320,000 cores using Kubernetes for HPC workloads, while IHME runs 20,000 cores for health analytics alongside Slurm in hybrid deployments.[2][8][11]

Slurm integrates with other frameworks to extend its scheduling capabilities. The Slurm Simulator enables parametric analysis of scheduling policies for optimization research. Nebius developed Soperator as a Kubernetes operator that automates Slurm cluster deployment in cloud environments. This allows ML and HPC teams to leverage Slurm while benefiting from Kubernetes-native management.[9][4]

Where it sits in the stack

view in atlas

What would break it

Kubernetes-based alternatives like Volcano target containerized HPC and AI workloads. Volcano is a specialized batch system built for Kubernetes, designed for environments where HPC, AI, and Big Data move into containerized infrastructure. Resource utilization improves from 40-60 percent in traditional HPC clusters to 70-85 percent with Kubernetes through intelligent bin-packing and dynamic scheduling. These efficiency gains pressure traditional batch schedulers to adapt or lose workloads.[3][11]

HTC-Grid, described as a successor to Slurm, achieved throughput exceeding 30,000 tasks per second using 12,000 workers. It completed a batch of 10 million zero-work tasks in under six minutes. HTC-Grid scaled from 1 worker to 20,000 workers on nearly 600 Amazon EC2 Spot Instances using Amazon EKS in approximately 20 minutes. The system backed by Amazon ECS scaled to 76,000 workers within 20 minutes using multiple ECS clusters.[16]

Flux serves as the scheduler and resource manager on CORAL2 systems like El Capitan and Tuolumne at Lawrence Livermore National Laboratory. These next-generation systems represent a shift away from Slurm for certain exascale deployments. Reinforcement learning algorithms are being developed to optimize co-scheduling policies by minimizing idle processing capacity while respecting QoS constraints. Such research may yield new scheduling approaches that displace established tools.[8][5]

What to watch

Hewlett Packard Enterprise holds patent US11695654B2 for high performance compute infrastructure as a service, granted July 4, 2023. Morgan Stanley holds patent US12190166B1 for observing and predicting data batch activity in real time, granted January 7, 2025. Rescale holds patent US12135989B2 for a compute recommendation engine, granted November 5, 2024. These patents indicate ongoing commercial interest in HPC scheduling technology.[20]

Related nodes

Container orchestration (Kubernetes and AI extensions)AI-native cluster and fabric managers

Sources

  1. vantagecompute.ai
  2. devopsschool.com
  3. scmgalaxy.com · 2026-01-22T06:04:42
  4. nebius.com · 2025-08-01T00:00:00
  5. arxiv.org · 2024-01-18T03:21:20
  6. sourceforge.net
  7. hpc-wiki.info
  8. hpc.llnl.gov
  9. arxiv.org
  10. en.wikipedia.org · 2024-01-24
  11. baculasystems.com
  12. aspsys.com
  13. schedmd.com
  14. hpcwire.com · December 16, 2019
  15. aws.amazon.com · 2022-09-14
  16. aws.amazon.com
  17. learn.microsoft.com
  18. oplexa.com
  19. ocaisa.github.io
  20. patents.google.com · 2023-07-04

Full scorecard, owner shares, supply edges and the full tracked roster are in the desk letter.

GET THE BRIEFING