The GPU resources at SubMIT have recently been updated to increase overall capacity and an express queue has been introduced. The changes are described below:
- Express queue for short GPU jobs
A new high-priority partition submit-gpu-express has been introduced to support rapid iteration and debugging. This partition has one dedicated node and higher priority on the nodes shared with submit-gpu. submit-gpu-express has a time limit of 1 hour. Jobs that require longer runtimes should continue to use the regular submit-gpu partition. The express partition can be used by adding “#SBATCH –partition=submit-gpu-express” to the batch script. - Recovered NVIDIA GTX 1080 nodes
8 GPU nodes equipped with NVIDIA GTX 1080 GPUs (4 GPUs per node) were recovered and returned to service under both the submit-gpu and submit-gpu-express partitions. This brings the total number of GTX 1080 nodes available to users to 11. - RTX 6000 node added for community use
A GPU node equipped with 2 RTX 6000 GPUs is now available for general community use. Jobs targeting this node can be scheduled through both the submit-gpu and submit-gpu-express partitions.
Updated from Users’ meeting on October 27
Benedikt Maier presented his experience using Submit’s GPU resources to train multimodal jet taggers for the CMS experiment, the slides can be found here. Based on a discussion at Users’ meeting, we added the list of nodes with GPUs to the User’s guide.
Please reach out to us with any questions about these changes or to request other features.


Leave a Reply
You must be logged in to post a comment.