Getting started#

These docs aren’t long. Please read all of them before running anything. It will save yourself and everyone else a lot of time.

Requesting access#

As mentioned in the guidelines, you need a MyDFKI/Intranet account to access our intranet, and a Pegasus account to access the Slurm cluster itself. Send request to cluster@dfki.de from your DFKI mail address.

Pegasus credentials will only ever be sent to the account holder’s DFKI address.

Connecting to the cluster#

Use ssh to connect to one of our head nodes:

SSH fingerprints:

  • login1.pegasus.kl.dfki.de
    SHA256:/Sjwh5nHthPCVol7XtEM53BKYy2jlTvOu+MFXnCx7M4 (ECDSA) SHA256:BPnYmJOFDWqJEKoUP4GoC1+k7olmFLhP0zRaxZ+pl4M (ED25519) SHA256:nEbr0eTd5v0Iis4/g2oqyZPTBvfa63MjB3qMM72eV0U (RSA)
  • login2.pegasus.kl.dfki.de
    SHA256:g9VaUAKpZYHvhMnM6j/bXeobih/WRD3KkrW2OKVDVGI (ECDSA) SHA256:oZqynX+tSPdHpSkryArsWHEVuls18o32mjuWYMYCMwY (ED25519) SHA256:a1PCII5MTIlcs2odAa34+rFFE21VMn58neEmm9DyHQ8 (RSA)
  • login3.pegasus.kl.dfki.de
    SHA256:grb2zDUzhkE9SpsxxxfMsrWfgzvEDzh+A4w3ncD1/Zk (ECDSA) SHA256:ui6zRhZbSZtmMVEDRbEoTL0QBGgM+mtQCoSZ9b1ZTJ8 (ED25519) SHA256:v6wLEv4W3af48omlydCVE+npc0bbZpi6z6B0cz905N8 (RSA)
ssh [username]@[head node]

Here you can, among other things, modify your home directory and schedule jobs with the srun command. See the Slurm Cluster section for more details on this.

Read the intro message as it may contain important information.

Do not run compute jobs or other resource intensive commands on head nodes! This includes connecting VSCode, Claude Code, and other remote IDEs!

SSH key authentication#

In case you don’t want to type in your credentials each time, feel free to set up public key auth like this (once):

# Create a ssh key pair on your local machine (ed25519 for secure and short keys)
ssh-keygen -t ed25519 -a 100 -C "$USER@$HOSTNAME"

# Copy your public key to remote machine
ssh-copy-id [username]@[head node]

# You should now be able to ssh using your key to authorize
ssh [username]@[head node]

Your $HOME is synchronized across all machines, so you only need to do this once.

Login node restriction#

To contradict the trend of using multiple login nodes simultaneously, and by that overly using their resources, a profile script is in place preventing logins to multiple login nodes. That means:

  • once you’re logged in to one of the login nodes, you can only use this particular node
  • multiple logins to the very same login node are possible
  • while you have a session (any kind of process under your account) on one login node, you cannot connect to any other login node (message with active sessions will be displayed)
  • to switch login nodes, you have to completely leave (end all processes) on the previous login node
  • wait times up to 60s may apply, until terminated sessions are recognized
  • in case you have trapped yourself in a deadlock, due to connecting to more than one login node too quickly (which you should not do anyway), contact Pegasus operators to resolve the issue

Takeaway: Choose one login node, stick to it, switch only for good reasons, terminate all processes on the previous one before connecting to another one.

Login node limits#

Be aware that any kind of resource hooging is not appreciated. To prevent accidental mistakes and to keep the login nodes usable, the following limits are in place for all login nodes:

  • max user sessions per node: 500 (basically any kind of shell)

The per user cgroup slice is:

  • CPU limit 6 cores
  • Mem limit 10G

That is the max amount of CPU cores and RAM a user can utilize on a login node.

There are cron jobs in place that terminate all user processes that were started 2 months ago. This is to prevent lingering and forgotten shells, screens and tmuxes eating up the memory. They run every night.

Moving data in and out#

While the cluster nodes all have shared access to NAS mounts such as /ds* (it’s actually not unlikely that you’ll find your standard dataset there already) or /netscratch, we don’t allow mounting these file systems directly on other machines (e.g., yours). This is due to the fact that performance and security matter.

Hence, if you want to get data in and out of the cluster, you’re left with options that are always there when you have ssh access:

  • scp, rsync
  • sftp clients. Often directly included in your favorite file-browser via sftp://... or fish:// (KDE), otherwise standalone tools such as: CyberDuck, WinSCP
  • sshfs (fuse) mounts

See the section on storage to understand where to put what kind of data.

If you’re planning to download a large dataset or import a lot of data (let’s say starting at 20 GB) or don’t have the right permissions (to put it in what you think is the right place), please contact us.

Running commands in the background#

When you disconnect from a ssh session, your shell will close and with it, all processes running inside said shell. Start a screen or tmux session to be able to disconnect and reconnect to your jobs later:

screen -S [descriptive name]

Detach from the screen with CTRL+A, followed by D. Then to reattach run:

screen -r [descriptive name]

Interactive jobs#

While it’s possible and sometimes useful to run interactive jobs on the cluster, e.g., to debug training code and avoid lengthy setup times or create custom enviroments we ask you to keep them to a minimum. From experience, we know that interactive sessions rarely utilize the requested resources while they’re in use and are promptly forgotten shortly after. Remember that jobs have exclusive access to the resources that they receive, so lingering jobs waste compute time. To limit their impact, interactive jobs must be started with a time limit of 4 hours or less (--time=04:00:00) and an immediate period of at most 3600 seconds (--immediate=3600).

Connecting via Saarbrücken VPN#

In case you’re connecting via VPN provided by DFKI Saarbrücken, be aware that only certain ports on cluster nodes are reachable. The standard ports for ssh (22), http (80), and https (443) as well as ports >= 10000 are reachable. So, if you want to connect to self deployed services, i.e. Jupyter Notebooks, Tensorboard etc, keep in mind to choose ports above 10000 for the service to listen on.

Organizing experiments#

Command lines to run jobs can get quite lengthy, so we recommend that you create shell scripts to reduce some (or all) of the boilerplate.

It also helps to reproduce experiments if all parameters are known. Use sacred or mlflow to automatically document your experiments in minute detail. Combined with our fixed software environments this provides excellent reproducibility.

Make sure to write your code in a way that jobs can be continued in case of any interruption. Create regular snapshots of model and optimizer state so you can resume later. Power outages and other service interruptions are rare, but they do happen eventually. With the ability to continue the job, a lot of time and resources can be saved compared to starting over.

Interactive tutorial#

We created an interactive, self-paced onboarding tutorial that teaches you how to use the Pegasus cluster.

# 1. Connect to a login node
ssh <your-username>@login1.pegasus.kl.dfki.de

# 2. Clone this tutorial into your home directory (code belongs in $HOME)
git clone git@git.opendfki.de:pegasus/tutorial.git
cd ./pegasus-tutorial

# 3. Start
./bin/tutorial