Logo Talent

Systems Engineer/Administrator (m/f/d) for AI Service Infrastructure

Leibniz-Rechenzentrum der Bayer. Akad. d. Wissenschaften Garching bei München, Bayern Remote Teilzeit Tag 1 · Seit 1 Tag online

Details zum Jobangebot

Looking for an employer you can count on? Join us!

We are looking to expand our team and are hiring:

Systems Engineer/Administrator (m/f/d) for AI Service Infrastructure

Your Role and Responsibilities:

  • Planning, deploying and operating an ever-growing AI service infrastructure for researchers and academic users as a member of a dedicated team within a friendly and open work environment
  • Running the day-to-day operations of high-performance cluster systems focused on AI and Data Analytics applications & workflows, following best-practices of system administration, automation and monitoring
  • Interacting with technology providers and developers (hardware/software) as well as national and European collaboration partners and projects
  • You actively contribute to the development and implementation of certified and transparent processes.

Your Qualifications:

*Required/Minimum Qualifications*

  • Master’s degree in computer science or other areas of scientific computing (e.g. physics, math), or Bachelor’s degree with multiple years (>3) of practical experience in system administration/engineering in research or enterprise environments

*Other Requirements*

  • Proficiency working in data center environments (incl. Linux, Git, Gitlab)
  • Advanced knowledge and experience in the administration/management of high-performance cluster systems for AI and Data Analytics, covering the majority of the following areas: 
  • Compute resources, including 
  • Accelerated server compute nodes (incl. GPUs)
  • Job scheduling and resource management (with Slurm Workload Manager)
  • Container orchestration (with Kubernetes)
  • Low latency fabrics (incl. InfiniBand), Ethernet and additional networking technologies (DNS, Firewall, etc.)
  • Fabric attached high-performance storage (e.g. GPFS, NFS, RDMA) 
  • Authentication/Authorization (incl. LDAP, Shibboleth)
  • Practical knowledge in architecting/designing, deploying and maintaining hardware and software systems based on components/technologies as listed above
  • Experience using automated tools for deployment and configuration management (i.e. infrastructure as code, e.g. Ansible), but not afraid to script on CLI, either (e.g. Bash, Python, Perl)
  • Knowledge and experience in software deployment and software lifecycle management (ideally based on principles of continuous integration/continuous deployment, CI/CD)
  • Willingness to keep up with current developments and to learn new technologies in the field of AI
  • Friendly