Submitting more applications increases your chances of landing a job.
Here’s how busy the average job seeker was last month:
Opportunities viewed
Applications submitted
Keep exploring and applying to maximize your chances!
Looking for employers with a proven track record of hiring women?
Click here to explore opportunities now!You are invited to participate in a survey designed to help researchers understand how best to match workers to the types of jobs they are searching for
Would You Be Likely to Participate?
If selected, we will contact you via email with further instructions and details about your participation.
You will receive a $7 payout for answering the survey.
Download the Bayt App to manage your real time conversation with the recruiter
We are seeking an experienced HPC (High-Performance Computing) System Engineer to design, implement, and manage cutting-edge HPC infrastructure using Dell servers, AMD GPUs (MI210), and Pure Storage systems. The ideal candidate will have expertise in Commvault backup systems, Kubernetes container orchestration, and multitenancy configurations, ensuring scalable, GPU-accelerated, and high-performance solutions tailored to enterprise and HPC workloads.
Key Responsibilities:
• Dell Servers:
• Architect and deploy HPC systems using Dell PowerEdge servers, ensuring high availability and optimized performance for compute-intensive applications.
• Manage server hardware lifecycle, including deployment, upgrades, and diagnostics.
• Configure HPC cluster nodes for seamless integration with Kubernetes and GPU workloads.
• AMD GPUs (MI210):
• Deploy and optimize AMD GPU-based servers to accelerate AI/ML, HPC, and data-intensive applications.
• Monitor GPU utilization, troubleshoot performance bottlenecks, and optimize workloads for GPU acceleration.
• Integrate GPUs into Kubernetes environments for containerized GPU-based applications.
Pure Storage:
• Design and manage Pure Storage solutions, including FlashBlade, to support HPC and data-intensive workloads.
• Implement multitenancy configurations for isolated, secure, and efficient resource utilization.
• Monitor storage health and ensure performance optimization for high-speed data access.
• Commvault Backup:
• Architect and manage enterprise-wide Commvault backup solutions, ensuring data integrity and readiness for disaster recovery.
• Implement backup and retention policies for HPC environments, including containerized and GPU-accelerated workloads.
Kubernetes Container Management:
• Deploy and manage Kubernetes clusters for HPC applications, ensuring scalability and fault tolerance.
• Configure persistent storage for containerized workloads and integrate storage with GPUs for high-performance data processing.
• Monitor cluster performance and troubleshoot HPC-specific Kubernetes challenges.
• System Optimization and Monitoring:
• Implement advanced monitoring solutions for servers, GPUs, storage, and Kubernetes clusters to ensure peak performance.
• Develop and enforce policies for system security, resource allocation, and compliance with industry standards.
• Lead capacity planning and scaling initiatives for HPC infrastructure.
Team Leadership and Collaboration:
• Mentor and guide junior engineers on HPC best practices, system design, and troubleshooting techniques.
• Collaborate with cross-functional teams, including data scientists and DevOps, to align infrastructure capabilities with organizational goals.
Qualifications:
• Technical Skills:
• Extensive experience with Dell PowerEdge servers in HPC or enterprise environments.
• Proven expertise in AMD GPUs (MI210), including their integration and optimization for AI/ML and HPC workloads.
• Advanced knowledge of Pure Storage systems, including multitenancy and high-performance configurations.
• Expertise in Commvault backup systems, including design, deployment, and disaster recovery.
• Strong proficiency in Kubernetes container orchestration, particularly for GPU-accelerated applications.
• Knowledge of high-performance interconnects (e.g., RDMA, InfiniBand) and networking for HPC.
Soft Skills:
• Strong problem-solving and analytical skills for addressing HPC-specific challenges.
• Effective communication and collaboration skills for technical and non-technical stakeholders.
• Leadership skills for mentoring and guiding junior team members.
Preferred Qualifications:
• Certifications in Dell EMC Proven Professional, AMD GPUs, Pure Storage, and Commvault.
You'll no longer be considered for this role and your application will be removed from the employer's inbox.