Job description
Job Purpose
Cloud System Engineer will function as a core infrastructure specialist within cloud engineering, integrating compute, storage, virtualization and Backup into multi-tenant cloud platforms, while enabling automation, orchestration, and service delivery via IaaS frameworks.
Roles & Responsibilities
- Deploy, configure, and manage enterprise-grade servers (Dell, HPE, Supermicro GPU/CPU platforms)
- Perform OS provisioning (Linux/Windows) and lifecycle management
- Monitor system performance and optimize CPU, memory, and IO utilization
- Implement firmware upgrades, patching, and hardware health monitoring
- Support workload sizing and capacity planning for cloud tenants
- Manage and administer storage platforms (Block, File, Object – e.g., PowerScale, SAN/NAS, SDS)
- Configure storage pools, volumes, quotas, and replication policies
- Ensure optimal performance through tiering, caching, and load balancing
- Implement data lifecycle management and storage efficiency (dedupe, compression)
- Monitor storage health, latency, and throughput KPIs
- Deploy and manage hypervisors (VMware, KVM, Hyper-V)
- Manage VM lifecycle: provisioning, cloning, migration, decommissioning
- Configure HA, DRS, clustering, and resource scheduling
- Optimize hypervisor performance and troubleshoot VM-related issues
- Support multi-tenant virtualization environments and cloud orchestration platforms
- Design and manage backup solutions (e.g., Commvault, Veeam)
- Configure backup policies for VMs, databases, file systems, and applications
- Monitor backup jobs, ensure SLA compliance, and troubleshoot failures
- Perform restore operations (file-level, VM-level, cross-platform)
- Implement DR strategies including replication, snapshot management, and failover testing
- Integrate compute, storage, and virtualization into cloud orchestration platforms
- Support IaaS platform provisioning and service catalog enablement
- Automate provisioning using scripts/tools (e.g., Ansible, Terraform )
- Ensure API-driven integration for cloud service consumption
- Monitor infrastructure using enterprise tools (alerts, logs, dashboards)
- Perform root cause analysis (RCA) for incidents and failures
- Participate in change management and release processes
- Ensure adherence to SLA, uptime, and service availability targets
- Implement hardening (OS, hypervisor, storage access)
- Ensure compliance with enterprise security policies
- Manage access control (RBAC, IAM integration)
- Support audit readiness and regulatory compliance
- Maintain LLDs, HLDs, SOPs, and runbooks
This job post has been translated by AI and may contain minor differences or errors.