Anzeige
Zurück
T

AI Infrastructure Engineer | Up to $7K - 0310

THE SUPREME HR ADVISORY PTE. LTD.
Singapore · 10294 km · vor 1 Tagen
VollzeitVor OrtS$5'000–7'000/Mt.
Matching nur mit Login und CV verfügbar
Lade deinen CV hoch und sieh sofort, wie gut jeder Job zu dir passt.
Anmelden

Gefragte Skills

Timely ExecutionKubernetesOnline TrainingData PipelineUtilization ManagementArchitectGPUContainer Orchestration
Anzeige

Stellenbeschreibung

AI Infrastructure Engineer

5 days, Mon - Fri 8.30am to 5.30pm

Salary: $5,000 to $7,000

Location: Kaki Bukit

Operating Systems:  Deep expertise in Linux systems administration, kernel tuning, and shell scripting (Bash/Python).

Accelerated Compute:  Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies.

Orchestration & Workload Scheduling:  Hands-on experience with Kubernetes (GPU operator, device plugins) and/or HPC schedulers (Slurm, Run:ai, Ray).

High-Speed Networking:  Proven experience with RDMA (RoCE v2 /InfiniBand), PFC (Priority Flow Control), and ECN configurations.

Storage Systems:  Familiarity with high-IOPS, low-latency shared storage architectures for AI datasets and model checkpoints.

Automation:  Proficiency in Infrastructure as Code (Terraform) and configuration management (Ansible).

Bachelor’s Degree in Computer Science, Information Technology, Computer

Engineering, or equivalent practical experience.

3–6+  years of hands-on experience in infrastructure engineering, high-performance computing (HPC), DevOps, or cloud infrastructure.

Relevant certifications are a plus (e.g., CKA/CKAD, NVIDIA Certified

Associate/Professional, AWS/Azure/GCP Solutions Architect).

Interested applicants can also send your resume to (Liwenpoh741#gmail.com) and allow our consultant to match you with our clients. No Charges will be incurred by Candidates for any service rendered.

WA ME  97•••019  for more AI Infrastructure Engineer   role

Job scopes:

Compute & Cluster Management

Architect, configure, and maintain high-density multi-GPU compute clusters (e.g. NVIDIA HGX/DGX architectures).

Implement and manage container orchestration platforms (Kubernetes, Slurm, or Ray) optimized for AI/ML distributed workloads.

Monitor GPU health, telemetry, utilization, and thermals; minimize idle compute time and prevent single-node bottlenecks.

High-Performance Networking & Storage

Design and optimize low-latency, lossless network fabrics supporting distributed training (InfiniBand, RoCE v2, NVLink, spine-leaf topologies).

Configure and scale high-throughput parallel file systems and object storage (e.g. Lustre, GPFS/IBM Spectrum Scale, Ceph, MinIO, NVMe-oF) to feed high-speed data pipelines.

Automation & Infrastructure as Code (IaC)

Build and manage automated deployment pipelines using Terraform, Ansible, Helm, or Pulumi.

Maintain standard golden images, Linux OS tuning (kernel parameters, NUMA node binding, GPU drivers, CUDA/cuDNN libraries), and firmware updates.

Operations, Observability & Performance

Set up end-to-end monitoring, alerting, and metrics dashboards (Prometheus, Grafana, DCGM exporter, NVIDIA System Management Interface).

Partner with AI/ML engineering teams to diagnose network bottlenecks, NCCL communication latency, and I/O wait states during distributed training jobs.

Lead incident response, root-cause analysis (RCA), and disaster recovery plans for mission-critical AI environments.

POH LI WEN REG NO: R25136683

THE SUPREMEHR ADVISORY PTE LTD EA NO: 14C7279

Quelle: mycareersfuture.gov.sg. Für die Inhalte der Inserate übernehmen wir keine Haftung.

Bewerbungstext erstellen

Wir erstellen aus deinem Profil und dieser Stelle einen Entwurf. Du prüfst und passt ihn an.

Nur mit Konto verfügbar.

Jetzt anmelden
Anzeige

Ähnliche Stellen