• AI / HPC Systems

    Meta (Menlo Park, CA)
    …fabric and host networking, comms lib and scheduling infrastructure. **Required Skills:** AI / HPC Systems Performance Engineer Responsibilities: 1. ... **Summary:** Meta's AI Training and Inference Infrastructure is growing exponentially...a loss-less fabric interconnect with minimal latency. To improve performance of these systems we constantly look… more
    Meta (04/20/25)
    - Related Jobs
  • Production Systems Engineer, Sustaining

    Meta (Menlo Park, CA)
    …hardware and software components, co-design 15. Experience in developing or debugging AI / HPC systems , performance optimizations, including familiarity ... or supporting production hardware at scale 9. Experience in deploying and productionizing AI / HPC systems and/or related components at scale 10. Experience in… more
    Meta (05/20/25)
    - Related Jobs
  • Senior AI - HPC Cluster Engineer

    NVIDIA (Santa Clara, CA)
    …to work effectively with diverse teams and individuals. + Experience analyzing and tuning performance for a variety of AI / HPC workloads. + Passion for ... GPU compute clusters that run demanding deep learning, high performance computing, and computationally intensive workloads. We seek a...storage systems like Lustre and GPFS for AI / HPC workloads + Familiarity with deep learning… more
    NVIDIA (04/02/25)
    - Related Jobs
  • Senior AI - HPC Storage Engineer

    NVIDIA (Santa Clara, CA)
    …designing and operating large scale storage infrastructure. + Experience analyzing and tuning performance for a variety of AI / HPC workloads. + Experience ... join us today! As a member of the GPU AI / HPC Infrastructure team, you will provide leadership...solutions to enable runs of demanding deep learning, high performance computing, and computationally intensive workloads. We seek an… more
    NVIDIA (05/07/25)
    - Related Jobs
  • Postdoctoral Appointee - HPC & AI

    Argonne National Laboratory (Lemont, IL)
    …on designing the communication infrastructure for next-generation High- Performance Computing ( HPC ) and Artificial Intelligence ( AI ) systems . This ... and optimize workload-specialized interconnects and network-aware communication strategies to enhance the performance of AI and HPC workloads. + Implement… more
    Argonne National Laboratory (06/13/25)
    - Related Jobs
  • Senior Observability Architect, AI

    NVIDIA (Santa Clara, CA)
    …looking for a technical leader to define a vision and roadmap for distributed observability systems for large-scale AI and HPC clusters and workloads and ... and visualization to spectacularly improve efficiency, performance , and productivity of AI and HPC workloads. You will lead technical teams to develop,… more
    NVIDIA (05/15/25)
    - Related Jobs
  • AI Infrastructure Engineer - HPC

    Cisco (Research Triangle Park, NC)
    …and technologies. Preferred Qualifications * Deep understanding of operating systems , computer networks, and high- performance applications. * Established ... Showcase the power of Cisco: our people, products, processes, systems , and data. Please join us and make this...and managing the internal NVIDIA DGX and Cisco-UCS based AI platforms at Cisco. You will provide leadership in… more
    Cisco (05/24/25)
    - Related Jobs
  • AI / HPC Network Engineer

    Meta (Menlo Park, CA)
    …requirements of RDMA workloads that expects a loss-less fabric interconnect. To improve performance of these systems we constantly look for opportunities across ... fabric and host networking, comms lib and scheduling infrastructure. **Required Skills:** AI / HPC Network Engineer Responsibilities: 1. Design, develop, test and… more
    Meta (05/08/25)
    - Related Jobs
  • HPC SRE Systems Engineer

    Ford Motor Company (Dearborn, MI)
    We are seeking a highly skilled and motivated HPC SRE Systems Engineer to join our growing team. You will be responsible for designing, building, and maintaining ... + Design, implement, and maintain a robust and scalable HPC infrastructure to support containerized AI /ML workloads...Troubleshoot and resolve complex technical issues related to Linux systems , networking, storage, and HPC applications. +… more
    Ford Motor Company (05/28/25)
    - Related Jobs
  • HPC Systems Admin

    General Dynamics Information Technology (Fairfax, VA)
    …High Speed Networks, Parallel File systems . . Experience running and optimizing HPC performance benchmarks or MPI codes would be a plus. . Experience ... Able to Obtain:** None **Public Trust/Other Required:** NACI (T1) **Job Family:** Systems Engineering **Skills:** High- Performance Computing ( HPC ) Systems more
    General Dynamics Information Technology (06/13/25)
    - Related Jobs