OpenAI

Network Operations Engineer, AI Networking

San Francisco·$157K–302K

Scaling · Datacenter Design

About the Team

OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads.

About the Role

We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network.

The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil.

Key Responsibilities

  • Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers.
  • Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), an
Read the rest on jobs.ashbyhq.com