MongoDB · Atlas
We are looking to speak to candidates who are based in Bengaluru for our hybrid working model.
Responsibilities
- Successfully co-ordinate with a global team of Cloud Operations Engineers who are tasked with ensuring our uptime guarantees to our Atlas customer base
- Help scale the worldwide Cloud Operations Engineering team with the strategic implementation of new processes and tools
- Assist in scoping, designing and deploying systems that reduce Mean Time to Resolve for customer incidents
- Monitor and detect emerging customer-facing incidents on the Atlas platform; assist in their proactive resolution
- Automate routine monitoring and troubleshooting tasks
- Diagnose live incidents, differentiate between platform issues versus usage issues, and take the next steps toward resolution
- Cooperate with our product management and cloud engineering organizations by identifying areas for improvement in the management applications powering the Atlas infrastructure
- Inform executive leadership and escalation management personnel of major outages
- Coordinate and participate in a weekly on-call rotation, where you will handle short term customer incidents (from direct surveillance or through alerts via our Technical Services Engineers)
Requirements
- Experience with being an on call DevOps, SRE, or Cloud Operations engineer (at least 8 years)
- Expertise with Linux system administration, configuration, troubleshooting
- Experience in monitoring, system performance data collection and analysis, and reporting
- Expertise with networking technologies like DNS, TCP/IP, etc
- Knowl