Himalayas Remote / WFH Teknologi & IT Full Time

Incident Response Engineer-Facility Operations Center

NVIDIA

Australia Gaji dirahasiakan Diposting 5 hari lalu
Lokasi Australia
Gaji Gaji dirahasiakan
Tipe Kerja Full Time · Remote
Negara Australia

Deskripsi Pekerjaan

Informasi lengkap tentang posisi dan persyaratan

Ringkasan Yukerja

Lowongan Incident Response Engineer-Facility Operations Center di NVIDIA kami kurasi dari Himalayas (kategori Teknologi & IT). Posisi ini ditandai sebagai remote — pastikan timezone dan syarat lokasi kandidat di deskripsi resmi. Yukerja.com bukan pemberi kerja — lamaran diproses di situs sumber resmi.

We are looking for a highly motivated and skilled Incident Response Engineer to join our Facility Operations Center (FOC) team. In this critical role, we are responsible for coordination and presentation within NVIDIA’s datacenters, with a specific focus on incident response, vendor support, and maintenance performance. You will be instrumental in ensuring the reliability and availability of our datacenter environments and minimizing the blast radius of incidents.

What you'll be doing:

  • Primary role is to perform coordination and communication across NVIDIA’s datacenter portfolio from an operations perspective regarding incidents, maintenance, and reporting/monitoring.

  • Develop standards and programs in support of reliability and operations initiatives, including Problem and Change Control, and define and maintain a health score for sites and environments, including testing methods to predict and isolate points of failure, assessing and advising on maintenance strategies, and providing related reporting and metrics.

  • Study failure data and work with machine learning and AI teams and tools to predict future failures, and facilitate reliability studies such as critical assessments, RAM models, and RCM studies. Identify and drive automation & process improvement opportunities across catalog quality workflows and reporting.

  • Coordinate disaster recovery tests, liaise during audits, collaborate with internal partners, and make vital progress to ensure business continuity and compliance. Perform risk assessments to ensure compliance with policies, procedures, rules & regulations, and data center standards.

  • Own and present end-to-end key business metrics related to incident response, including ownership and representation of internal and external tooling. Lead root cause analysis for outages and adjust documentation, workflows, and operating procedures to avoid future incidents.

  • Assess process improvement & transformation opportunities and partner with process owners & collaborators to scope opportunities, define problem statements and objectives, and structure projects and teams.

  • Work multi-functionally with other team members and groups within the organization, and develop strong, productive relationships across peer organizations that further the organization's business objectives. This will incorporate training, coaching, and mentoring Operations teams as needed to empower them to use operations tools and systems to meet daily business needs.

  • Other projects and duties as assigned.

What we need to see:

  • Bachelor’s degree in a related field (e.g., Electrical Engineering, Mechanical Engineering, Industrial Engineering, Computer Engineering, Telecommunication Engineering, Computer Science, or business-related field) or equivalent experience.

  • 5+ years of operations or environmental, health, and safety experience within data centers.

  • Proficient in developing and driving reliability activities (modeling predictions, life cycle testing, stress testing, etc.).

  • Commercial and financial awareness, with a full comprehension of the impact of failure in translation to business costs, production targets, and fulfillment of customer orders.

  • Highly developed numeracy, statistical, and reporting skills; ability to analyze, interpret, and apply information, data, and trends.

  • Enthusiastic about achieving goals and maintaining organization, capable of strategizing and meeting set objectives. Demonstrated ability to be meticulous, organized, and capable of consolidating data analyses for presentation to large-scale groups.

  • Proficient in the use of asset database and DCIM solutions to extract data and develop meaningful insights.

  • Experience in designing, deploying, or maintaining large-scale datacenter infrastructure (whether ACSMEP or networking) or the ability to create strategic infrastructure roadmaps including on-premise, hybrid, and cloud technologies.

  • Demonstrated knowledge and advanced proficiency working with Microsoft Office Suite software and G-Suite software.

Ways to stand out from the crowd:

  • Proven experience in reliability engineering related to electrical or mechanical cooling systems.

  • Certifications such as CDCMP, CMRP, CRL, CRE in Maintenance and Reliability.

  • Knowledge of relevant ISO standards and their implementation.

  • Demonstrated expertise in statistics, forecasting, and management information methods and techniques.

  • Strong IT systems knowledge and skills including sophisticated Office/G-Suite skills and the ability to learn new software packages.

Originally posted on Himalayas

Disclaimer: Yukerja.com adalah agregator lowongan kerja, bukan pemberi kerja. Lowongan ini diagregasi dari Himalayas. Proses lamaran dilakukan di situs resmi perusahaan atau portal sumber. Kami tidak bertanggung jawab atas keakuratan informasi lowongan.

Tips Melamar Incident Response Engineer-Facility Operations Center

  1. Baca deskripsi lengkap dan pastikan skill Anda match sebelum melamar ke NVIDIA.
  2. Sesuaikan CV dan cover letter dengan kata kunci dari job description — terutama untuk kategori Teknologi & IT.
  3. Klik Lamar Sekarang untuk diarahkan ke Himalayas. Proses rekrutmen sepenuhnya di situs sumber.
  4. Siapkan portfolio atau LinkedIn yang update jika diminta di tahap screening.
  5. Waspadai permintaan transfer uang — lowongan resmi tidak memungut biaya.

Artikel terkait: CV ATS · Blog Karir & Tips