Data Center Facility Operations Lead
San Francisco, CA | New York City, NYFull-TimeLeadOperations
About Anthropic
- Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
About Anthropic
- Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
About the role
- As the Facility Operations Lead within the Data Center Infrastructure organization, you will define and manage end-to-end operations across Anthropic’s growing fleet of data center infrastructure. You will be accountable for the global operations management of critical mechanical, electrical, plumbing, and control systems. As a subject matter expert in data center infrastructure and operations, you will develop scalable processes for execution, help manage data center partners & vendors, lead deployment schedules and maintain close coordination with related disciplines such as design, engineering, and capacity planning. If you are experienced in operations, are passionate about HPC data centers, & enjoy working in complex, fast-paced environments, we welcome you to apply.
Responsibilities
- Define the Facility Operations strategy including SLAs, incident management, reporting methods, and approaches for reporting.
- Oversee third-party deployment, commissioning, and ongoing maintenance of facility infrastructure such as switch gear, technology cooling systems, and computer room air handlers.
- Create standards for monitoring data center systems operation, minimizing risk, and effectively responding to/managing excursions and incidents.
- Develop innovative approaches to improve time to market and equipment availability while minimizing cost and complexity.
- Enable scale-out of operations through industry best practices optimization of bill of materials, logistics, and physical processes.
- Manage partner relationships and maintain alignment with Anthropic goals.
- Define the Facility Operations strategy including SLAs, incident management, reporting methods, and approaches for reporting.
- Oversee third-party deployment, commissioning, and ongoing maintenance of facility infrastructure such as switch gear, technology cooling systems, and computer room air handlers.
- Create standards for monitoring data center systems operation, minimizing risk, and effectively responding to/managing excursions and incidents.
- Develop innovative approaches to improve time to market and equipment availability while minimizing cost and complexity.
- Enable scale-out of operations through industry best practices optimization of bill of materials, logistics, and physical processes.
- Manage partner relationships and maintain alignment with Anthropic goals.
You may be a good fit if you
- Have 10+ years of experience in data center facility operations as a manager, engineer, technical lead or related role.
- Can demonstrate a proven track record overseeing a portfolio of mission critical systems.
- Are a subject matter expert in facility operations process development including deployment, commissioning, & maintenance.
- Are experienced in third-party partner relationship management.
- Bachelor's degree in relevant domain or equivalent practical experience.
- Have 10+ years of experience in data center facility operations as a manager, engineer, technical lead or related role.
- Can demonstrate a proven track record overseeing a portfolio of mission critical systems.
- Are a subject matter expert in facility operations process development including deployment, commissioning, & maintenance.
- Are experienced in third-party partner relationship management.
- Bachelor's degree in relevant domain or equivalent practical experience.
It's a bonus if you have
- 15+ years of experience managing operations for large scale data centers.
- An understanding of AI/ML workloads, including power, cooling, and network connectivity
- Experience supporting technology cooling solutions such as rear door heat exchangers and/or cooling distribution units,
- Experience managing third-party datacenters such as colo operators and/or contract staffing vendors.
- Familiarity with data center facilities commissioning processes including developing commissioning plans, creating criteria, and conducting hand-over reviews
- 15+ years of experience managing operations for large scale data centers.
- An understanding of AI/ML workloads, including power, cooling, and network connectivity
- Experience supporting technology cooling solutions such as rear door heat exchangers and/or cooling distribution units,
- Experience managing third-party datacenters such as colo operators and/or contract staffing vendors.
- Familiarity with data center facilities commissioning processes including developing commissioning plans, creating criteria, and conducting hand-over reviews
- The annual compensation range for this role is listed below.
- For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.
How we're different
- We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills.
- The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.
