Data Center Hardware Operations Lead
San Francisco, CA | New York City, NYFull-TimeLeadOperations
About Anthropic
- Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
About Anthropic
- Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
About the role
- As the Hardware Operations Lead within the Data Center Infrastructure organization, you will define and manage end-to-end operations across Anthropic’s growing fleet of data center hardware. You will be accountable for all data center physical operations including material movement, physical deployment, bring-up, maintenance, and incident response for complex network and high performance compute platforms. As a subject matter expert in data center infrastructure and operations, you will develop scalable processes for execution, help manage data center partners & vendors, lead deployment schedules and maintain close coordination with related disciplines such as design, engineering, and capacity planning. If you are experienced in operations, are passionate about HPC data centers, & enjoy working in complex, fast-paced environments, we welcome you to apply.
Responsibilities
- Define the Hardware Operations strategy including SLAs, incident management, reporting methods, and approaches for reporting.
- Oversee third-party delivery and bring-up of network and compute systems from material delivery through QA testing, handover to production, and ongoing maintenance.
- Drive development of tools in support of operations workflows, traceability, and vendor coordination.
- Develop innovative approaches to improve time to market & equipment availability for data center hardware.
- Enable scale-out of operations through industry best practices optimization of bill of materials, logistics, and physical processes.
- Manage partner relationships and maintain alignment with Anthropic goals.
- Define the Hardware Operations strategy including SLAs, incident management, reporting methods, and approaches for reporting.
- Oversee third-party delivery and bring-up of network and compute systems from material delivery through QA testing, handover to production, and ongoing maintenance.
- Drive development of tools in support of operations workflows, traceability, and vendor coordination.
- Develop innovative approaches to improve time to market & equipment availability for data center hardware.
- Enable scale-out of operations through industry best practices optimization of bill of materials, logistics, and physical processes.
- Manage partner relationships and maintain alignment with Anthropic goals.
You may be a good fit if you
- Have 7+ years of experience in data center operations as a manager, engineer, technical lead or related role.
- Can demonstrate a proven track record overseeing a portfolio of mission critical systems.
- Are a subject matter expert in operational processes including deployment, maintenance, material movement, and warehousing.
- Can demonstrate subject matter expertise common data analytics/dashboarding tools such as PowerBI, Atlassian Analytics, Tableau, or similar.
- Are experienced in third-party partner relationship management.
- Bachelor's degree in relevant domain or equivalent practical experience.
- Have 7+ years of experience in data center operations as a manager, engineer, technical lead or related role.
- Can demonstrate a proven track record overseeing a portfolio of mission critical systems.
- Are a subject matter expert in operational processes including deployment, maintenance, material movement, and warehousing.
- Can demonstrate subject matter expertise common data analytics/dashboarding tools such as PowerBI, Atlassian Analytics, Tableau, or similar.
- Are experienced in third-party partner relationship management.
- Bachelor's degree in relevant domain or equivalent practical experience.
It's a bonus if you have
- 10+ years of experience managing operations and logistics for large scale data centers.
- An understanding of AI/ML workloads, including power, cooling, and network connectivity
- Experience in supply chain management including bill of material definition, sourcing, procurement, and vendor onboarding/management.
- Experience managing third-party datacenters such as colo operators and/or contract operations staffing vendors.
- Experience in data center facility infrastructure such as mechanical, electrical, and plumbing systems.
- 10+ years of experience managing operations and logistics for large scale data centers.
- An understanding of AI/ML workloads, including power, cooling, and network connectivity
- Experience in supply chain management including bill of material definition, sourcing, procurement, and vendor onboarding/management.
- Experience managing third-party datacenters such as colo operators and/or contract operations staffing vendors.
- Experience in data center facility infrastructure such as mechanical, electrical, and plumbing systems.
- The annual compensation range for this role is listed below.
- For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.
How we're different
- We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills.
- The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.
