LogoKode$word
SpaceX logo
Verified Tech Organization

Careers at SpaceX

Browse and filter through all verified positions currently open at SpaceX.

Total Company Roles516
Matching Filter516
spacex.comHQ: Hawthorne, California, United States

SpaceX designs, manufactures and launches the world's most advanced rockets and spacecraft. The company was founded in 2002 by Elon Musk to revolutionize space transportation, with the ultimate goal of making life multiplanetary. SpaceX has gained worldwide attention for a series of historic milestones. It is the only private company ever to return a spacecraft from low-Earth orbit, which it first accomplished in December 2010. The company made history again in May 2012 when its Dragon spacecraft attached to the International Space Station, exchanged cargo payloads, and returned safely to Earth - a technically challenging feat previously accomplished only by governments. Since then Dragon has delivered cargo to and from the space station multiple times, providing regular cargo resupply missions for NASA. For more information, visit www.spacex.com.

Sector:aerospaceaviationaviation and aerospacedracospace travel

All Openings (516)

Ordered by most recently published

Site Reliability Engineer (Application Software)

On-sitefull timeMid-LevelCalifornia, United States
Apply Now

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SITE RELIABILITY ENGINEER (APPLICATION SOFTWARE) The application software team is the central nervous system of SpaceX. We build mission-critical platforms that accelerate vehicle software delivery, testing, and operations for every Falcon 9, Starship, and Dragon mission all while powering Starlink’s global growth. This position will have a meaningful impact on Starship by significantly reducing safety-critical build and test times for vehicle software. We are looking for a Site Reliability Engineer who brings a strong SRE mindset, cares deeply about safety, quality, and attention to detail, and possesses the ability to understand the big picture before writing code. The ideal candidate fully understands what they are building, enjoys hard problem solving, thinks strategically, and is decisive, organized, and self-critical. SpaceX relies on our vehicle software being built quickly and correctly, tested rigorously, and rapidly iterated on. You will build and maintain the tools that make this possible. Every time a Falcon 9 or Starship launches, a Dragon capsule docks with the ISS, or a Starlink satellite connects a new community, the software responsible for it was created with the tools you design, improve, and scale. Aerospace experience is not required. We value smart, motivated, collaborative engineers who treat teammates with fairness, respect, and support, and who want to take full ownership of challenging problems to help make humanity multi-planetary. RESPONSIBILITIES: Deploy, upgrade, operate, maintain, and scale our suite of mission-critical products and services Manage our underlying infrastructure as code and use modern observability tools to provide a complete picture of application health Closely collaborate with software engineers to design and build highly operable, maintainable, and testable systems Engage in and improve the entire software development lifecycle — from inception and design through deployment, operation, and continuous refinement Practice sustainable incident response and blameless postmortems Provide high-quality end-user support to vehicle software engineers Participate in the team’s on-call rotation Identify and eliminate performance bottlenecks using measurement and creative engineering BASIC QUALIFICATIONS: Bachelor’s degree in computer science, information systems, or an engineering discipline; OR 3+ years of professional experience in SRE or DevOps in lieu of a degree 1+ years of experience with Python and Python-based development frameworks Experience with Linux operating systems PREFERRED SKILLS AND EXPERIENCE: Experience with build systems (Bazel, Buck, Make, etc.) Experience with both container and virtualization technologies (Docker, Kubernetes, vSphere, QEMU, KVM, etc.) Experience with databases and data modeling (Postgres, MySQL, ClickHouse, etc.) Experience with infrastructure as code (IaC) tools for managing fleets of servers Experience with Terraform, Ansible, Puppet, or similar automation frameworks Knowledge of the technologies that predate and underpin modern cloud infrastructure, with the ability to translate high-level developer experiences into specific implementations from first principles Ability to work with mission-critical and sensitive systems with appropriate urgency and care Ability to communicate effectively with customers, peers, and management in both formal and informal settings Experience with full-stack development (the team primarily uses Python, JavaScript, and C#; end users primarily use C++) ADDITIONAL REQUIREMENTS: Must be able to work extended hours and weekends as needed COMPENSATION AND BENEFITS: Pay Range: Level I: $125,000.00 - $160,000.00/per year Level II: $145,000.00 - $195,000.00/per year Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience. Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock, stock options, or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan. You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation & will be eligible for 10 or more paid holidays per year. Employees accrue paid sick leave pursuant to Company policy which satisfies or exceeds the accrual, carryover, and use requirements of the law. ITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here . SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status. Applicants wishing to view a copy of SpaceX’s Affirmative Action Plan for veterans and individuals with disabilities, or applicants requiring reasonable accommodation to the application/interview process should reach out to EEOCompliance@spacex.com .

View more...
Cloud, DevOps & SRE
VerifiedToday

Site Reliability Engineer, GNC

On-sitefull timeMid-LevelCalifornia, United States
Apply Now

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SITE RELIABILITY ENGINEER, GNC SpaceX’s mission is to make humanity multiplanetary by developing fully and rapidly reusable launch systems capable of launching Starship multiple times per day while continuing to scale the Starlink constellation. To support these goals, we are seeking a Site Reliability Engineer to operate and scale custom-built, mission-critical products for the Guidance, Navigation, and Control (GNC) teams. GNC teams at SpaceX are responsible for vehicle design, trajectory design and optimization, high-fidelity vehicle simulation, software and control algorithm development, while also supporting both launch and on-orbit operations across multiple vehicle programs. In this role, you will work closely with GNC teams across SpaceX to maintain and improve a suite of critical GNC-focused tools and infrastructure that must scale reliably to enable a multiplanetary future. These systems include on-prem services, large-scale Monte Carlo simulations on our high-performance computing (HPC) cluster, automated data analysis pipelines, continuous integration systems for rocket and simulation software, GNC analysis infrastructure, and vehicle configuration verification tools. The ideal candidate is flexible, possesses broad skills spanning product operations and software development, and thrives in a fast-paced, high-impact environment. RESPONSIBILITIES: Deploy, upgrade, operate, and scale a suite of mission-critical GNC products and services Provision and maintain virtual and physical servers Work with SpaceX HPC team to monitor and maintain an HPC cluster consisting of tens of thousands of CPUs. Closely collaborate with GNC software engineers to create highly operable and maintainable products Monitoring and incident response for web applications and services Manage the underlying computational infrastructure of GNC in collaboration with IT stakeholders Engage in and improve the whole lifecycle of services from whiteboard to operational Make data-driven recommendations for future hardware purchases Practice sustainable incident response and postmortems Provide end-user support to GNC engineering for products by becoming an expert on analysis applications and support users in troubleshooting and pointing to features Configure automated deployment pipelines for web apps Develop or improve GNC web apps and tools for better usability, maintainability, and robustness Demo and document new software changes such as operating system upgrades, shared filesystem changes, or major tool rollouts Focus on performance bottlenecks and performance improvement techniques BASIC QUALIFICATIONS: Bachelor’s degree in computer science, information systems/IT, engineering, math, or scientific discipline and 2+ years of software development experience OR 4+ years of professional experience building software with site reliability or DevOps in lieu of a degree 1+ years of experience with Linux operating systems 1+ years of experience with Python and Python based development frameworks PREFERRED SKILLS AND EXPERIENCE: 2+ years of systems administration, site reliability engineering, or DevOps experience 2+ years of experience with Python and Python-based development frameworks 2+ years of Linux experience Expertise with Docker, Vagrant, and Kubernetes or similar technologies Extensive Experience with configuration management tools such as Ansible, Puppet, Terraform Experience with build systems (Make, Bazel / Pants / Buck, Gradle) and package management tools (pip, npm) Strong understanding of virtualization and hypervisor technologies Understanding of databases and data modeling Experience with automatically managing dozens or hundreds of servers Strong networking knowledge of TCP/IP Experience scaling web applications and optimizing applications for performance Experience with managing on-prem infrastructure, including direct experience managing GPU fleets Experience with high-performance computing systems or large-scale data analysis systems Must be comfortable working with mission-critical and sensitive systems, with a sense of urgency appropriate to the responsibilities Ability and willingness to obtain a Top Secret clearance ADDITIONAL REQUIREMENTS: An active clearance may provide the opportunity for you to work on sensitive SpaceX missions; if so, you will be subject to pre-employment drug and random drug and alcohol testing Willing to work extended hours and weekends when needed to meet critical deadlines COMPENSATION AND BENEFITS: Pay Range: Site Reliability Engineer/Level I: $125,000.00 - $145,000.00/per year Site Reliability Engineer/Level II: $145,000.00 - $175,000.00/per year Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience. Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock, stock options, or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan. You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per year. Employees accrue paid sick leave pursuant to Company policy which satisfies or exceeds the accrual, carryover, and use requirements of the law. ITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here . SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status. Applicants wishing to view a copy of SpaceX’s Affirmative Action Plan for veterans and individuals with disabilities, or applicants requiring reasonable accommodation to the application/interview process should reach out to EEOCompliance@spacex.com .

View more...
Cloud, DevOps & SRE
VerifiedToday

Site Reliability Engineer (High Performance Computing)

On-sitefull timeMid-LevelCalifornia, United States
Apply Now

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SITE RELIABILITY ENGINEER (HIGH PERFORMANCE COMPUTING) SpaceX HPC is a shared compute platform used across the company — vehicle and structures simulation, machine learning, AI inference, and more. We support every program at SpaceX to design and operate the worlds most advanced rockets and satellites. This role exists to put a real Site Reliability Engineer operating model on these capabilities and accelerating the world class engineering at SpaceX: toil reduction, automation, observability, and a sustainable incident process. We are looking for a Site Reliability Engineer who wants to own everything from Linux machines and our Infrastructure as Code, storage, and user facing applications – the whole ecosystem as a product, not as a ticket queue. You do not need a prior HPC title. You do need production instincts — you have operated real infrastructure, you write code to delete toil, and you care about whether users can actually get work done, not just whether nodes ping. You’ll work alongside HPC systems engineers who design and commission clusters to help make them more reliable and provide world class services for world class engineers. Aerospace experience is not required. We value engineers who treat teammates with fairness and respect, who are self-critical, and who will take ownership of hard production problems. RESPONSIBILITIES: Participate in the team's on-call rotation; practice sustainable incident response and blameless postmortems Manage node lifecycle with infrastructure as code: OS images, firmware, configuration management, kernel and driver stack Build observability for both HPC administrators and end users — cluster, node, and storage health for operators, and job/workflow-level signal for the people running work on the platform Reduce toil with automation; split time between operating production systems and writing the software that makes that work smaller Sustainably manage resources, including compute and storage Lead capacity planning with users across the company: understand what they will need next, and turn that into a concrete picture of tomorrow's compute and storage Collaborate with HPC systems engineers and with engineers across all disciplines across the company on operable, maintainable infrastructure BASIC QUALIFICATIONS: Bachelor's degree in computer science, engineering, math, or a scientific discipline; OR 2+ years of professional experience operating production infrastructure in lieu of a degree 2+ years of experience with Linux operating systems in production 2+ years of experience operating production infrastructure (servers, services, or networks), including monitoring, debugging, and repairing what you own PREFERRED SKILLS AND EXPERIENCE: 2+ years of professional experience in SRE, DevOps, or production infrastructure engineering Experience with monitoring and alerting (Prometheus, Grafana, Nagios, or similar) Experience deploying and maintaining configuration management or infrastructure as code (Ansible, Puppet, Terraform, or similar) Experience writing scripts/code (eg. Python or similar languages) to automate common tasks Experience with containers (Docker, Podman, Singularity/Apptainer) Experience with Kubernetes administration for on-premise deployment Experience with distributed or high-performance storage (VAST or similar), including capacity, performance, and lifecycle management Familiarity with HPC clusters, schedulers (Slurm, PBS, LSF), or GPU compute — not required; we will teach this Familiarity with scientific computing (CFD, FEA) and/or ML training workloads (PyTorch, TensorFlow, CUDA) and/or AI inference workloads Good understanding of version control, testing, continuous integration, build, deployment and monitoring Ability to communicate clearly with users, peers, and vendors in both incident and design settings Comfortable working with mission-critical and sensitive systems, with a sense of urgency appropriate to the responsibilities Eligibility for access to classified material up to TS/SCI with polygraph ADDITIONAL REQUIREMENTS: Position is based in Hawthorne, CA and is primarily on-site Must be able to participate in an on-call rotation Must be willing to work extended hours and weekends as needed for incidents, cluster bring-up, and time-critical failures COMPENSATION AND BENEFITS: Pay Range: Level 1: $125,000.00 - $160,000.00 Level 2: $145,000.00 - $195,000.00 Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience. Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan. You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per year. Employees accrue paid sick leave pursuant to Company policy which satisfies or exceeds the accrual, carryover, and use requirements of the law. ITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here . SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status. Applicants wishing to view a copy of SpaceX’s Affirmative Action Plan for veterans and individuals with disabilities, or applicants requiring reasonable accommodation to the application/interview process should reach out to EEOCompliance@spacex.com .

View more...
Cloud, DevOps & SRE
VerifiedToday

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SITE RELIABILITY ENGINEER — HPC & AUTOMATION (SILICON ENGINEERING) At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most advanced broadband internet system. Starlink is the world’s largest satellite constellation and is providing fast, reliable internet to millions of users worldwide. We design, build, test, and operate all parts of the system – thousands of satellites, consumer receivers that allow users to connect within minutes of unboxing, and the software that brings it all together. We’ve only begun to scratch the surface of Starlink’s potential global impact and are looking for best-in-class engineers to help maximize Starlink’s utility for communities and businesses around the globe. We are seeking a motivated, proactive, and intellectually curious engineer who will work alongside world-class cross-disciplinary teams (systems, firmware, architecture, design, validation, product engineering, ASIC implementation). As a Site Reliability Engineer on the Silicon Engineering team you will get the opportunity to design, operate, scale, and automate the high performance computing infrastructure we use to develop the chips powering the world's largest satellite constellation and a global internet service. This position will have a meaningful impact on Starlink silicon by enabling faster design-iterations, simulations, and regression turnaround times that gate how fast our chip teams can ship. RESPONSIBILITIES: Deploy, upgrade, operate, maintain, and scale our suite of clusters and services Collaborate with engineers to develop automated, full turnkey solutions for silicon simulation workflows to speed up project timelines Manage our underlying infrastructure as code and use modern observability tools to provide a complete picture of cluster and infrastructure health Operate the continuous integration pipeline, build and release systems, and version control across the environment Identify and eliminate performance bottlenecks using measurement and creative engineering BASIC QUALIFICATIONS: Bachelor’s degree in computer science, information systems, or an engineering discipline; OR 2+ years of professional experience in system administration, high performance computing, or site reliability engineering 1+ years of development experience with Bash, Python, and/or other programming languages 1+ years of experience with Linux operating systems PREFERRED SKILLS AND EXPERIENCE: Familiarity with containerization technologies (i.e. Docker, Kubernetes) Knowledge in computer system concepts (computer architecture, computer organization, operating systems and concurrency) Experience with databases and data modeling (e.g., MySQL, PostgreSQL, SQLite) Networking knowledge of TCP/IP Experience with high performance computing and workload managers (e.g., Slurm, LSF) Experience with Terraform, Ansible, Puppet, or similar automation frameworks Experience building monitoring and alerting as code (e.g., Grafana, Prometheus, custom exporters) Experience with CI/CD automation at scale (e.g., Jenkins, Bamboo, build systems) Experience with infrastructure as code (IaC) tools for managing fleets of servers Experience with using & building REST API clients/servers Experience with enterprise/networked storage automation (e.g., NetApp ONTAP REST API/CLI, NFS) Experience with ASIC design flows and tools (e.g., Cadence, Synopsys, Ansys, Keysight, Siemens) Strong desire to find performance bottlenecks and performance improvement techniques Excellent communication skills with the ability to communicate with customers, peers, management, etc. in both formal and informal situations Ability to quickly learn new tools and frameworks Interest in or experience with AI/LLM-assisted tooling (e.g., Grok, Claude Code) ADDITIONAL REQUIREMENTS: Ability to work extended hours and weekends as needed to meet critical milestones COMPENSATION AND BENEFITS: Pay Range: Level 1: $125,000.00 - $150,000.00 Level 2: $145,000.00 - $175,000.00 Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience. Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan. You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per year. Employees in Washington State accrue paid sick time in compliance with state and federal law. Company shuttles are offered to employees for roundtrip travel from select Seattle locations to the SpaceX Redmond office Monday to Friday. ITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here . SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status. Applicants wishing to view a copy of SpaceX’s Affirmative Action Plan for veterans and individuals with disabilities, or applicants requiring reasonable accommodation to the application/interview process should reach out to EEOCompliance@spacex.com .

View more...
Cloud, DevOps & SRE
VerifiedToday

Site Reliability Engineer, Kubernetes Platform (Starshield)

On-sitefull timeMid-LevelRedmond, United States
Apply Now

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SITE RELIABILITY ENGINEER, KUBERENTES PLATFORM (STARSHIELD) At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy the Starshield constellation. Starshield is the world’s largest US government satellite constellation and is tasked with providing immediate access to critical intelligence and national security data for the US government anywhere on the globe. We design, build, test, and operate all parts of the system – receivers that allow users to connect within minutes, and the software that brings it all together. We’ve only begun to scratch the surface of Starshield's global impact and are looking for best-in-class engineers to help us further our ambitious goals. As an engineer focused on Starshield's software and network infrastructure, you will design, operate and scale the infrastructure we use to run the world’s largest government satellite constellation. These positions cover a variety of areas ranging from Site Reliability Engineering, Developer Operations, and our internal Kubernetes platforms. You will develop automation to deploy and manage on-premise compute resources, create highly scalable and maintainable software products, and directly collaborate with engineering across the board. RESPONSIBILITES: Develop automation to deploy and manage on-premise Kubernetes clusters Deploy and manage core infrastructure such as databases, monitoring and distributed storage Closely collaborate with software engineers to create highly scalable, operable, and maintainable products Engage in and improve the whole lifecycle of services -- from inception and design, through deployment, operation and refinement Monitoring and alerting supporting systems to have high availability Hands-on integration and troubleshooting across the entire Starshield stack Identify areas for improvement and create innovative solutions that enable high system availability BASIC QUALIFICATIONS: Bachelor’s degree in computer science, information systems/IT, or an engineering discipline and 1+ years of professional experience in site reliability engineering or DevOps; OR 3+ years of professional experience in site reliability engineering or DevOps in lieu of a degree 1+ years of professional experience with Linux operating systems Experience with Terraform, Ansible, or other infrastructure tools Experience with containerization technologies (i.e. OCI containers, Kubernetes) Experience scripting in Bash, Python, or other similar languages Development experience in Python, C++, or Go PREFERRED SKILLS AND EXPERIENCE: 1+ years of experience with Python and Python-based development frameworks Experience managing Kubernetes clusters, not just using them Knowledge of Linux boot process and systems configuration Deep understanding of testing, continuous integration, build, deployment & continuous monitoring Understanding of relevant build technologies, such as Bazel and Makefiles Focus on performance bottlenecks and performance improvement techniques Understanding of distributed databases and data modeling Experience with automatically managing dozens, hundreds, or thousands of servers (eg: Terraform or Ansible) Strong networking knowledge of TCP/IP Excellent communications skills with the ability to communicate with customers, peers, management etc. in both formal and informal situations Active Top Secret, Top Secret SCI, or DOE Level Q clearance ADDITIONAL REQUIREMENTS: Must be willing to work extended hours and weekends as needed This position requires successfully obtaining and maintaining a Top Secret Security Clearance as a condition of employment. While the clearance may not be immediately necessary upon hire, we encourage you to initiate the application process promptly upon accepting this offer. Your ability to secure the necessary clearance is essential for fulfilling key responsibilities of the role. Should you be unable to obtain it, SpaceX reserves the right to modify or terminate your employment to align with operational needs. COMPENSATION AND BENEFITS: Pay Range: Level 1: $125,000.00 - $160,000.00 Level 2: $145,000.00 - $195,000.00 Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience. Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan. You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per year. Employees in Washington State accrue paid sick time in compliance with state and federal law. Company shuttles are offered to employees for roundtrip travel from select Seattle locations to the SpaceX Redmond office Monday to Friday. ITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here . SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status. Applicants wishing to view a copy of SpaceX’s Affirmative Action Plan for veterans and individuals with disabilities, or applicants requiring reasonable accommodation to the application/interview process should reach out to EEOCompliance@spacex.com .

View more...
Cloud, DevOps & SRE
VerifiedToday

Site Reliability Engineer, Kubernetes Platform (Starshield)

On-sitefull timeMid-LevelCalifornia, United States
Apply Now

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SITE RELIABILITY ENGINEER, KUBERENTES PLATFORM (STARSHIELD) At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy the Starshield constellation. Starshield is the world’s largest US government satellite constellation and is tasked with providing immediate access to critical intelligence and national security data for the US government anywhere on the globe. We design, build, test, and operate all parts of the system – receivers that allow users to connect within minutes, and the software that brings it all together. We’ve only begun to scratch the surface of Starshield's global impact and are looking for best-in-class engineers to help us further our ambitious goals. As an engineer focused on Starshield's software and network infrastructure, you will design, operate and scale the infrastructure we use to run the world’s largest government satellite constellation. These positions cover a variety of areas ranging from Site Reliability Engineering, Developer Operations, and our internal Kubernetes platforms. You will develop automation to deploy and manage on-premise compute resources, create highly scalable and maintainable software products, and directly collaborate with engineering across the board. RESPONSIBILITES: Develop automation to deploy and manage on-premise Kubernetes clusters Deploy and manage core infrastructure such as databases, monitoring and distributed storage Closely collaborate with software engineers to create highly scalable, operable, and maintainable products Engage in and improve the whole lifecycle of services -- from inception and design, through deployment, operation and refinement Monitoring and alerting supporting systems to have high availability Hands-on integration and troubleshooting across the entire Starshield stack Identify areas for improvement and create innovative solutions that enable high system availability BASIC QUALIFICATIONS: Bachelor’s degree in computer science, information systems/IT, or an engineering discipline and 1+ years of professional experience in site reliability engineering or DevOps; OR 3+ years of professional experience in site reliability engineering or DevOps in lieu of a degree 1+ years of professional experience with Linux operating systems Experience with Terraform, Ansible, or other infrastructure tools Experience with containerization technologies (i.e. OCI containers, Kubernetes) Experience scripting in Bash, Python, or other similar languages Development experience in Python, C++, or Go PREFERRED SKILLS AND EXPERIENCE: 1+ years of experience with Python and Python-based development frameworks Experience managing Kubernetes clusters, not just using them Knowledge of Linux boot process and systems configuration Deep understanding of testing, continuous integration, build, deployment & continuous monitoring Understanding of relevant build technologies, such as Bazel and Makefiles Focus on performance bottlenecks and performance improvement techniques Understanding of distributed databases and data modeling Experience with automatically managing dozens, hundreds, or thousands of servers (eg: Terraform or Ansible) Strong networking knowledge of TCP/IP Excellent communications skills with the ability to communicate with customers, peers, management etc. in both formal and informal situations Active Top Secret, Top Secret SCI, or DOE Level Q clearance ADDITIONAL REQUIREMENTS: Must be willing to work extended hours and weekends as needed This position requires successfully obtaining and maintaining a Top Secret Security Clearance as a condition of employment. While the clearance may not be immediately necessary upon hire, we encourage you to initiate the application process promptly upon accepting this offer. Your ability to secure the necessary clearance is essential for fulfilling key responsibilities of the role. Should you be unable to obtain it, SpaceX reserves the right to modify or terminate your employment to align with operational needs. COMPENSATION AND BENEFITS: Pay Range: Level 1: $125,000.00 - $160,000.00 Level 2: $145,000.00 - $195,000.00 Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience. Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan. You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per year. Employees accrue paid sick leave pursuant to Company policy which satisfies or exceeds the accrual, carryover, and use requirements of the law. ITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here . SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status. Applicants wishing to view a copy of SpaceX’s Affirmative Action Plan for veterans and individuals with disabilities, or applicants requiring reasonable accommodation to the application/interview process should reach out to EEOCompliance@spacex.com .

View more...
Cloud, DevOps & SRE
VerifiedToday

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SITE RELIABILITY ENGINEER, KUBERNETES PLATFORM (TOP SECRET CLEARANCE) As a member of the Classified IT Systems Engineering team, the Site Reliability Engineer is involved in designing scalable systems capable of supporting a growing volume of data products being generated in mass. We build tools that enable us to work more efficiently, and that help us build software systems that are secure, reliable, and autonomous. Our engineers are responsible for the life cycle of the systems they create, including development, testing, and operational support. RESPONSIBILITIES: Develop automation to deploy and manage compute resources both on-premises and in the cloud Build, maintain, and scale on-premises hardware systems designed to host GPU-accelerated machine learning workloads Deploy and manage core infrastructure such as databases, monitoring and storage Closely collaborate with software engineers to create highly scalable, operable and maintainable products Engage in and improve the whole lifecycle of services -- from inception and design, through deployment, operation and refinement BASIC QUALIFICATIONS: Bachelor’s degree in computer science, information systems, or an engineering discipline; OR 2+ years of professional experience in software, DevOps, or site reliability engineering in lieu of a degree 1+ year of experience with Kubernetes 1+ year of experience with Linux operating systems Experience in Bash, Python, and/or other scripting languages Experience building, maintaining, and scaling on-premises and/or cloud systems Active Top Secret, Top Secret SCI, or DOE Level Q clearance PREFERRED SKILLS AND EXPERIENCE: Experience hosting and pushing the state of the art in inferential model benchmarks Experience with systems administration, site reliability engineering, or DevOps engineering Experience with Python and Python-based development frameworks Experience with virtualization and hypervisor technologies Experience with automatically managing dozens or hundreds of servers Knowledge of performance bottlenecks and performance improvement techniques Excellent communications skills with the ability to communicate with customers, peers, management etc. in both formal and informal situations Ability to quickly learn new tools and frameworks. ADDITIONAL REQUIREMENTS: An active clearance may provide the opportunity for you to work on sensitive SpaceX missions; if so, you will be subject to pre-employment drug and random drug and alcohol testing Must be willing to work extended hours and weekends as needed COMPENSATION AND BENEFITS: Pay Range: Level 3: $145,000.00 - $195,000.00 Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience. Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan. You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per year. Employees accrue paid sick leave pursuant to Company policy which satisfies or exceeds the accrual, carryover, and use requirements of the law. Those with an active clearance will receive a 10% differential, up to an additional $20,000 annually, once officially briefed into a classified program. ITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here . SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status. Applicants wishing to view a copy of SpaceX’s Affirmative Action Plan for veterans and individuals with disabilities, or applicants requiring reasonable accommodation to the application/interview process should reach out to EEOCompliance@spacex.com .

View more...
Cloud, DevOps & SRE
VerifiedToday

Site Reliability Engineer (Manufacturing Infrastructure)

On-sitefull timeMid-LevelTexas, United States
Apply Now

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SITE RELIABILITY ENGINEER (MANUFACTURING INFRASTRUCTURE) The application software team is the central nervous system of SpaceX. Manufacturing is how SpaceX turns designs into hardware. The compute, storage, and networking that run our factories must be as reliable as the products we build. This team owns infrastructure supporting Starship, Starlink, Starshield, and Terafab. This position will have a direct impact on factory uptime, throughput, and production scale across programs. The ideal candidate has strong software engineering fundamentals and a passion for infrastructure: reliability, stability, proactive maintenance, and scalability. You understand the system before you change it, solve hard problems, communicate clearly with stakeholders and teammates, and take ownership of work that manufacturing depends on. Aerospace experience is not required. We value smart, motivated, collaborative engineers who treat teammates with fairness, respect, and support, and who want to take full ownership of challenging problems to help make humanity multi-planetary. RESPONSIBILITIES: Deploy, upgrade, operate, maintain, and scale compute, storage, and networking for manufacturing systems across Starship, Starlink, Starshield, and Terafab Manage infrastructure as code and use observability to provide a complete picture of platform health Design for reliability, stability, and scale; find and remove bottlenecks with measurement and engineering Practice proactive maintenance: capacity planning, lifecycle management, and reducing toil before it becomes an incident Partner with software engineers, manufacturing stakeholders, and site teams to build operable, maintainable systems Improve the full lifecycle—from design through deployment, operation, and continuous refinement Practice sustainable incident response and blameless postmortems Provide high-quality support to manufacturing and engineering users Communicate clearly with stakeholders and teammates Participate in on-call and travel to sites as needed for deployments, incidents, and cross-site reliability BASIC QUALIFICATIONS: Bachelor’s degree in computer science, information systems, or an engineering discipline; OR 3+ years of professional experience in SRE or DevOps in lieu of a degree 1+ years of software development experience Experience with Linux operating systems PREFERRED SKILLS AND EXPERIENCE: Experience with compute, storage, and/or networking infrastructure in production Infrastructure as code (Terraform, Ansible, Puppet, or similar) Containers and virtualization (Docker, Kubernetes, vSphere, QEMU, KVM, etc.) Databases and data modeling (Postgres, Clickhouse, etc.) Ability to translate high-level requirements into implementations from first principles Comfort with mission-critical systems and appropriate urgency and care Skillful communication with customers, peers, and management Comfort operating across multiple sites and manufacturing programs ADDITIONAL REQUIREMENTS: Must be able to work extended hours and weekends as needed Must be able to travel to different sites (Hawthorne, CA; Redmond, WA; Cape Canaveral, FL; Starbase, TX) Ability to pass Air Force background check for Cape Canaveral This role requires you to be onsite. Remote and/or hybrid work will not be considered ITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here . SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status. Applicants wishing to view a copy of SpaceX’s Affirmative Action Plan for veterans and individuals with disabilities, or applicants requiring reasonable accommodation to the application/interview process should reach out to EEOCompliance@spacex.com .

View more...
Cloud, DevOps & SRE
VerifiedToday

Site Reliability Engineer (Raptor)

On-sitefull timeMid-LevelCalifornia, United States
Apply Now

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SITE RELIABILITY ENGINEER (RAPTOR) SpaceX is looking for a Site Reliability Engineer with a strong drive to solve challenging problems in the Raptor engine organization. You will be empowered to solve a wide range of systems engineering problems – including High Performance Computing, System performance, networking, manufacturing infrastructure – with the singular goal of accelerating the pace of rocket engine development and delivering excellent results. You will work with engineering, analysis, IT, and facilities teams to identify and solve fundamental challenges to scaling the department engineering output. RESPONSIBILITIES: Manage server infrastructure, HPC systems, storage systems, networks, and high-speed interconnect (Infiniband). Design, procure, and integrate infrastructure systems. Work with propulsion engineering staff to solve critical bottlenecks. Coordinate and communicate with company -wide infrastructure and IT teams. Support application deployment (ANSYS, StarCCM+, manufacturing software) for best performance of software on real-world systems. BASIC QUALIFICATIONS: 1+ years of hands-on experience with client and server hardware/software, management tools, enterprise networking, virtualization, and security technologies. Bachelor's degree in computer science, engineering, math, or scientific discipline; OR 2+ years of professional experience building software in lieu of a degree. Experience with Linux and Windows server software. PREFERRED SKILLS AND EXPERIENCE: 1+ year of systems engineering experience Experience with scripting languages (Bash, Python), automation (Puppet, Ansible), and other common sysadmin tools. Experience building, deploying, and troubleshooting large-scale compute systems. Familiarity with resource development and management (Kubernetes, Docker). Familiarity with engineering and analysis applications, such as CFD and FEA. Familiarity with diagnosing bottlenecks and designing systems for performance. Able to work effectively in a dynamic environment while assuming high levels of responsibility and demonstrating accountability for rocket engine-level outcomes. COMPENSATION AND BENEFITS: Pay range: Site Reliability Engineer/Level 1: $125,000.00 - $150,000.00/per year Site Reliability Engineer/Level 2: $145,000.00 - $175,000.00/per year Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience. Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock, stock options, or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan. You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per year. Employees accrue paid sick leave pursuant to Company policy which satisfies or exceeds the accrual, carryover, and use requirements of the law. ITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here . SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status. Applicants wishing to view a copy of SpaceX’s Affirmative Action Plan for veterans and individuals with disabilities, or applicants requiring reasonable accommodation to the application/interview process should reach out to EEOCompliance@spacex.com .

View more...
Cloud, DevOps & SRE
VerifiedToday

AI Engineer, Platform Infrastructure, Special Programs

On-sitefull timeMid-LevelPalo Alto, United States
Apply Now

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. AI ENGINEER, PLATFORM INFRASTRUCTURE, SPECIAL PROGRAMS As an AI Engineer, Platform Infrastructure you will build the tooling, and work with our cleared engineers at classified sites execute it. The quality of what you build directly determines how effectively those engineers can operate in environments where they cannot call you for help. Your tooling must be deterministic, complete, well-tested, and foolproof. RESPONSIBILITIES: Design and build the deployment generator. Create and maintain platform documentation. Build the bundle pipeline. Implement profile-driven rendering for public cloud, enterprise on-prem, and classified air-gap targets based on profile selection. Build and maintain the cross-profile CI matrix. Build the testing and validation framework. Integrate with the existing Supercompute team's operators, Helm charts, and CRDs by consuming their work as inputs to the generator without modifying production infrastructure. Own the monitoring stack migration. BASIC QUALIFICATIONS: Bachelor's degree in computer science, mathematics, computer engineering, physics, data science, or engineering discipline. 1+ years of experience in Go Lang and/or Python. 1+ years of experience working with Kubernetes or similar tooling for automating, developing and scaling containerized applications. PREFERRED SKILLS AND EXPERIENCE: Ability to obtain and maintain a Top Secret or Top Secret SCI clearance. Experience building CI/CD pipelines and deployment automation (Buildkite, GitHub Actions, or equivalent). Familiarity with container image management. Experience with Infrastructure-as-Code (Pulumi or Terraform) and understanding of state management tradeoffs. Understanding of Linux systems. Experience building air-gapped or disconnected deployment tooling where all dependencies must be pre-staged. Familiarity with GPU infrastructure, including NVIDIA drivers, CUDA, NCCL, InfiniBand/RoCE networking. Experience with PXE boot, cloud-init, squashfs, or other bare metal provisioning technologies. Own the bare metal provisioning pipeline. ADDITIONAL REQUIREMENTS: Must be willing to work extended hours and weekends as needed. 20% travel may be required to government sites. This position requires successfully obtaining and maintaining a Top Secret Security Clearance as a condition of employment. While the clearance may not be immediately necessary upon hire, we encourage you to initiate the application process promptly upon accepting this offer. Your ability to secure the necessary clearance is essential for fulfilling key responsibilities of the role. Should you be unable to obtain it, SpaceX reserves the right to modify or terminate your employment to align with operational needs. COMPENSATION AND BENEFITS: Pay range: AI Engineer: $135,000 - $230,000/per year Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience. Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock, stock options, or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan. You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per year. Employees accrue paid sick leave pursuant to Company policy which satisfies or exceeds the accrual, carryover, and use requirements of the law. ITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here . SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status. Applicants wishing to view a copy of SpaceX’s Affirmative Action Plan for veterans and individuals with disabilities, or applicants requiring reasonable accommodation to the application/interview process should reach out to EEOCompliance@spacex.com .

View more...
AI / ML & Data Science
VerifiedToday

Page 24 of 52