LogoKode$word
Cerebras logo
Verified Tech Organization

Careers at Cerebras

Browse and filter through all verified positions currently open at Cerebras.

Total Company Roles23
Matching Filter23
cerebras.aiHQ: Sunnyvale, CA, USCEO: Andrew D. Feldman708 employees

Cerebras Systems Inc. is a leading innovator in artificial intelligence infrastructure. The company develops and manufactures an advanced AI compute platform, integrating proprietary hardware systems and software. This platform is delivered in rack-mountable units, suitable for deployment in data centers, scaling all the way up to supercomputer-level capabilities. At its core is the groundbreaking Wafer-Scale Engine (WSE), a unique chip that encompasses an entire silicon wafer. This innovation is specifically engineered to deliver superior performance and speed compared to conventional GPUs, addressing the intensive computational demands of inference, Generative AI, and a broad spectrum of other AI applications. Cerebras serves a diverse clientele, including leading hyperscalers, advanced foundation model laboratories, AI-native and digital-first businesses, large enterprises, and key players in Sovereign AI initiatives. With operations spanning the United States, Europe, the Middle East, Africa, and other international markets, Cerebras maintains a significant global footprint. The company was established in 2015 and is headquartered in Sunnyvale, California.

Sector:Semiconductors

All Openings (23)

Ordered by most recently published

Senior ASIC RTL Design Engineer, AI Hardware

On-sitefull timeLead / StaffSunnyvale, United States
Apply Now

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. About The Role: As a senior front-end design engineer, you will be a key part of the world-class team designing and developing the next generations of the Cerebras Wafer Scale Engine (WSE). This role requires deep expertise in RTL design and integration, with a strong focus on delivering high-performance, power-efficient, and scalable solutions. You will collaborate closely with the design verification, physical design, software and system teams to bring innovative semiconductor architectures from concept to production, addressing the unique challenges of building WSE systems. Responsibilities: Drive all aspects of chip design, including Functional Specification, Micro-architecture, RTL development, Synthesis. Work closely with PD team members for design closure to meet PPA goals. Work closely with Design verification and DFT teams for achieving the best functional and test coverage. Work with software and system teams to understand opportunities to deliver optimal performance and feature set for the product. Debug silicon-level functional, timing, and power issues during bring up. Requirements Master’s degree in Computer Science, Electrical Engineering, or equivalent. Can work in a hybrid work environment. 8+ years of experience in delivering complex, high performance high quality RTL designs. Experience with Front End Chip integration and third-party IP integration. Demonstrated experience in high-performance computing, GPU/CPU designs, machine learning or related fields. Proven track record of multiple silicon success. Experience collaborating with external vendors. Networking stack experience including TCP/IP, RDMA and Ethernet. Knowledge of PCIe, CPU interfaces and Serdes technology. Working knowledge of scripting tools : Python, TCL. Assets: Experience with FPGA development toolchain, including Place and Route, Floor planning and Timing Analysis is a plus. Experience managing external ASIC vendor through product development cycle. Location: Sunnyvale, CA (Open to Remote) The base salary range for this position is $200,000 to $290,000 annually. Actual compensation may include bonus and equity, and will be determined based on factors such as experience, skills, and qualifications. Why Join Cerebras People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras: Build a breakthrough AI platform beyond the constraints of the GPU. Publish and open source their cutting-edge AI research. Work on one of the fastest AI supercomputers in the world. Enjoy job stability with startup vitality. Our simple, non-corporate work culture that respects individual beliefs. Find out more about what it's like to work at Cerebras here ! Apply today and become part of the forefront of groundbreaking advancements in AI! Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them. This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

View more...
Hardware & EmbeddedVia Ashby
Verified4 days ago

Staff Software Engineer, Inference API

On-sitefull timeLead / StaffToronto, Canada
Apply Now

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. About the Role Cerebras is building a new generation of disaggregated AI inference systems that combine GPU-accelerated prefill with ultra-fast decode on the Cerebras Wafer-Scale Engine. We are hiring a Software Engineer to build and evolve the ML API layer that makes this heterogeneous serving system accessible, reliable, and easy to use. You will work across our inference APIs, model integration layer, request-routing services and Cerebras inference platform to deliver a consistent experience across models and accelerator backends. This role sits at the intersection of machine learning systems, API design, model serving, and distributed systems. You will enable new model architectures and inference capabilities, define stable user-facing behavior, and ensure that features such as streaming, sampling, tool use, structured outputs, multimodal inputs, and model configuration behave correctly and consistently in production. You will work closely with model enablement, compiler, runtime, cloud infrastructure, product, customer-facing teams, and customers directly. This is a hands-on software engineering role for someone who enjoys turning rapidly evolving ML capabilities into durable, production-quality APIs. Responsibilities Build production ML inference APIs. Design, implement, and maintain APIs for chat completions, text generation, streaming, model configuration, tool calling, structured outputs, multimodal inputs, and other emerging inference capabilities. Deliver a unified serving experience. Create consistent request and response semantics across GPU prefill, Cerebras decode, and other heterogeneous inference backends. Enable new models and capabilities. Integrate emerging foundation models, tokenizers, prompt formats, sampling methods, attention variants, multimodal inputs, and model-specific features into the serving platform. Own API compatibility and evolution. Maintain compatibility with widely adopted inference interfaces while designing Cerebras-specific extensions. Establish clear versioning, deprecation, validation, and backward compatibility practices. Integrate with model-serving runtimes. Extend and integrate custom inference services with vLLM, PyTorch, Hugging Face libraries, the AMD ROCm stack, and Cerebras runtime components. Support disaggregated inference. Build the control and data paths required to coordinate GPU prefill with Cerebras decode, including request routing, state transfer, error handling, retries, and lifecycle management. Improve serving performance. Optimize streaming behavior, time to first token, request latency, throughput, batching, serialization, tokenization, scheduling, and communication between serving components. Ensure functional and numerical correctness. Build validation systems for tokenization, sampling, logits, generated outputs, precision changes, model upgrades, determinism, and compatibility across serving backends. Strengthen reliability and observability. Define end-to-end service indicators and build structured logging, tracing, metrics, dashboards, health checks, and diagnostic tooling for production inference traffic. Develop testing and qualification infrastructure. Create conformance tests, workload-replay tools, model-validation suites, performance benchmarks, integration tests, and release gates. Improve developer experience. Build intuitive configuration, SDKs, documentation, examples, debugging tools, and self-service workflows for internal developers, customers, and partners. Collaborate across the stack. Partner with compiler, runtime, kernel, cloud, product, and solutions teams to translate model and customer requirements into scalable serving capabilities. Minimum Qualifications 5+ years of software engineering experience, including substantial individual-contributor ownership of production software or distributed systems. Strong programming ability in Python and Go plus experience developing performance-sensitive or highly concurrent services in C++, Rust, or a similar systems language. Experience building stable APIs with clear validation, error handling, observability, compatibility, and versioning practices. Experience integrating software across service, framework, runtime, and infrastructure boundaries. Experience designing or maintaining OpenAI-compatible, gRPC, REST, or streaming inference APIs. Experience with Linux, containers, Kubernetes or comparable orchestration systems, CI/CD, and operating latency-sensitive services in production. Ability to diagnose correctness, reliability, and performance issues across multiple components of a distributed serving system. Strong communication and cross-functional execution skills, with the ability to turn ambiguous model or product requirements into production-quality software. Bachelor's degree in computer science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience. Preferred Qualifications Experience modifying or contributing to vLLM, SGLang, PyTorch, Hugging Face Transformers, Triton, TensorRT-LLM, or another open-source ML systems project. Experience creating API conformance, model-quality, numerical-comparison, determinism, or performance-regression test systems Experience building SDKs, developer tools, model registries, configuration systems, or self-service ML platforms. Experience with multi-model or multi-tenant inference platforms, including routing, admission control, fairness, quotas, rate limiting, and capacity-aware scheduling Understanding of model-specific tokenization, chat templates, generation configuration, logits processing, stopping criteria, tool calling, structured generation, and constrained decoding. Experience with disaggregated prefill/decode architectures, KV-cache transfer, prefix caching, chunked prefill, memory-aware admission control, or request scheduling. Experience designing, building, or operating production APIs and services for machine learning, large language models, or other data-intensive applications. Familiarity with reduced-precision inference and quantization formats such as BF16, FP8, FP4, INT8, or INT4. Why Join Cerebras People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras: Build a breakthrough AI platform beyond the constraints of the GPU. Publish and open source their cutting-edge AI research. Work on one of the fastest AI supercomputers in the world. Enjoy job stability with startup vitality. Our simple, non-corporate work culture that respects individual beliefs. Find out more about what it's like to work at Cerebras here ! Apply today and become part of the forefront of groundbreaking advancements in AI! Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them. This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

View more...
Software EngineeringVia Ashby
Verified5 days ago

Staff AI Engineer – Business Systems

On-sitefull timeLead / StaffSunnyvale, United States
Apply Now

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. Hands-on AI engineering, solution architecture and compliance-by-design for enterprise Finance, operations and business systems Responsibilities The role is accountable for hands-on delivery and architecture within its layer, with shared governance across BIS, Finance, Business Operations, IT and Security and active partnership with other enterprise functions. AI solution architecture Design end-to-end agentic solutions and determine when a use case should query a source system directly versus use the unified data model. Partner with stakeholders to identify high-value use cases, translate requirements into controlled AI workflows and select AI, conventional automation or no new technology. Create reusable architecture patterns for agents, tools, APIs, MCP servers, prompts, evaluations and human-review workflows. Produce solution designs, security flows, deployment patterns and technical standards. AI engineering and system enablement Build AI agents, orchestration services, enterprise applications and reusable platform components. Deliver workflows for close and reporting, procurement, forecasting, billing and compliance monitoring where AI adds measurable value. Establish secure, primarily read-only AI connections to approved business systems, beginning with NetSuite and extending to adjacent Finance and enterprise platforms as priorities evolve. Preserve source-system authentication, authorization, user-level entitlements, rate limits and audit trails. Implement citations, evidence links, deterministic checks, exception handling and safe action boundaries. Prototype-to-enterprise delivery Assess business-built or rapidly developed prototypes for value, architecture, security, maintainability and control readiness. Refactor or rebuild approved prototypes into tested, monitored and supportable enterprise applications. Establish development, test and production environments, release pipelines, incident response and rollback controls. AI platform strategy Evaluate AI models, agent frameworks, connectors and enterprise platforms on a regular cadence. Run structured proofs of concept and assess security, accuracy, integration, scalability, experience, cost and vendor viability. Maintain platform standards and recommend adoption, retention, replacement or retirement decisions. Organizational enablement and adoption Create clear documentation, reusable patterns and reference architectures; coach teams on effective agent design, prompts, evaluation practices and safe operating boundaries. Establish feedback loops with users and process owners; use adoption, task success, efficiency, trust and support signals to guide iteration. Finance, SOX and compliance Translate Finance, Security, Privacy, SOX and SSDLC requirements into technical architecture and application controls. Implement least privilege, segregation of duties, logging, retention, evaluation, change control and audit evidence. Require deterministic validation and reconciliation for financially material outputs. Support SOX walkthroughs, control testing, audits, risk assessments and remediation while escalating formal approval to control owners. CANDIDATE PROFILE Qualifications, success measures and boundaries Required capabilities are calibrated for a Staff-level hands-on engineer with solution-architecture responsibilities. Required qualifications 8+ years in software, platform, integration, solution engineering or enterprise applications, including meaningful hands-on production ownership in complex environments. Strong Python and/or TypeScript skills; experience with APIs, MCP or comparable tool protocols, enterprise authentication and distributed-system design. Practical experience building production AI systems using agents, tool use, retrieval, structured outputs, evaluations and monitoring. Practical familiarity with leading LLM platforms and agent frameworks, such as OpenAI, Anthropic, Gemini, LangChain, Semantic Kernel or comparable technologies, including prompt and context engineering. Strong solution-architecture judgment across security, reliability, performance, cost, observability and supportability. Working knowledge of enterprise Finance processes such as general ledger, close, reporting, procure-to-pay, order-to-cash, forecasting and management reporting. Working knowledge of compliance-by-design, including access, segregation of duties, change management, interfaces, automated controls, completeness and accuracy, and audit evidence. Ability to communicate with engineers, Finance leaders, control owners, Security and executives. Preferred qualifications Experience with ERPs, data platforms, frontier AI platforms, agent frameworks or comparable enterprise technologies. Experience building internal enterprise applications. Hands-on experience implementing SOX controls or operating in a public-company or audit-regulated environment. Success measures Time from approved use case to controlled production and sustained adoption, with evidence of measurable business value. Reduction in manual effort and business-process cycle time; improvement in decision quality or service levels. Accuracy, groundedness, reconciliation success and production reliability of deployed agents. User adoption, task success, stakeholder trust and support burden for production workflows. Latency, operating cost and cost per successful task for deployed agents and applications. Reuse of approved architecture patterns and components across use cases. Number of viable prototypes transitioned into governed enterprise solutions. Security, SOX and audit findings; evidence completeness; incident rate and remediation time. Quality and timeliness of AI platform evaluations and roadmap recommendations. Why Join Cerebras People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras: Build a breakthrough AI platform beyond the constraints of the GPU. Publish and open source their cutting-edge AI research. Work on one of the fastest AI supercomputers in the world. Enjoy job stability with startup vitality. Our simple, non-corporate work culture that respects individual beliefs. Find out more about what it's like to work at Cerebras here ! Apply today and become part of the forefront of groundbreaking advancements in AI! Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them. This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

View more...
AI / ML & Data ScienceVia Ashby
Verified7 days ago

Network Security Engineer (Remote)

Remotefull timeSeniorUnited States (Remote)
Apply Now

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. Cerebras is seeking a Network Security Engineer to build, operate, and secure the network infrastructure supporting our data centers, cloud environments, corporate network, and AI systems. This is a remote, U.S.-based individual contributor role working across network security and network operations. You will support firewalls, segmentation, routing and switching, network access, and automation while collaborating with on-site data center and infrastructure teams. Occasional travel to Cerebras offices or data centers may be required. Responsibilities Operate firewalls, ACLs, segmentation, and network security controls across data center, corporate, and AWS environments. Remotely support data center network deployment, configuration, troubleshooting, maintenance, and lifecycle management in partnership with on-site teams. Implement and maintain segmentation, firewall, and ACL policies across corporate, compute, customer, and infrastructure environments. Manage inbound and outbound network controls, including public exposure, NAT, load balancers, DNS filtering, proxies, and egress policies. Implement and operate VPN, ZTNA, NAC, Wi-Fi, and vendor or partner connectivity. Perform recurring firewall rule reviews, segmentation audits, and remediation of unnecessary network exposure. Automate network and security workflows using Terraform, Ansible, GitOps, Python, and policy as code. Support network telemetry, detection, investigation, and containment in partnership with Security Operations. Troubleshoot routing, switching, TCP/IP, DNS, TLS, and cloud networking issues. Maintain network architecture documentation, procedures, and runbooks. Coordinate remote changes and incident response across distributed infrastructure and engineering teams. Skills and Qualifications 7+ years of experience in network security, network engineering, cloud security, or infrastructure security. Strong hands-on experience with firewalls, ACLs, segmentation, routing, switching, VPNs, and network troubleshooting. Experience supporting data center and on-premises network infrastructure, including working effectively with remote hands. Experience with AWS networking, including VPCs, transit gateways, security groups, and load balancers. Experience with Palo Alto, Juniper, Cloudflare, or similar platforms. Proficiency with Terraform, Ansible, Python, GitOps, or similar automation tools. Strong understanding of ZTNA, NAC, egress controls, TCP/IP, DNS, and TLS. Strong written communication and documentation skills. Ability to operate independently in a remote environment and collaborate across time zones. Ability to travel occasionally to Cerebras offices or data centers as business needs require. Relevant Experience Experience in several of the following is valuable: AI, HPC, or large-scale compute environments. Data center networking and security. AWS network security and segmentation. ZTNA and VPN architectures. DNS filtering, SWG, proxy, or SASE platforms. Vendor and partner connectivity. Public exposure and attack surface remediation. Network detection and incident response. Remote operation of geographically distributed network infrastructure. Why Join Cerebras People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras: Build a breakthrough AI platform beyond the constraints of the GPU. Publish and open source their cutting-edge AI research. Work on one of the fastest AI supercomputers in the world. Enjoy job stability with startup vitality. Our simple, non-corporate work culture that respects individual beliefs. Find out more about what it's like to work at Cerebras here ! Apply today and become part of the forefront of groundbreaking advancements in AI! Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them. This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

View more...
CybersecurityVia Ashby
Verified8 days ago

Distributed Systems Security Engineer

On-sitefull timeLead / StaffSunnyvale, United States
Apply Now

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. Cerebras is seeking a Distributed Systems Security Engineer to help build out the next generation of security controls and services in our on-prem and cloud clusters. In this role you will be responsible for driving hardening efforts across every layer of the cluster stack, covering a range of systems at the bare metal, virtualized, and k8s containerized levels. The ideal candidate brings a combination of wide breadth of knowledge on a variety of systems and networking topics, a pragmatic and results-driven outlook on security, and the flexibility and attention to detail to work in a fast-paced development environment without compromising our security standards. Responsibilities Define core security requirements for distributed systems environments, covering everything from secure boot configurations to workload identity in k8s pods Partner with platform, infrastructure, networking, and hardware teams to identify gaps and deliver end-to-end solutions improving the overall security maturity of the environment Develop threat models for each part of the cluster stack, outlining trust boundaries and the interplay between first and third-party mechanisms Work with engineering leads and architects to bake secure design in from the beginning Balance engineering outcomes with security controls, enabling the company to deliver securely quickly Document security posture and strategy for technical and nontechnical audiences from both 1p and 3p perspectives Skills and Qualifications 8-10 years of experience with heterogenous compute environments, kubernetes hardening, and general networked system security 10+ years with HPCs (high performance clusters) Strong hands-on engineering abilities in on-prem and cloud environments Ability to work in fragmented ecosystems, hardening both modern k8s deployments as well as bare metal instances Familiarity with AWS, Kubernetes, and distributed compute systems a plus Strong written communication skills, with the ability to make security concepts approachable for non-specialists. Why Join Cerebras People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras: Build a breakthrough AI platform beyond the constraints of the GPU. Publish and open source their cutting-edge AI research. Work on one of the fastest AI supercomputers in the world. Enjoy job stability with startup vitality. Our simple, non-corporate work culture that respects individual beliefs. Find out more about what it's like to work at Cerebras here ! Apply today and become part of the forefront of groundbreaking advancements in AI! Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them. This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

View more...
CybersecurityVia Ashby
Verified8 days ago

Software Supply Chain Security Engineer

On-sitefull timeLead / StaffSunnyvale, United States
Apply Now

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. Cerebras is seeking a Software Supply Chain Security Engineer to help build a secure foundation for our build and development workflows. In this role you will be responsible for assessing and hardening every layer of our source to artifact pipeline, from how we store and manage code to what we deploy into our clusters. The ideal candidate brings a combination of wide breadth of knowledge on a variety of software supply chain topics throughout the whole SDLC process, a pragmatic and results-driven outlook on security, and the flexibility and attention to detail to work in a fast-paced development environment without compromising our security standards. Responsibilities Define core security requirements for artifact build pipelines including maturing SLSA processes, and see them through from ideation to implementation Partner with platform, infrastructure, networking, and hardware teams to identify supply chain gaps and deliver end-to-end solutions for ensuring confidence in CI/CD artifacts Develop threat models for each layer of the build and deploy stack, outlining trust boundaries and the interplay between first and third-party mechanisms Work with engineering leads and devops architects to build in secure supply chain principles from the beginning Balance engineering outcomes with security controls, enabling the company to deliver securely quickly Document security posture and strategy for technical and nontechnical audiences from both 1p and 3p perspectives Skills and Qualifications 8-10 years of experience with software supply chain and SDLC topics including secure code management, build pipeline hardening, artifact signing and attestation, and the end-to-end integration thereof. Masters in Computer Science or PhD (preferred) Strong hands-on engineering abilities in devops and SDLC contexts Ability to work in fragmented and legacy build environments, experience with modernizing stacks is a plus Familiarity with AWS, Jenkins, SLSA, and best practices for secure artifact provenance and builds are a plus Strong written communication skills, with the ability to make security concepts approachable for non-specialists. Direct work with Agentic workflows and deep understanding of LLM and reasoning cycles. Why Join Cerebras People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras: Build a breakthrough AI platform beyond the constraints of the GPU. Publish and open source their cutting-edge AI research. Work on one of the fastest AI supercomputers in the world. Enjoy job stability with startup vitality. Our simple, non-corporate work culture that respects individual beliefs. Find out more about what it's like to work at Cerebras here ! Apply today and become part of the forefront of groundbreaking advancements in AI! Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them. This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

View more...
CybersecurityVia Ashby
Verified8 days ago

Software Engineer, Kernel Reliability

On-sitefull timeMid-LevelUnited States
Apply Now

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. About The Role We're looking for a deeply technical, hands-on software engineer to join our on-field Kernel Reliability team. You'll help tackle a critical challenge: improving the reliability of our advanced compute clusters and the underlying inference, training, and internal production services. In this role, you'll work close to the code and design solutions that will scale with our rapidly growing system production and software service offerings. If you have strong fundamentals in systems, debugging, and failure analysis—and enjoy building tools and solving hard reliability problems—we want to hear from you. New college graduates are welcome. Responsibilities Contribute to the technical roadmap and execution for kernel-centric reliability of our internal and customer-facing systems. Partner with System and Cluster Operations teams to reduce system and service downtime after failure through tooling, analysis, and hands-on debugging support. Work with the Debug Team to enhance debug tools with the goal of speeding up failure analysis. Collaborate with software teams to improve the software stack—including kernels—to improve on-field debugging and failure analysis. Work with ASIC and hardware architecture teams to co-design next-generation architectures with reliability and ease of debug in mind. Participate in incident response, root-cause analysis, and post-mortems; drive follow-ups that measurably improve reliability over time. Skills & Qualifications We recognize great engineers come from different backgrounds. If you're excited about the role, we encourage you to apply even if you don't meet every qualification. Required (or demonstrated through projects/internships/coursework): Strong programming skills in C/C++ and Python. Solid foundations in operating systems, computer architecture, and systems programming fundamentals. Ability to debug complex issues using logs, traces, and standard debugging workflows; interest in root-cause analysis. Preferred Skills & Qualifications Exposure to parallel and distributed programming (message passing, multicore, GPU, embedded, etc.). Experience building or using debug/diagnostic tools (debuggers, core dump handling, tracing, sanitizers, profilers, etc.). Familiarity with debugging distributed and parallel applications (deadlocks, livelocks, race conditions, etc.). Knowledge of computer architecture concepts (instruction pipelining, multithreading, networking, memory systems, etc.). Operations & Monitoring: familiarity with monitoring, incident response, and post-mortem culture. Why Join Cerebras People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras: Build a breakthrough AI platform beyond the constraints of the GPU. Publish and open source their cutting-edge AI research. Work on one of the fastest AI supercomputers in the world. Enjoy job stability with startup vitality. Our simple, non-corporate work culture that respects individual beliefs. Find out more about what it's like to work at Cerebras here ! Apply today and become part of the forefront of groundbreaking advancements in AI! Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them. This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

View more...
Software EngineeringVia Ashby
Verified14 days ago

Software Engineer - Host and Network IO

On-sitefull timeMid-LevelSunnyvale, United States
Apply Now

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. About The Role The Host and Network IO Team develops the full IO path implementation between a distributed system of server nodes, through the cluster, down to the custom RoCE network stack implemented in Cerebras' system, and over the proprietary IOs onto the WSE. As a software developer on the team, you will interface between AI application-level IO teams, cluster architecture teams, and FPGA/ASIC teams to develop solutions that optimize bandwidth and latency while minimizing congestion, pauses, pause spreading, unfairness, etc. Strong skills in socket programming will enable you to deploy robust management operations, while deftness in RDMA Verbs will enable you to optimize CPU resources and shape network traffic to deliver real world impact on AI performance metrics, as well as developing tools for gaining insight and visibility into network behavior. Meticulous analysis and rigour are key tenants of this role, harnessing that together with a deeply-understood mental model of the server, NIC, protocol, switch, and custom hardware behavior will enable you to lead network debug, optimize traffic patterns, and prescribe architectural changes. Responsibilities Develop x86 & ARM software to expose next-generation hardware IO capabilities for AI/HPC application teams Govern a generic IO API with multiple internal users. Develop control and configuration subsystems directly interacting with Cerebras hardware Drive network performance debug of large AI clusters Gather and analyze network statistics and packet traces to root cause and alleviate bottlenecks and sub-optimalities. Develop tools/telemetry for increasing visibility into the network and IO datapath. Optimize cpu/mem utilization leveraging kernel bypass and zero-copy techniques Integrate leading edge networking technologies and protocols Lead cross-functional technical projects spanning multiple teams and integrating diverse software and hardware components to deliver an improved network IO solution. Foster clear and effective communication across teams and stakeholders. Skills & Qualifications Master's/PhD in Computer Science or Electrical Engineering + 1 year industry experience, OR 3+ years industry experience. Experience in large software environments. Embedded systems, HW/SW co-design, and some driver development. Network protocol familiarity (TCP, RoCE) and network debug tools such as Wireshark, or willingness to learn Some network switch environment familiarity or willingness to learn (Arista, Juniper, etc.). Detail-oriented but keen to learn the bigger picture and step out of comfort zone to embrace the unknown. Why Join Cerebras People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras: Build a breakthrough AI platform beyond the constraints of the GPU. Publish and open source their cutting-edge AI research. Work on one of the fastest AI supercomputers in the world. Enjoy job stability with startup vitality. Our simple, non-corporate work culture that respects individual beliefs. Find out more about what it's like to work at Cerebras here ! Apply today and become part of the forefront of groundbreaking advancements in AI! Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them. This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

View more...
Software EngineeringVia Ashby
Verified18 days ago

Software Engineer - Host and Network IO

On-sitefull timeMid-LevelToronto, Canada
Apply Now

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. About The Role The Host and Network IO Team develops the full IO path implementation between a distributed system of server nodes, through the cluster, down to the custom RoCE network stack implemented in Cerebras' system, and over the proprietary IOs onto the WSE. As a software developer on the team, you will interface between AI application-level IO teams, cluster architecture teams, and FPGA/ASIC teams to develop solutions that optimize bandwidth and latency while minimizing congestion, pauses, pause spreading, unfairness, etc. Strong skills in socket programming will enable you to deploy robust management operations, while deftness in RDMA Verbs will enable you to optimize CPU resources and shape network traffic to deliver real world impact on AI performance metrics, as well as developing tools for gaining insight and visibility into network behavior. Meticulous analysis and rigour are key tenants of this role, harnessing that together with a deeply-understood mental model of the server, NIC, protocol, switch, and custom hardware behavior will enable you to lead network debug, optimize traffic patterns, and prescribe architectural changes. Responsibilities Develop x86 & ARM software to expose next-generation hardware IO capabilities for AI/HPC application teams Govern a generic IO API with multiple internal users. Develop control and configuration subsystems directly interacting with Cerebras hardware Drive network performance debug of large AI clusters Gather and analyze network statistics and packet traces to root cause and alleviate bottlenecks and sub-optimalities. Develop tools/telemetry for increasing visibility into the network and IO datapath. Optimize cpu/mem utilization leveraging kernel bypass and zero-copy techniques Integrate leading edge networking technologies and protocols Lead cross-functional technical projects spanning multiple teams and integrating diverse software and hardware components to deliver an improved network IO solution. Foster clear and effective communication across teams and stakeholders. Skills & Qualifications Master's/PhD in Computer Science or Electrical Engineering + 1 year industry experience, OR 3+ years industry experience. Experience in large software environments. Embedded systems, HW/SW co-design, and some driver development. Network protocol familiarity (TCP, RoCE) and network debug tools such as Wireshark, or willingness to learn Some network switch environment familiarity or willingness to learn (Arista, Juniper, etc.). Detail-oriented but keen to learn the bigger picture and step out of comfort zone to embrace the unknown. Why Join Cerebras People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras: Build a breakthrough AI platform beyond the constraints of the GPU. Publish and open source their cutting-edge AI research. Work on one of the fastest AI supercomputers in the world. Enjoy job stability with startup vitality. Our simple, non-corporate work culture that respects individual beliefs. Find out more about what it's like to work at Cerebras here ! Apply today and become part of the forefront of groundbreaking advancements in AI! Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them. This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

View more...
Software EngineeringVia Ashby
Verified18 days ago

Staff GPU Inference SDET

On-sitefull timeLead / StaffSunnyvale, United States
Apply Now

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. About the Role As a Staff GPU Inference SDET, you will be the founding quality, reliability, and validation lead for a new GPU Inference Development team. Working closely with engineering leads and cross-functional systems infrastructure teams, you will design, build, and scale the end-to-end release qualification and automated test ecosystem for our GPU inference stack and rack-scale accelerated compute fleets. In this high-impact role, you will be responsible for building automated test suites to validate multi-node GPU cluster bring-up, verifying prefill worker optimizations, testing open-source and custom serving engines, and ensuring numerical correctness and performance stability under real-world streaming workloads. You will be the primary technical anchor ensuring production-grade reliability, fault isolation, and peak inference performance across accelerated GPU infrastructure. WHAT YOU’LL DO Build GPU Release Qualification Systems : Design and implement automated test automation frameworks, regression gates, and release qualification pipelines for the complete GPU inference stack—spanning custom API services, model-serving workers, container runtimes, serving engines, driver stacks, and firmware. Inference Serving & Workload Validation : Benchmark and stress-test distributed LLM serving frameworks, focusing on prefill vs. decode worker performance, continuous batching, prefix caching, KV-cache efficiency, and tensor/expert parallelism. Performance & Performance Modeling Verification : Build automated workload replay and benchmarking tools to validate GPU performance models. Track critical serving metrics including Time-to-First-Token (TTFT), Inter-Token Latency (ITL), request throughput, tail latency (P99), and capacity efficiency. Numerical Correctness & Quality Gates : Build validation infrastructure to ensure model accuracy, precision stability (FP16/FP8/quantization), determinism, and output correctness across software updates, kernel fusions, and hardware revisions. Fault Injection & Fleet Resilience : Engineer chaos engineering and fault-injection suites to simulate node failures, inter-node network degradation, GPU memory leaks, driver/firmware mismatches, and automated recovery paths for multi-node GPU clusters. Observability & CI/CD Integration : Integrate automated test pipelines with telemetry tools (e.g., Prometheus, Grafana) to turn one-off investigations into repeatable engineering gates and continuous performance monitoring. REQUIREMENTS: 8+ years of software engineering experience as an SDET, Infrastructure Quality Lead, or Systems Test Engineer. GPU & Cluster Infrastructure Expertise : Hands-on experience bringing up, provisioning, and validating multi-node GPU clusters (NVIDIA or AMD ecosystem) across public cloud infrastructure or enterprise data center environments. Inference Stack Knowledge : Deep understanding of LLM serving engines and distributed runtimes, including prefill vs. decode disaggregation, KV-cache management, and dynamic batching. Automation & Scripting : Expert-level Python programming skills with extensive experience designing custom test automation frameworks, diagnostic tooling, and CI/CD integration. Orchestration & Networking : Strong proficiency with container orchestration tools (e.g., Kubernetes, Slurm, Ray) and high-performance cluster interconnects (e.g., InfiniBand, RoCE, NCCL). Failure Analysis & Debugging : Proven background in root-cause analysis across software/hardware boundaries, stress testing, and node failure simulation in distributed systems. NICE TO HAVES: Direct experience with either AMD (ROCm / HIP) or NVIDIA software stacks. Experience building workload replay tools, ML evaluation pipelines, or MLPerf Inference benchmark suites. Familiarity with low-level kernel profiling tools (PyTorch Profiler, NVTX, ROCm profilers) or C++ Why Join Cerebras People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras: Build a breakthrough AI platform beyond the constraints of the GPU. Publish and open source their cutting-edge AI research. Work on one of the fastest AI supercomputers in the world. Enjoy job stability with startup vitality. Our simple, non-corporate work culture that respects individual beliefs. Find out more about what it's like to work at Cerebras here ! Apply today and become part of the forefront of groundbreaking advancements in AI! Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them. This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

View more...
QA & AutomationVia Ashby
Verified19 days ago

Page 1 of 3

PreviousNext