LogoKode$word
Togetherai logo
Verified Tech Organization

Careers at Togetherai

Browse and filter through all verified positions currently open at Togetherai.

Total Company Roles30
Matching Filter30

All Openings (30)

Ordered by most recently published

Senior Software Engineer - AI Compute, Together Cloud

On-sitefull timeSeniorAmsterdam, Netherlands
Apply Now

About the Role Together AI is building the AI Native Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art GPU cloud infrastructure. The Together Cloud team builds the [Together GPU Clusters](https://www.together.ai/gpu-clusters) flagship IaaS product that provides high-performance, AI-ready GPU clusters through a self-serve cloud console, along with the virtualized infrastructure layer powering Together's inference, RL, and fine-tuning products. As a Senior Software Engineer focusing on AI Compute in the Together Cloud org, you will build and own major components of the next generation AI cloud platform – a highly available, global cloud infrastructure with cutting-edge virtualization of the latest ML hardware: GB300s/VRs, BlueField DPUs, InfiniBand and dual/quad-plane RoCEv2 fabrics. That virtualized computing platform powers our own SaaS products – inference, RL, and fine-tuning – and serves external cloud customers through self-serve offerings such as on-demand/reserved Kubernetes/Slurm clusters, across dozens of data centers and hundreds of thousands of GPUs. The hard problems in rapidly scaling heterogeneous GPU fleets are software problems: fully automated bootstrapping of GPU data centers, high-performance virtualization of GPU compute and DC networking without compromising isolation or portability, and fault-tolerant decentralized control planes. We solve them by building global and in-DC services, Kubernetes operators, and high-performance SDN libraries, forking hypervisors and Linux kernels, and building infra tailored to inference and fine-tuning. You'll own massive greenfield projects across their full lifecycle – scoping the problem and writing PRDs with our PMs, designing the system, and building it through to GA launch. Responsibilities Build the GPU and network virtualization stack: the hypervisor, kernel, and SDN work that makes GPU compute and DC networking high-performance, portable, and strongly isolated across heterogeneous hardware. Build and maintain our in-DC IaaS layer: the services, Kubernetes operators, and libraries that provision and manage compute, storage, and networks in our data centers, including VMs, parallel filesystems, VPCs, and InfiniBand partitions. Implement and harden the bring-up path for a new Vera Rubin data center with thousands of GPUs. Scale the distributed GPU scheduling and global management plane: the control-plane services behind on-demand and reserved clusters, including the automation that onboards new capacity and raises per-cluster limits. Harden the monitoring and automated remediation layer for fault tolerance: automated detection, isolation, and recovery of failed nodes that keeps distributed pretraining and large-scale inference fault-tolerant. Own your components end-to-end: write the design docs, break the work into milestones that ship incrementally, and improve the reliability of what's already in production. Raise the bar around you: code review, design feedback, and mentoring junior engineers. Build the tooling other teams rely on: testing frameworks, developer tools, and documentation that make our systems robust and usable across teams, plus contributions to the core, open-source Together AI platform. To be successful, you'll need to be deeply technical, ready to own projects truly end-to-end, and an excellent communicator — strong software development fundamentals, strong systems knowledge and troubleshooting instincts, and the collaboration and diplomacy skills to work across teams. Requirements 5+ years of professional software development experience, with strong proficiency in at least one backend programming language (Golang desired). Demonstrated ownership of large-scale projects driven end-to-end to completion, writing high-performance, well-tested, production-quality code. Demonstrated experience building and operating high-performance and/or globally distributed micro-service architectures across one or more cloud providers (AWS, Azure, GCP). Strong systems knowledge across compute, networking, and storage, including concurrency, memory management, performant I/O, and scale. Excellent communication skills — able to write clear design docs and work effectively with both technical and non-technical team members. Experience building and operating reliable, customer-facing production systems at scale, and owning the infrastructure automation (Terraform, Ansible), observability (Prometheus, Grafana), and CI/CD (GitHub Actions, ArgoCD) that keep them healthy. Preferred Qualifications: Proficiency with Kubernetes internals, such as implementing non-trivial Kubernetes operators, device/storage/network plugins, custom schedulers, or patches to Kubernetes itself Proficiency with VMs/hypervisors, such as QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV Proficiency with DC networking tech + solutions, such as VLAN, VXLAN, VPN, VPC, OVS/OVN Experience with Cluster API or similar Experience working on high-performance compute, networking, and/or storage Experience virtualizing GPUs and/or InfiniBand Experience building IaaS or PaaS systems at scale Experience with DPUs/SmartNICs GPU programming, NCCL, CUDA knowledge About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
Software EngineeringVia Greenhouse
Verified4 days ago

Staff Software Engineer - AI Compute, Together Cloud

On-sitefull timeLead / StaffSan Francisco, United States
Apply Now

About the Role Together AI is building the AI Native Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art GPU cloud infrastructure. The Together Cloud team builds the [Together GPU Clusters](https://www.together.ai/gpu-clusters) flagship IaaS product that provides high-performance, AI-ready GPU clusters through a self-serve cloud console, along with the virtualized infrastructure layer powering Together's inference, RL, and fine-tuning products. As a Staff Software Engineer focusing on AI Compute in the Together Cloud org, you will set technical direction for and build major components of the next generation AI cloud platform – a highly available, global cloud infrastructure with cutting-edge virtualization of the latest ML hardware: GB300s/VRs, BlueField DPUs, InfiniBand and dual/quad-plane RoCEv2 fabrics. That virtualized computing platform powers our own SaaS products – inference, RL, and fine-tuning – and serves external cloud customers through self-serve offerings such as on-demand/reserved Kubernetes/Slurm clusters, across dozens of data centers and hundreds of thousands of GPUs. This is an architect-and-build role. Fully automated bootstrapping of GPU data centers, high-performance virtualization of GPU compute and DC networking without compromising isolation or portability, and fault-tolerant decentralized control planes — you'll set the architecture for these across our global and in-DC services, and be a key owner of the hardest parts, in the code as well as the design. Your designs will span the IaaS layer of a greenfield Vera Rubin data center up to the global management plane that schedules capacity across all of them. At this level the job is as much leverage as code: the standards you set and the engineers you grow decide how fast the rest of Together Cloud ships. Responsibilities Own the GPU and network virtualization stack: the hypervisor, kernel, and SDN work that keeps GPU compute and DC networking high-performance, portable, and strongly isolated across heterogeneous hardware. Own the in-DC IaaS layer: architect and roadmap the services, Kubernetes operators, and libraries that provision and manage compute, storage, and networks in our data centers — VMs, parallel filesystems, VPCs, and InfiniBand partitions; lead its build-out for a new Vera Rubin data center with thousands of GPUs, from hardware bring-up to customer-facing API. Design the GPU scheduling and global management plane: the distributed control plane behind on-demand and reserved clusters across dozens of data centers, including the systems that scale per-cluster limits and automate the onboarding of new capacity. Architect monitoring and automated remediation for fault tolerance: the strategy for automated detection, isolation, and recovery of failed nodes that keeps distributed pretraining and large-scale inference running through hardware failures. Set technical direction across teams: lead design reviews, resolve cross-cutting architectural disagreements, unblock cross-team dependencies and integration risks, and define the standards other engineers build against — measured in cluster reliability, time-to-first-GPU on new capacity, and quality at scale. Grow the team: mentor senior and junior engineers, deepen the team's expertise in virtualization, DC networking, and GPU infrastructure, and help raise the hiring bar for Together Cloud. Set the engineering bar: create the testing frameworks, tools, and developer documentation that make our systems robust and usable by other teams, and shape the core, open-source Together AI platform. To be successful you'll need to be deeply technical and an excellent communicator — expert software development fundamentals, deep systems knowledge and troubleshooting instincts, and the leadership and diplomacy skills to align teams that don't report to you. Much of this work starts ambiguous, and we expect you to define the scope yourself and drive it to production. Requirements 7+ years of professional software development experience, with expert-level proficiency in at least one backend language (Golang desired), writing high-performance, well-tested, production-quality code. Track record of owning the architecture of large distributed systems from blank page to production at scale, including the judgment calls that could not be reversed cheaply. Deep experience building and operating globally distributed, high-performance microservice architectures across one or more cloud providers (AWS, Azure, GCP). Expert systems knowledge across compute, networking, and storage — including concurrency, memory management, performant I/O, and scale at a global level. Demonstrated technical leadership beyond your own commits: mentoring senior engineers, leading design reviews, and driving alignment across teams that do not report to you. Excellent communication and diplomacy skills — able to write design docs that settle arguments, and to work effectively with technical and non-technical stakeholders. Experience building and operating reliable, customer-facing production systems at scale, and owning the infrastructure automation (Terraform, Ansible), observability (Prometheus, Grafana), and CI/CD (GitHub Actions, ArgoCD) that keep them healthy. Preferred Qualifications (not must haves) Deep Kubernetes internals experience, such as implementing non-trivial Kubernetes operators, device/storage/network plugins, custom schedulers, or patches to Kubernetes itself Deep experience with VMs/hypervisors, such as QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV Deep experience with DC networking tech + solutions, such as VLAN, VXLAN, VPN, VPC, OVS/OVN Experience with Cluster API or similar Experience working on high-performance compute, networking, and/or storage Experience virtualizing GPUs and/or InfiniBand Experience building IaaS or PaaS systems at scale Experience with DPUs/SmartNICs GPU programming, NCCL, CUDA knowledge About Together AI Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers in our journey in building the next generation AI infrastructure. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $260,000 - $300,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
Software EngineeringVia Greenhouse
Verified5 days ago

Senior Software Engineer - AI Compute, Together Cloud

On-sitefull timeSeniorSan Francisco, United States
Apply Now

About the Role Together AI is building the AI Native Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art GPU cloud infrastructure. The Together Cloud team builds the [Together GPU Clusters](https://www.together.ai/gpu-clusters) flagship IaaS product that provides high-performance, AI-ready GPU clusters through a self-serve cloud console, along with the virtualized infrastructure layer powering Together's inference, RL, and fine-tuning products. As a Senior Software Engineer focusing on AI Compute in the Together Cloud org, you will build and own major components of the next generation AI cloud platform – a highly available, global cloud infrastructure with cutting-edge virtualization of the latest ML hardware: GB300s/VRs, BlueField DPUs, InfiniBand and dual/quad-plane RoCEv2 fabrics. That virtualized computing platform powers our own SaaS products – inference, RL, and fine-tuning – and serves external cloud customers through self-serve offerings such as on-demand/reserved Kubernetes/Slurm clusters, across dozens of data centers and hundreds of thousands of GPUs. The hard problems in rapidly scaling heterogeneous GPU fleets are software problems: fully automated bootstrapping of GPU data centers, high-performance virtualization of GPU compute and DC networking without compromising isolation or portability, and fault-tolerant decentralized control planes. We solve them by building global and in-DC services, Kubernetes operators, and high-performance SDN libraries, forking hypervisors and Linux kernels, and building infra tailored to inference and fine-tuning. You'll own massive greenfield projects across their full lifecycle – scoping the problem and writing PRDs with our PMs, designing the system, and building it through to GA launch. Responsibilities Build the GPU and network virtualization stack: the hypervisor, kernel, and SDN work that makes GPU compute and DC networking high-performance, portable, and strongly isolated across heterogeneous hardware. Build and maintain our in-DC IaaS layer: the services, Kubernetes operators, and libraries that provision and manage compute, storage, and networks in our data centers, including VMs, parallel filesystems, VPCs, and InfiniBand partitions. Implement and harden the bring-up path for a new Vera Rubin data center with thousands of GPUs. Scale the distributed GPU scheduling and global management plane: the control-plane services behind on-demand and reserved clusters, including the automation that onboards new capacity and raises per-cluster limits. Harden the monitoring and automated remediation layer for fault tolerance: automated detection, isolation, and recovery of failed nodes that keeps distributed pretraining and large-scale inference fault-tolerant. Own your components end-to-end: write the design docs, break the work into milestones that ship incrementally, and improve the reliability of what's already in production. Raise the bar around you: code review, design feedback, and mentoring junior engineers. Build the tooling other teams rely on: testing frameworks, developer tools, and documentation that make our systems robust and usable across teams, plus contributions to the core, open-source Together AI platform. To be successful, you'll need to be deeply technical, ready to own projects truly end-to-end, and an excellent communicator — strong software development fundamentals, strong systems knowledge and troubleshooting instincts, and the collaboration and diplomacy skills to work across teams. Requirements 5+ years of professional software development experience, with strong proficiency in at least one backend programming language (Golang desired). Demonstrated ownership of large-scale projects driven end-to-end to completion, writing high-performance, well-tested, production-quality code. Demonstrated experience building and operating high-performance and/or globally distributed micro-service architectures across one or more cloud providers (AWS, Azure, GCP). Strong systems knowledge across compute, networking, and storage, including concurrency, memory management, performant I/O, and scale. Excellent communication skills — able to write clear design docs and work effectively with both technical and non-technical team members. Experience building and operating reliable, customer-facing production systems at scale, and owning the infrastructure automation (Terraform, Ansible), observability (Prometheus, Grafana), and CI/CD (GitHub Actions, ArgoCD) that keep them healthy. Preferred Qualifications: Proficiency with Kubernetes internals, such as implementing non-trivial Kubernetes operators, device/storage/network plugins, custom schedulers, or patches to Kubernetes itself Proficiency with VMs/hypervisors, such as QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV Proficiency with DC networking tech + solutions, such as VLAN, VXLAN, VPN, VPC, OVS/OVN Experience with Cluster API or similar Experience working on high-performance compute, networking, and/or storage Experience virtualizing GPUs and/or InfiniBand Experience building IaaS or PaaS systems at scale Experience with DPUs/SmartNICs GPU programming, NCCL, CUDA knowledge About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $220,000 - $270,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
Software EngineeringVia Greenhouse
Verified5 days ago

GTM Data Analytics Engineer

On-sitefull timeMid-LevelSan Francisco, United States
Apply Now

About the Role We are seeking a highly analytical and detail-oriented Data Analytics Engineer to join our data team. This role will report into the Revenue Operations team. The ideal candidate will help build and maintain the data infrastructure and reporting used across go-to-market, finance, and product teams, working closely with analytics colleagues across both the GTM and core data teams to deliver reliable, well-documented data products. Responsibilities Develop and maintain dashboards and reports primarily supporting GTM stakeholders, with additional use by Finance, Product, and other cross-functional teams; assist with deep-dive analyses on business performance and trends. Pull and analyze data from source systems (Salesforce, Amplitude, production systems, billing) to support reporting and ad hoc requests. Use SQL to extract, clean, and analyze data, with attention to correctness and efficiency. Build foundational SQL and business knowledge and quickly grow into contributing to dbt models within Snowflake, following established dimensional modeling standards. Partner with GTM, Finance, and Product stakeholders to understand data needs and keep reporting and definitions consistent. Requirements Bachelor's degree in Business Analytics, Data Science, Statistics, Computer Science, or a related quantitative field. 1-3 years of experience in a Data Analyst, BI, or Analytics Engineering role. Strong SQL skills required, with the ability to write complex, multi-step queries involving joins and aggregations; experience with dbt and/or Python is a strong plus. Familiarity with a cloud data warehouse (Snowflake preferred) and BI tools (Hex, Metabase). Strong attention to detail and ability to manage multiple priorities in a fast-paced environment. Excellent communication skills, with the ability to explain data findings to non-technical stakeholders. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $ 120K - $150K + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy .

View more...
Data Engineering & BIVia Greenhouse
Verified6 days ago

About the Role We're looking for a Software Engineer to build the systems that treat infrastructure as software. This role owns the software state machines that provision hardware, bring it into service, and manage its full lifecycle — turning racks of GPUs into running inference clusters without a human touching a runbook. The Research and Inference team is your customer: today they file tickets and wait; the target state is that they issue a single API call to stand up, scale, or tear down a cluster, and the system takes care of the rest. The platform is manifest-driven such that teams declare the desired state of a cluster or host — shape, topology, software stack — and the system is responsible for reconciling reality to that manifest, continuously, through every stage of its lifecycle. You will design the engines that manifest the schema, the engines that execute against it, and the workflows that carry a piece of hardware or a cluster from one state to the next—taking it from bare metal to a fully functioning AI cluster for training or inference. You'll write production code which is typed, tested, versioned, and deployed through CI/CD that models infrastructure state and reconciles it, the same way a Kubernetes controller reconciles a cluster's desired state. Success looks like eliminating manual provisioning work, not documenting it better. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine: design and implement the software that models the full lifecycle of a physical host from discovery, inference bring-up to GPU driver/CUDA stack, health validation, and decommission/RMA — as explicit, versioned states and transitions. Build the self-service API: design declarative APIs and a control plane so the inference team can request, scale, and tear down inference clusters with one API call — no ticket, no human in the loop. Automate self-healing: detect degraded or failed nodes, drain them safely, trigger repair or replacement, and reintroduce healthy capacity into the pool automatically. Own reliability of the pipeline: idempotency, retries, rollback, and drift detection so the provisioning system is as dependable as any other production service. Partner with the inference/ML platform team: understand the cluster shapes they need — topology, interconnect, scheduling constraints — and encode them as first-class abstractions in the platform. Engineer it like software: strong typing, automated tests, code review, versioning, and CI/CD for infrastructure code — this is a product, not a collection of Ansible playbooks. Requirements Core requirements (all levels): Strong software engineering background in Go, Python, Rust, or similar — you write and test real software for a living. Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent to run long-lived, manifest-driven workflows that survive failures and resume mid-execution. Experience building software control planes or orchestration systems that model state and reconcile it over time (e.g., Kubernetes controllers/operators, custom reconciliation loops, workflow engines). Experience with event-driven systems — designing and building software around message queues, event streams, or pub/sub (e.g., Kafka, NATS, SQS) rather than polling or cron-driven scripts. A product mindset. You’ve built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. Nice to have: Exposure to bare-metal provisioning (PXE/iPXE, Redfish/IPMI, BMC) and/or networking fundamentals (VLANs, BGP, fabric design), or GPU/accelerator infrastructure. Experience with GPU cluster software stacks (NCCL, CUDA, InfiniBand/RoCE). Prior work at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization. Systems programming in Rust or Go. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy .

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Machine Learning Engineer - Inference

On-sitefull timeMid-LevelSan Francisco, United States
Apply Now

About the Role Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role involves working with state-of-the-art large language models models and ensuring they run efficiently and effectively at scale. If you are passionate about AI inference, PyTorch, and developing high-performance systems, we want to hear from you. This position offers the chance to collaborate closely with AI researchers and engineers to create cutting-edge AI solutions. Join us in shaping the future at Together AI! Responsibilities Design and build the production systems that power the Together AI inference engine, enabling reliability and performance at scale. Develop and optimize runtime inference services for large-scale AI applications. Collaborate with researchers, engineers, product managers, and designers to bring new features and research capabilities to the world. Conduct design and code reviews to ensure high standards of quality. Create services, tools, and developer documentation to support the inference engine. Implement robust and fault-tolerant systems for data ingestion and processing. Requirements 3+ years of experience writing high-performance, well-tested, production-quality code. Proficiency with Python and PyTorch. Demonstrated experience in building high performance libraries and tooling. Excellent understanding of low-level operating systems concepts including multi-threading, memory management, networking, storage, performance, and scale. Preferred: Knowledge of existing AI inference systems such as TGI, vLLM, TensorRT-LLM, Optimum Preferred: Knowledge of AI inference techniques such as speculative decoding. Preferred: Knowledge of CUDA/Triton programming. Nice to have: Knowledge of Rust, Cython and compilers. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other competitive benefits. The US base salary range for this full-time position is $200,000 - $300,000 + equity + benefits. Our salary ranges are determined by location, level, and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunities to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
AI / ML & Data ScienceVia Greenhouse
Verified6 days ago

About the Role We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine: design and implement the software that models the full lifecycle of a physical host from discovery, inference bring-up to GPU driver/CUDA stack, health validation, and decommission/RMA — as explicit, versioned states and transitions. Build the self-service API: design declarative APIs and a control plane so the inference team can request, scale, and tear down inference clusters with one API call — no ticket, no human in the loop. Automate self-healing: detect degraded or failed nodes, drain them safely, trigger repair or replacement, and reintroduce healthy capacity into the pool automatically. Own reliability of the pipeline: idempotency, retries, rollback, and drift detection so the provisioning system is as dependable as any other production service. Partner with the inference/ML platform team: understand the cluster shapes they need — topology, interconnect, scheduling constraints — and encode them as first-class abstractions in the platform. Engineer it like software: strong typing, automated tests, code review, versioning, and CI/CD for infrastructure code — this is a product, not a collection of Ansible playbooks. Requirements Core requirements (all levels): Strong software engineering background in Go, Python, Rust, or similar — you write and test real software for a living. Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent to run long-lived, manifest-driven workflows that survive failures and resume mid-execution. Experience building software control planes or orchestration systems that model state and reconcile it over time (e.g., Kubernetes controllers/operators, custom reconciliation loops, workflow engines). Experience with event-driven systems — designing and building software around message queues, event streams, or pub/sub (e.g., Kafka, NATS, SQS) rather than polling or cron-driven scripts. A product mindset. You’ve built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. Nice to have: Exposure to bare-metal provisioning (PXE/iPXE, Redfish/IPMI, BMC) and/or networking fundamentals (VLANs, BGP, fabric design), or GPU/accelerator infrastructure. Experience with GPU cluster software stacks (NCCL, CUDA, InfiniBand/RoCE). Prior work at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization. Systems programming in Rust or Go. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $160,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy .

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Senior Backend Engineer, Inference Platform

On-sitefull timeSeniorSan Francisco, United States
Apply Now

About the Role Together AI is building the Inference Platform that brings the most advanced generative AI models to the world. Our platform powers multi-tenant serverless workloads and dedicated endpoints, enabling developers, enterprises, and researchers to harness the latest LLMs, multimodal models, image, audio, video, and speech models at scale. If you get a thrill from optimizing latency down to the last millisecond, this is your playground. You’ll work hands-on with tens of thousands of GPUs (H100s, H200s, GB200s, and beyond), figuring out how to fully utilize every FLOP and every gigabyte of memory. You’ll collaborate directly with research teams to bring frontier models into production, making breakthroughs usable in the real world. Our team also works closely with the open source community, contributing to and leveraging projects like SGLang, vLLM, and NVIDIA Dynamo to push the boundaries of inference performance and efficiency. Shape the core inference backbone that powers Together AI’s frontier models. Solve performance-critical challenges in global request routing, load balancing, and large-scale resource allocation. Work with state-of-the-art accelerators (H100s, H200s, GB200s) at global scale. Partner with world-class researchers to bring new model architectures into production. Collaborate with and contribute to the open source community, shaping the tools that advance the industry. A culture of deep technical ownership and high impact — where your work makes models faster, cheaper, and more accessible. Competitive compensation, equity, and benefits. Responsibilities Build and optimize global and local request routing, ensuring low-latency load balancing across data centers and model engine pods. Develop auto-scaling systems to dynamically allocate resources and meet strict SLOs across dozens of data centers. Design systems for multi-tenant traffic shaping, tuning both resource allocation and request handling — including smart rate limiting and regulation — to ensure fairness and consistent experience across all users. Engineer trade-offs between latency and throughput to serve diverse workloads efficiently. Optimize prefix caching to reduce model compute and speed up responses. Collaborate with ML researchers to bring new model architectures into production at scale. Continuously profile and analyze system-level performance to identify bottlenecks and implement optimizations. Requirements 5+ years of demonstrated experience building large-scale, fault-tolerant, distributed systems and API microservices. Strong background in designing, analyzing, and improving efficiency, scalability, and stability of complex systems. Excellent understanding of low-level OS concepts: multi-threading, memory management, networking, and storage performance. Expert-level programming in one or more of: Rust, Go, Python, or TypeScript. Knowledge of modern LLMs and generative models and how they are served in production is a plus. Experience working with the open source ecosystem around inference is highly valuable; familiarity with SGLang, vLLM, or NVIDIA Dynamo will be especially handy. Experience with Kubernetes or container orchestration is a strong plus. Familiarity with GPU software stacks (CUDA, Triton, NCCL) and HPC technologies (InfiniBand, NVLink, MPI) is a plus. Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $200,000 - $290,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Senior Developer Productivity Engineer

On-sitefull timeSeniorSan Francisco, United States
Apply Now

About the Role At Together AI, you'll own the systems and tooling that let our engineers ship quickly and reliably: CI/CD pipelines, release infrastructure, local dev environments, and the testing systems the whole engineering team depends on. This is a high-leverage, build-heavy role. You'll take on big projects and drive them end to end: building a full-stack release service with deployment orchestration, rebuilding the front-end build pipeline, standing up testing infrastructure that kills flaky builds. The goal is straightforward: engineers should spend their time building products, not fighting their tooling. This is an engineering role focused on building systems, not a QA or manual testing role. You build the infrastructure that makes manual effort unnecessary. Responsibilities Build and own the production systems engineers rely on: release services with deployment orchestration (canary, blue/green), reusable CI/CD pipelines, and self-service tooling. Build testing infrastructure: the frameworks and harnesses that make tests fast and reliable, and track down flaky tests and slow builds at the systems level. Build and maintain local dev environments and containerized workflows so engineers can spin up and iterate fast. Create starter templates and shared tooling (CLIs, codegen, IDE integrations) so teams can stand up new services in minutes. Be the go-to engineer for GitHub Actions and GitOps: defining problems, untangling pipeline issues, and improving how the org works with them. Improve front-end build tooling and developer workflows alongside product engineers. Work across the org to find the biggest sources of friction and fix them at the root. Document what you build so others can use it. Requirements 5+ years building software in production, with a strong track record of owning systems or services end to end, from design through operation. Strong software engineering fundamentals. You write production-quality code and have built and operated real services, not just scripts or test automation. Proficiency in Python, Go, and JavaScript/TypeScript for tooling, automation, and enablement. Deep experience with CI/CD systems (GitHub Actions, ArgoCD, GitOps) and building scalable, reusable pipelines. Experience with containerized development workflows and local dev tooling (e.g., Skaffold). Experience building starter templates and scaffolding, with engineering teams, to accelerate new service creation. Hands-on experience with AI coding tools (e.g., Claude Code, Cursor, Copilot) and a habit of using them to move faster and ship more. A builder's mindset and strong ownership. You like creating the tools and systems other people depend on. Rigor in diagnosing systemic issues: flaky tests, build latency, deployment failures. Nice to have Kubernetes expertise (EKS, K3s) and experience optimizing containerized builds. Infrastructure as Code (Terraform, Ansible, Pulumi). Front-end tooling familiarity (React, Next.js, Jest) to optimize front-end dev workflows. Monitoring and observability (Prometheus, Grafana, Honeycomb) to debug bottlenecks. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $180,000 - $250,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Platform Engineer, Model Shaping

On-sitefull timeMid-LevelSan Francisco, United States
Apply Now

About the Role The Model Shaping team at Together AI works on products and research for tailoring open foundation models to downstream applications. We build services that allow machine learning developers to choose the best models for their tasks and further improve these models using domain-specific data. In addition to that, we develop new methods for more efficient model training and evaluation, drawing inspiration from a broad spectrum of ideas across machine learning, natural language processing, and ML systems. As a Platform Engineer in Model Shaping, you will work at the intersection of backend engineering and infrastructure, building the foundational layers of Together’s platform for model customization and evaluation. You will design, develop, and operate both the backend services and the underlying systems that enable us to sustainably and reliably scale production workflows launched by our users, as well as internal research experiments. You will operate in a cross-functional environment, collaborating with other engineers and researchers in the team to improve the infrastructure based on the needs of projects they work on. You will also interact with other engineering teams at Together (such as Commerce, Data Engineering, and Cloud Infrastructure) to integrate the services developed by Model Shaping with systems developed by those teams. Responsibilities Design and build Together’s systems and infrastructure for model customization, including user-facing features and internal improvements Contribute to reliability improvements for the platform, participating in an on-call rotation and improving processes for incident response Create and improve internal tooling for deployment, continuous integration, and observability Build a job orchestration platform spanning multiple datacenters, supporting a highly heterogeneous hardware landscape Partner with teams developing internal services, co-designing these services and incorporating them in systems built within Together Requirements 3+ years of experience in building infrastructure or backend components of production services Extensive experience designing, operating, and troubleshooting production Linux environments and Kubernetes-based platforms Strong software engineering background in Python or Go Experienced with infrastructure automation tools (Terraform, Ansible), monitoring/observability stacks (Prometheus, Grafana), and CI/CD pipelines (GitHub Actions, ArgoCD) Cloud environment (e.g., AWS/GCP/Azure) administration experience, preferably with a hybrid bare-metal/cloud environment Strong communication skills, be willing to document systems and processes and collaborate with peers of varying technical expertise Comfortable operating across the stack, from cluster operations and infrastructure automation to backend service development Experience in any of the following will make you stand out: Developing large-scale production systems with high reliability requirements Pipeline orchestration frameworks (e.g., Kubeflow, Argo Workflows, Flyte) Managing GPU workloads on HPC clusters, ideally with hands-on experience in operating NVIDIA’s networking stack (e.g., NCCL, Mellanox firmware, GPUDirect RDMA) Deployment of services for AI training or inference Networking fundamentals, including TCP/IP, DNS, routing, load balancing, TLS, and network debugging tools Maintaining or contributing to open-source projects About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is $200,000 - $290,000. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
Cloud, DevOps & SREVia Greenhouse
Verified6 days ago

Page 1 of 3

PreviousNext