LogoKode$word
Togetherai logo
Verified Tech Organization

Careers at Togetherai

Browse and filter through all verified positions currently open at Togetherai.

Total Company Roles30
Matching Filter30

All Openings (30)

Ordered by most recently published

Software Development In Test Intern

On-siteinternshipInternshipSan Francisco, United States
Apply Now

Role Overview As a Software Development in Test (SDET) Intern, you’ll have the opportunity to be a key player in setting a high quality bar for our users. You’ll work on designing and implementing automated testing processes while getting exposure to a key function at Together AI. Across teams like Cluster Management and Inference Platform, our work centers on automating and testing the critical flows behind our infrastructure. This is a rare opportunity to gain deep insight into how AI infrastructure is provisioned, managed, and scaled or to get hands-on experience benchmarking the newest cutting edge open source models. This role is on-site at our HQ in San Francisco, CA. Responsibilities Developing automated test scripts for functionality, performance, and reliability testing across the website and services Write clean, efficient, and well-documented code with a focus on long-term maintainability Extend and improve test automation frameworks to increase efficiency and overall coverage across the platform. Collaborate with engineering and product teams to understand project requirements and contribute to defining test plans and quality standards Minimum Qualifications Actively pursuing a degree in Computer Science, Software Engineering, or a related field, earned or expected by Summer 2028 Excellent programming skills in Typescript, Go or Python Knowledge of automation testing methodologies, tools, and best practices Ability to solve problems creatively and communicate trade-offs effectively Preferred Qualifications Prior SDET experience through internships, hackathons, or projects Experience in API Testing, AI infrastructure, and/or Git workflows and CI automation. Experience with Playwright or Cypress About Together AI Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancements such as FlashAttention, Mamba, FlexGen, Petals, Mixture of Agents, and RedPajama. Internship Program Details Our internship program runs 12 to 14 weeks, giving you the opportunity to work alongside industry-leading engineers and researchers across multiple teams. This cohort's internship dates span either May 17th to August 6th or June 14th to September 3rd. Compensation We offer competitive compensation, housing stipends, and other competitive benefits. The estimated US hourly rate for this role is $58 an hour. Our hourly rates are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Software Engineer Intern

On-siteinternshipInternshipSan Francisco, United States
Apply Now

Role Overview As a Software Engineer Intern, you’ll have the opportunity to work on a variety of projects, from designing scalable systems, building product features, to optimizing performance-critical code. You’ll collaborate with cross-functional teams to build robust, user-focused solutions and contribute to our mission of delivering cutting-edge technology. This role is on-site at our HQ in San Francisco, CA and you’ll be placed into one of our engineering teams spanning Platform Engineering, Infrastructure, Inference, and more! Responsibilities Design, develop, and maintain high-quality software across our tech stack from low-level system to customer-facing UIs Collaborate with product managers, designers, and engineers to deliver features and improvements Write clean, efficient, and well-documented code with a focus on maintainability Participate in code reviews, debugging, and performance optimization Contribute ideas to shape the direction of our products and technical infrastructure Qualifications Actively pursuing a degree in Computer Science, Software Engineering, or a related field, earned or expected by Summer 2028 Excellent programming skills Experience with version control systems (e.g., Git) and collaborative development workflows Ability to solve problems creatively and communicate trade-offs effectively Bonus: Experience with web development, databases, distributed systems, or cloud platforms, OSS contributions. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Internship Program Details: Our internship program spans over 12 weeks where you’ll have the opportunity to work with industry-leading engineers building a cloud from the ground up and possibly contribute to influential open source projects. Our internship dates are May 17th to August 6th or June 14th to September 3rd. Compensation We offer competitive compensation, housing stipends, and other competitive benefits. The estimated US hourly rate for this role is $58/hr. Our hourly rates are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Software Engineer Intern

On-siteinternshipInternshipSan Francisco, United States
Apply Now

Role Overview As a Software Engineer Intern, you’ll have the opportunity to work on a variety of projects, from designing scalable systems, building product features, to optimizing performance-critical code. We’re looking for engineers looking to build meaningful projects end-to-end from initial scope into production. You’ll collaborate with cross-functional teams to build robust, user-focused solutions and contribute to our mission of delivering cutting-edge technology. This role is on-site at our HQ in San Francisco, CA and you must be available for our Winter Internship Dates between January to April. You’ll also be placed into one of our engineering teams spanning Platform Engineering, Infrastructure, SREs, and Inference. Responsibilities Design, develop, and maintain high-quality software across our tech stack from low-level system to customer-facing UIs Collaborate with product managers, designers, and engineers to deliver features and improvements Write clean, efficient, and well-documented code with a focus on maintainability Participate in code reviews, debugging, and performance optimization Contribute ideas to shape the direction of our products and technical infrastructure Qualifications Actively pursuing a degree in Computer Science, Software Engineering, or a related field, earned or expected by Summer 2028 Excellent programming skills with experience in Python, Go, or Typescript Ability to solve problems creatively and communicate trade-offs effectively Prior experiences through internships, hackathons, or projects Bonus: Experience with web development, databases, distributed systems, or cloud platforms, OSS contributions. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Internship Program Details: Our internship program spans over 12 to 14 weeks where you’ll have the opportunity to work with industry-leading engineers building a cloud from the ground up and possibly contribute to influential open source projects. Our internship dates span between January 4th to April 9th. Compensation We offer competitive compensation, housing stipends, and other competitive benefits. The estimated US hourly rate for this role is $58 an hr. Our hourly rates are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Software Engineer, New Grad

On-sitefull timeEntry / JuniorSan Francisco, United States
Apply Now

About the Role As an early career Software Engineer, you’ll work on a variety of projects, from designing scalable systems, building product features, to optimizing performance-critical code. You’ll collaborate with cross-functional teams to build robust, user-focused solutions and contribute to our mission of delivering cutting-edge technology. This role is ideal for recent graduates with a strong foundation in software development and a passion for learning. This role is on-site at our HQ in San Francisco, CA and you’ll be placed into one of our engineering teams spanning Machine Learning, Platform Engineering, Infrastructure, and Inference. Responsibilities Design, develop, and maintain high-quality software across our tech stack from low-level system to customer-facing UIs. Collaborate with product managers, designers, and engineers to deliver features and improvements. Take high autonomy over your work - own projects end to end Participate in code reviews, debugging, and performance optimization. Contribute ideas to shape the direction of our products and technical infrastructure. Work on real, meaningful projects and learn from experienced professionals along the way. Requirements Actively pursuing a degree in Computer Science, Software Engineering, or a related field, earned or expected by Summer 2027 Solid CS fundamentals (data structures, algorithms, systems thinking) Past hands-on experience through internships, personal projects, or previous roles Experience with version control systems (e.g., Git) and collaborative development workflows. Ability to solve problems creatively and communicate trade-offs effectively. Enthusiasm for learning and thriving in a fast-paced, collaborative environment. Bonus: Experience with web development, databases, distributed systems, or cloud platforms, OSS contributions. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $150,000 - $160,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Staff Machine Learning Engineer, Voice AI

On-sitefull timeLead / StaffSan Francisco, United States
Apply Now

About the Role Together AI is building the best inference infrastructure for voice applications. Our Voice AI platform powers production-grade, real-time voice agents and applications — serving speech-to-text and text-to-speech models with best-in-class latency and reliability. We're looking for a Staff ML Engineer to drive the model serving layer for voice workloads. You'll work hands-on with inference engines like TRT-LLM and SGLang to optimize how we serve models like Whisper, Parakeet, Orpheus, and Kokoro — pushing latency and throughput to the frontier. You'll profile GPU utilization, design batching strategies for streaming audio, and ensure new model architectures can go from research to production quickly. This is a foundational hire on a small, high-impact team. Voice inference has unique challenges — streaming audio, tokenization, real-time latency budgets — that require dedicated ML engineering focus. You'll shape how Together serves voice models as the industry moves from pipeline architectures (ASR → LLM → TTS) toward end-to-end speech-to-speech. Own the model serving stack that powers Together's voice platform across STT, TTS, and speech-to-speech. Work directly with state-of-the-art accelerators (H100s, H200s, B200s) to optimize voice model inference. Collaborate with model partners (Cartesia, Deepgram, Rime, and others) to bring their models to production on Together's infrastructure. Build quality evaluation frameworks that guide model selection for customers and inform the roadmap. Join a small, early-stage team with outsized impact on a fast-growing product area. Responsibilities Own the voice inference roadmap end-to-end — define and execute the technical strategy for optimizing STT, TTS, and speech-to-speech models across Together's infrastructure, with a clear-eyed view of where the field is heading and how to position the platform ahead of it. Drive best-in-class inference performance — architect and implement systems targeting leading TTFB, throughput, and GPU utilization for voice workloads; set the performance bar others in the industry measure against, not just catch up to. Lead productionization of voice models at scale — design the serving architecture for serverless and dedicated endpoints, including batching strategies, streaming inference pipelines, and memory management tailored to real-time audio; own reliability and latency SLAs. Build the voice evaluation platform — design a rigorous, extensible evaluation framework covering WER across accents, languages, and noise conditions for STT; naturalness, latency, and pronunciation fidelity for TTS; establish the internal benchmark methodology that informs model selection and roadmap decisions. Shape the architecture for next-generation model support — anticipate and enable emerging model paradigms — audio-native LLMs, codec-based architectures (SNAC, Encodec), and end-to-end speech-to-speech systems — before they're mainstream, not after. Serve as the technical DRI for model partner integrations — lead deep collaboration with partners such as Cartesia, Deepgram, and Rime; own the full lifecycle from integration to optimization to ongoing performance accountability. Diagnose and resolve the hardest performance problems in the stack — conduct systematic profiling and root-cause analysis from GPU kernel behavior to framework-level bottlenecks; drive shipped improvements with documented, measurable impact. Influence platform architecture across the organization — partner with platform engineering leadership to ensure the serving layer is built for the latency and reliability demands of real-time voice APIs; your technical decisions should raise the ceiling for the whole team. Define and scale voice fine-tuning capabilities — lead the technical direction for enabling customers to fine-tune STT and TTS models on Together's infrastructure, establishing the primitives for differentiated voice experiences. Lay technical foundations for a category-defining product surface — architect systems with enough foresight that they support multiple new voice products with minimal rework; think in terms of platforms, not point solutions. Requirements 8+ years of ML engineering experience, with a demonstrated focus on model serving, inference optimization, or ML infrastructure at production scale — including systems you've owned from design through live traffic. Deep, practical expertise in LLM serving engines (vLLM, SGLang, TensorRT-LLM, or equivalent) — you've modified engine internals, debugged edge cases under load, and contributed improvements back; you don't stop at the API surface. Expert-level Python and PyTorch proficiency, with a strong command of GPU optimization — CUDA kernels, memory hierarchies, profiling toolchains — and a track record of turning that knowledge into shipped latency or throughput wins. Proven system design judgment — you've made architectural decisions that held up at scale and influenced how a team or platform evolved; you can articulate the tradeoffs you made and why. Strong technical leadership — you operate with high autonomy, define the right problems before solving them, and raise the bar for engineering quality around you without requiring process overhead. Sharp product intuition for developer tooling — you understand what voice application developers actually need to ship great products, and you let that shape your technical priorities, not just the other way around. Proven ability to move fast in ambiguous environments — you've thrived on early-stage or platform teams where scope is wide, ownership is deep, and the roadmap you build is the one you execute. Strong foundation in speech and audio ML (ASR/TTS architectures, audio signal processing) — directly relevant experience is strongly preferred; exceptional ML engineering fundamentals with genuine curiosity about the domain is also considered. Familiarity with audio codec and tokenization schemes (SNAC, Encodec, DAC) is a meaningful plus at this level. Experience training or fine-tuning speech models at scale is a significant advantage. Bachelor's or Master's in Computer Science, Electrical Engineering, or related field — or equivalent depth demonstrated through your work. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $220,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
AI / ML & Data ScienceVia Greenhouse
Verified6 days ago

Staff Platform Engineer, Service Infrastructure

On-sitefull timeLead / StaffSan Francisco, United States
Apply Now

About the Role Together AI is hiring a Staff Platform Engineer to join the Product Foundations engineering organization and drive its service infrastructure strategy. Product Foundations builds and operates Together’s mission-critical product platforms that support all cloud products, including API Platform (non-Inference), web UI Platform, Billing, and customer-facing IAM. These services sit on the critical path for customers and internal systems. This is a hands-on Staff role focused on evolving Product Foundations’ core infrastructure strategy from the inside: understanding service team needs, turning repeated infrastructure problems into reusable patterns, and coordinating across platform owners so Product Foundations services are reliable, repeatable, and built on the right company-wide foundations. Responsibilities Own the technical direction for service infrastructure within Product Foundations, including Kubernetes, AWS, Terraform, CDNs, ALBs, DNS, IAM, service networking, and related operational patterns. Up-level existing Product Foundations services by improving reliability, operability, deployment safety, infrastructure consistency, and production readiness. Partner deeply with API Platform and UI Platform on networking, DNS, CDN, load balancing, delivery, and gateway patterns for critical customer-facing interfaces. Work closely with Infrastructure, Networking, and Security teams to bring company-wide platform standards into Product Foundations and contribute PF requirements back into shared frameworks. Help drive cross-company infrastructure initiatives that Product Foundations depend on or help maintain, including Terraform CI/CD, Kubernetes networking, zero-trust service communication, policy-as-code, and cross-DC/provider networking. Build and evolve reusable service infrastructure primitives, including Helm charts, Terraform modules, GitHub Actions/GitOps workflows, service scaffolding, runbooks, and documentation. Establish durable technical standards through design docs, architecture reviews, mentorship, and hands-on implementation that help Together scale services across teams, regions, and cloud environments. Requirements 7+ years of professional experience in platform engineering, service infrastructure, SRE, distributed systems, cloud infrastructure, or related roles. Deep production experience with Kubernetes, including EKS, Helm, ArgoCD/Argo Rollouts, ingress, autoscaling, secrets, service identity, networking, and progressive delivery. Strong Terraform experience, including module design, infrastructure CI/CD, policy enforcement, production applies, and safe self-service workflows. Experience operating networking and edge infrastructure such as CDNs, ALBs/NLBs, DNS, TLS, ingress/egress controls, and traffic management. Proficiency in one or more programming languages used for infrastructure tooling and automation, such as Go, Python, TypeScript, or similar. AWS experience, ideally including EKS, IAM, VPC networking, load balancing, Route 53, CloudFront, ECR, and related service infrastructure. Direct experience with observability systems, including metrics, logs, traces, dashboards, alerting, SLOs, and incident response. Proven ability to lead cross-functional technical initiatives across product engineering, infrastructure, networking, and security teams. Strong written communication skills, with experience producing clear design docs, migration plans, operational guidance, and technical standards. Staff-level judgment: you can define ambiguous problems, make pragmatic tradeoffs, influence without authority, and leave both systems and teams better than you found them. Nice to Have Experience building internal developer platforms or paved-path service frameworks used by many engineering teams. Experience embedding infrastructure best practices into product engineering teams at scale. Experience with service mesh or zero-trust infrastructure such as mTLS, SPIFFE/SPIRE, Cilium, Istio, Linkerd, Envoy, or similar. Experience with OPA, Gatekeeper, Kyverno, Sentinel, or other policy-as-code systems. Experience with multi-region, multi-cluster, hybrid-cloud, or cross-provider service networking. Experience with supply-chain security, image signing, SBOMs, vulnerability management, or compliance automation. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $240,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
Cloud, DevOps & SREVia Greenhouse
Verified6 days ago

AI Infrastructure Systems Engineer

On-sitefull timeMid-LevelSan Francisco, United States
Apply Now

About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardware and software. A passion for solving complex infrastructure challenges through software. An automation-first mindset —if a task is repeated, your instinct is to build a system to eliminate it. Bonus Experience GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch InfiniBand or RoCE networking Bare-metal provisioning and lifecycle management Large-scale AI training or inference clusters Hardware health monitoring and predictive failure detection Distributed storage systems AI agents and autonomous infrastructure operations About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $190,000 - $270,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy .

View more...
Cloud, DevOps & SREVia Greenhouse
Verified6 days ago

About the Role We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine: design and implement the software that models the full lifecycle of a physical host from discovery, inference bring-up to GPU driver/CUDA stack, health validation, and decommission/RMA — as explicit, versioned states and transitions. Build the self-service API: design declarative APIs and a control plane so the inference team can request, scale, and tear down inference clusters with one API call — no ticket, no human in the loop. Automate self-healing: detect degraded or failed nodes, drain them safely, trigger repair or replacement, and reintroduce healthy capacity into the pool automatically. Own reliability of the pipeline: idempotency, retries, rollback, and drift detection so the provisioning system is as dependable as any other production service. Partner with the inference/ML platform team: understand the cluster shapes they need — topology, interconnect, scheduling constraints — and encode them as first-class abstractions in the platform. Engineer it like software: strong typing, automated tests, code review, versioning, and CI/CD for infrastructure code — this is a product, not a collection of Ansible playbooks. Requirements Core requirements (all levels): Strong software engineering background in Go, Python, Rust, or similar — you write and test real software for a living. Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent to run long-lived, manifest-driven workflows that survive failures and resume mid-execution. Experience building software control planes or orchestration systems that model state and reconcile it over time (e.g., Kubernetes controllers/operators, custom reconciliation loops, workflow engines). Experience with event-driven systems — designing and building software around message queues, event streams, or pub/sub (e.g., Kafka, NATS, SQS) rather than polling or cron-driven scripts. A product mindset. You’ve built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. Nice to have: Exposure to bare-metal provisioning (PXE/iPXE, Redfish/IPMI, BMC) and/or networking fundamentals (VLANs, BGP, fabric design), or GPU/accelerator infrastructure. Experience with GPU cluster software stacks (NCCL, CUDA, InfiniBand/RoCE). Prior work at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization. Systems programming in Rust or Go. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy .

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

AI Infrastructure Systems Engineer (Amsterdam & London)

On-sitefull timeMid-LevelAmsterdam, Netherlands
Apply Now

About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Hybrid at our office in Amsterdam or Remote in the UK. Responsibilites Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardware and software. A passion for solving complex infrastructure challenges through software. An automation-first mindset —if a task is repeated, your instinct is to build a system to eliminate it. Bonus Experience GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch InfiniBand or RoCE networking Bare-metal provisioning and lifecycle management Large-scale AI training or inference clusters Hardware health monitoring and predictive failure detection Distributed storage systems AI agents and autonomous infrastructure operations About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy .

View more...
Cloud, DevOps & SREVia Greenhouse
Verified6 days ago

GTM Data Analytics Engineer

On-sitefull timeMid-LevelSan Francisco, United States
Apply Now

About the Role We are seeking a highly analytical and detail-oriented Data Analytics Engineer to join our data team. This role will report into the Revenue Operations team. The ideal candidate will help build and maintain the data infrastructure and reporting used across go-to-market, finance, and product teams, working closely with analytics colleagues across both the GTM and core data teams to deliver reliable, well-documented data products. Responsibilities Develop and maintain dashboards and reports primarily supporting GTM stakeholders, with additional use by Finance, Product, and other cross-functional teams; assist with deep-dive analyses on business performance and trends. Pull and analyze data from source systems (Salesforce, Amplitude, production systems, billing) to support reporting and ad hoc requests. Use SQL to extract, clean, and analyze data, with attention to correctness and efficiency. Build foundational SQL and business knowledge and quickly grow into contributing to dbt models within Snowflake, following established dimensional modeling standards. Partner with GTM, Finance, and Product stakeholders to understand data needs and keep reporting and definitions consistent. Requirements Bachelor's degree in Business Analytics, Data Science, Statistics, Computer Science, or a related quantitative field. 1-3 years of experience in a Data Analyst, BI, or Analytics Engineering role. Strong SQL skills required, with the ability to write complex, multi-step queries involving joins and aggregations; experience with dbt and/or Python is a strong plus. Familiarity with a cloud data warehouse (Snowflake preferred) and BI tools (Hex, Metabase). Strong attention to detail and ability to manage multiple priorities in a fast-paced environment. Excellent communication skills, with the ability to explain data findings to non-technical stakeholders. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $ 120K - $150K + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy .

View more...
Data Engineering & BIVia Greenhouse
Verified6 days ago

Page 3 of 3