LogoKode$word
Togetherai logo
Verified Tech Organization

Careers at Togetherai

Browse and filter through all verified positions currently open at Togetherai.

Total Company Roles30
Matching Filter30

All Openings (30)

Ordered by most recently published

Senior Machine Learning Engineer, Voice AI

On-sitefull timeSeniorSan Francisco, United States
Apply Now

About the Role Together AI is building the best inference infrastructure for voice applications. Our Voice AI platform powers production-grade, real-time voice agents and applications — serving speech-to-text and text-to-speech models with best-in-class latency and reliability. We're looking for a Senior ML Engineer to drive the model serving layer for voice workloads. You'll work hands-on with inference engines like TRT-LLM and SGLang to optimize how we serve models like Whisper, Parakeet, Orpheus, and Kokoro — pushing latency and throughput to the frontier. You'll profile GPU utilization, design batching strategies for streaming audio, and ensure new model architectures can go from research to production quickly. This is a foundational hire on a small, high-impact team. Voice inference has unique challenges — streaming audio, tokenization, real-time latency budgets — that require dedicated ML engineering focus. You'll shape how Together serves voice models as the industry moves from pipeline architectures (ASR → LLM → TTS) toward end-to-end speech-to-speech. Own the model serving stack that powers Together's voice platform across STT, TTS, and speech-to-speech. Work directly with state-of-the-art accelerators (H100s, H200s, B200s) to optimize voice model inference. Collaborate with model partners (Cartesia, Deepgram, Rime, and others) to bring their models to production on Together's infrastructure. Build quality evaluation frameworks that guide model selection for customers and inform the roadmap. Join a small, early-stage team with outsized impact on a fast-growing product area. Responsibilities Optimize inference performance for voice models (STT, TTS, speech-to-speech) — targeting best-in-class TTFB, throughput, and GPU utilization across our curated model set. Productionize voice models on serverless and dedicated endpoints, including batching strategies, streaming inference, and memory management tailored to audio workloads. Build and maintain a voice model evaluation framework — measuring WER across accents, languages, and noise conditions for STT; naturalness, latency, and pronunciation accuracy for TTS. Enable new model architectures in our serving stack as the field evolves, including audio-native LLMs, codec-based models (SNAC), and speech-to-speech systems. Collaborate with model partners to integrate and optimize their models (Cartesia, Deepgram, Rime, and others) running on Together's infrastructure. Profile and debug performance across the full inference stack — from GPU kernels to framework-level bottlenecks — and ship measurable improvements. Work with the platform engineering side of the team to ensure the serving layer meets the latency and reliability requirements of real-time voice APIs. Contribute to voice model fine-tuning capabilities (STT and TTS) as we enable customers to build differentiated voice experiences on Together. Lay the groundwork for multiple new products down the line. Requirements 5+ years of experience in ML engineering, with a focus on model serving, inference optimization, or ML infrastructure. Hands-on experience with LLM serving engines (vLLM, SGLang, TensorRT-LLM, or similar) — comfortable reading and modifying engine internals, not just using APIs. Strong proficiency in Python and PyTorch; experience with GPU profiling and optimization (CUDA, memory management, kernel-level debugging). Track record of shipping ML systems to production with measurable performance improvements. Strong product sense — you think about what developers building voice apps actually need, not just what's technically interesting. Comfort working on a small, early-stage team where you'll wear multiple hats and move fast. Experience with speech and audio ML (ASR, TTS architectures, audio signal processing) is a strong plus but not required — you can learn this quickly if you have strong ML engineering fundamentals. Familiarity with audio codecs and tokenization schemes (SNAC, Encodec, DAC) is a plus. Experience training or fine-tuning speech models is a plus. Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field, or equivalent practical experience About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $200,000 - $260,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
AI / ML & Data ScienceVia Greenhouse
Verified6 days ago

Senior Product Engineer, Fullstack

On-sitefull timeSeniorSan Francisco, United States
Apply Now

About the Role Together AI's product is what developers touch every day — Playground, Model Garden, Fine-Tuning, File Management, Batch Processing, Model Evals. These surfaces are how users experience the platform, and they need to be fast, intuitive, and something developers genuinely love using. That's what this role is about. We're growing the Product Engineering team and looking for a fullstack engineer who's motivated by shipping things users genuinely value — not features that just look good in a slide deck. You'll plug directly into new and existing workstreams and start shipping quickly, working across our core product surfaces to improve existing capabilities and build new ones. This is a high-ownership role on a small, fast-moving team, where one strong engineer moves the needle visibly. You'll see your work in production, instrument and measure it, and use that to shape what comes next. Responsibilities Ship improvements and new capabilities across Together AI's core product surfaces — Playground, Model Garden, Fine-Tuning, File Management, Batch Processing, and Model Evals Own projects and features end-to-end, from data model and API integration through the user interface Collaborate closely with other engineers to build clean, typed, reusable interfaces in our web application Instrument and measure what you ship — add product analytics, monitor system health, and prove users are getting value with data Work across design, product, and engineering to ship things that unlock user value Raise the bar on frontend and fullstack quality across the team through code review and collaborative architecture decisions Requirements 4–7 years of experience building and shipping real product to real users Strong JavaScript/TypeScript fundamentals; TypeScript experience highly valued Strong React expertise in a large-scale, production environment Full-stack fluency: you connect the dots from data model to API design to state management to the UI, and know how to make the right call at each layer in service of clean, maintainable code and an intuitive user experience A pragmatic approach to engineering — you know when to move fast and when to slow down, and you make that call based on user value Curiosity about AI and LLMs — you want to understand the models and products you're building on top of, not just the code around them Adept at building and improving agentic workflows, with good judgment for when to lean on AI and when not to Cross-functional "we ship it" mentality — no narrow ownership, no finger-pointing Next.js, Tailwind, and shadcn/ui experience is a plus Familiarity with product analytics (e.g., Amplitude) and systems monitoring (e.g., Grafana) is a plus Experience with CI/CD workflows is a plus Experience with a backend language like Go or Python About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $160,000 - 230,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy .

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Senior Product Manager, Model APIs & Developer Experience

On-sitefull timeSeniorSan Francisco, United States
Apply Now

About the Role Together serves a broad portfolio of open-weight models, backed by infrastructure designed to run them quickly and efficiently. The interfaces around those models shape the entire customer experience: how easily a developer can migrate an application, how reliably an agent can use a model, and how a team runs large asynchronous workloads. As Senior Product Manager for Model APIs and Developer Experience, you will own those interfaces. Your scope will extend beyond chat to include the APIs and developer experience for image, video, voice, and passthrough models. You will also own compatibility with the wider developer and agent ecosystem, the voice consumption experience and associated model partnerships, and the evolution of our batch inference product. These areas are connected by a single goal: make Together's models easy to adopt and reliable to build on for developers and autonomous agents. This is a hands-on product role. You will read API specifications, test integrations, inspect code when useful, and build prototypes or demo applications to sharpen your thinking. You will work closely with engineering and research, own the product direction, make thoughtful tradeoffs, ship, and learn from how customers use what we build. Responsibilities Set the product direction and roadmap across chat and multimodal APIs, ecosystem compatibility, voice, and batch inference. Close the API and behavioral gaps that make it harder for customers to move workloads from proprietary model providers to Together. Develop a grounded view of how Together's APIs perform in the developer and agent ecosystem, and turn the most important gaps into clear product priorities. Define an intuitive API and consumption experience for real-time and agentic voice applications. Build productive relationships with voice model partners and align internal and external teams around a strong joint product experience. Reimagine batch inference across the job lifecycle, developer experience, completion guarantees, pricing, and packaging. Build lightweight prototypes and demo applications to test product ideas and reduce uncertainty before committing significant engineering resources. Work closely with engineering and research to understand technical constraints, make product tradeoffs, and deliver reliable customer experiences. Talk with customers and study product usage to separate isolated requests from patterns that should shape the platform. Create a clear, active roadmap and communicate why we are making specific investments and what we expect them to change. Requirements Meaningful experience with APIs and developer tools, whether you designed them, built them, or owned them as a product manager. Strong product judgment and the ability to make clear decisions in an evolving market. A willingness to get close to the work by reading specifications and code, testing integrations directly, and building working prototypes with AI development tools. A track record of earning trust with engineering and external partners through preparation, sound technical judgment, clear communication, and follow-through. A bias toward experimentation and iteration. You know how to gather enough evidence to make a decision, ship, and adjust based on what you learn. Strong communication and relationship-building skills, including the ability to align teams and move work forward across company boundaries.. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $200 - 280k + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy .

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Senior Software Engineer, Customer Insights

On-sitefull timeSeniorSan Francisco, United States
Apply Now

About the Role Together.ai is looking for a Senior Software Engineer to join the Customer Insights team. Customer Insights owns the customer-facing visibility layer of Together's Cloud — the systems that help customers understand what's happening in their environment, trace back what happened and why, and act on it with confidence. Every time a customer needs to understand their own activity — what happened, who did it, when, and whether it needs attention — they rely on the systems we build. Whether it's a developer inspecting recent activity, an enterprise team auditing their organization, or an operator diagnosing an anomaly, we make that visibility clear and trustworthy: simple for everyday cases, robust enough for complex organizational needs. We're turning today's fragmented visibility patterns into coherent product and platform foundations, and building toward a next generation of insight tooling that summarizes activity, explains anomalies, and correlates signals across surfaces — for both human operators and autonomous agents. You'll take ambiguous problems and turn them into a clear technical direction. You'll set the technical approach for significant pieces of the system, weigh in on architecture beyond your immediate area, and help less experienced engineers navigate unfamiliar problem spaces. You'll have real autonomy over how problems get solved, paired with a team that reviews, debates, and learns in the open. Responsibilities Design and ship core features of Together's Customer Insights platform — the systems customers rely on to see, trace, and act on their own activity. Drive the technical design for new capabilities, including proposing architecture, writing design docs, and getting buy-in across teams. Partner with Data Platform and Observability engineering teams to leverage Together’s core data stack and capabilities to build out Customers Insights data collection and stream processing systems. Own significant areas of the system end-to-end — from design through implementation, testing, rollout, and long-term health. Write and maintain critical-path backend and product code used by multiple teams and customer-facing surfaces. Identify risks and blockers early, drive resolution across team boundaries. Requirements 4+ years building and operating production distributed systems or customer-facing backend platforms. Experience designing and scaling data pipelines for high-volume ingestion and real-time querying specifically (the "at scale" specialization, distinct from #1's general systems experience). Strong programming skills in one or more of Go, Python, Java,TypeScript, C++, or similar production languages. Excellent communication skills – able to write clear design docs and work effectively with both technical and non-technical team members Strong data modeling instincts — comfort working with schemas, queries, and relational and/or non-relational data at scale. Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience. Nice to Have Interest in, or some exposure to, stream processing, data pipelines, notification systems, or analytical systems (this is a great team to learn them on) About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $180k - 250k + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy .

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Senior Software Engineer — Infra Agent Systems

On-sitefull timeSeniorSan Francisco, United States
Apply Now

About the Role Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure. We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling. You’ll work across two areas: Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack. Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve. We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure. This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation , solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure. Why this Role You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective. You’ll have the opportunity to build foundational systems from the ground up, work on infrastructure at massive scale, and help define how self-improving AI agents operate real-world AI infrastructure. Responsibilities Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets. Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents. Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions. Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs. Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations. Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops. Turn what agents learn in production into reliable, reviewed software and automation. Requirements 5+ years of experience building production backend systems, distributed systems, or infrastructure platforms. Strong systems design skills and experience owning significant systems from design through production. Depth in at least one of the following: AI agent systems, orchestration, tool use, evaluation, or grounding Knowledge graphs or graph data modeling Search, retrieval, ranking, RAG, or semantic search systems Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across complex systems. Experience with Kubernetes, GitOps such as ArgoCD, infrastructure-as-code, and cloud platforms. Comfortable working across languages such as Go, TypeScript, Python, or Rust. Experience in the following is a plus: GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers Graph databases Event-driven systems and messaging platforms such as NATS or Kafka Observability platforms such as Prometheus and Grafana Building evaluation frameworks or improving the quality and reliability of LLM-powered systems About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $250,000 - $300,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy .

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Senior Software Engineer — Infra Agent Systems

On-sitefull timeSeniorAmsterdam, Netherlands
Apply Now

About the Role Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure. We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling. You’ll work across two areas: Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack. Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve. We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure. This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation , solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure. responsible for delivering the software but also for operating and supporting it in production. Why this Role You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective. You’ll have the opportunity to build foundational systems from the ground up, work on infrastructure at massive scale, and help define how self-improving AI agents operate real-world AI infrastructure. Hybrid in Amsterdam Responsibilities Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets. Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents. Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions. Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs. Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations. Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops. Turn what agents learn in production into reliable, reviewed software and automation. Requirements 5+ years of experience building production backend systems, distributed systems, or infrastructure platforms. Strong systems design skills and experience owning significant systems from design through production. Depth in at least one of the following: AI agent systems, orchestration, tool use, evaluation, or grounding Knowledge graphs or graph data modeling Search, retrieval, ranking, RAG, or semantic search systems Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across complex systems. Experience with Kubernetes, GitOps such as ArgoCD, infrastructure-as-code, and cloud platforms. Comfortable working across languages such as Go, TypeScript, Python, or Rust. Experience in the following is a plus: GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers Graph databases Event-driven systems and messaging platforms such as NATS or Kafka Observability platforms such as Prometheus and Grafana Building evaluation frameworks or improving the quality and reliability of LLM-powered systems About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy .

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

About the Role Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure. We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling. You’ll work across two areas: Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack. Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve. We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure. This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation , solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure. responsible for delivering the software but also for operating and supporting it in production. Why this Role You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective. You’ll have the opportunity to build foundational systems from the ground up, work on infrastructure at massive scale, and help define how self-improving AI agents operate real-world AI infrastructure. Remote based in India Responsibilities Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets. Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents. Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions. Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs. Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations. Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops. Turn what agents learn in production into reliable, reviewed software and automation. Requirements 5+ years of experience building production backend systems, distributed systems, or infrastructure platforms. Strong systems design skills and experience owning significant systems from design through production. Depth in at least one of the following: AI agent systems, orchestration, tool use, evaluation, or grounding Knowledge graphs or graph data modeling Search, retrieval, ranking, RAG, or semantic search systems Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across complex systems. Experience with Kubernetes, GitOps such as ArgoCD, infrastructure-as-code, and cloud platforms. Comfortable working across languages such as Go, TypeScript, Python, or Rust. Experience in the following is a plus: GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers Graph databases Event-driven systems and messaging platforms such as NATS or Kafka Observability platforms such as Prometheus and Grafana Building evaluation frameworks or improving the quality and reliability of LLM-powered systems About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy .

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Senior Software Engineer — Infra Agent Systems UK

On-sitefull timeSeniorLondon, United Kingdom
Apply Now

About the Role Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure. We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling. You’ll work across two areas: Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack. Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve. We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure. This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation , solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure. responsible for delivering the software but also for operating and supporting it in production. Why this Role You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective. You’ll have the opportunity to build foundational systems from the ground up, work on infrastructure at massive scale, and help define how self-improving AI agents operate real-world AI infrastructure. Remote based in the UK Responsibilities Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets. Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents. Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions. Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs. Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations. Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops. Turn what agents learn in production into reliable, reviewed software and automation. Requirements 5+ years of experience building production backend systems, distributed systems, or infrastructure platforms. Strong systems design skills and experience owning significant systems from design through production. Depth in at least one of the following: AI agent systems, orchestration, tool use, evaluation, or grounding Knowledge graphs or graph data modeling Search, retrieval, ranking, RAG, or semantic search systems Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across complex systems. Experience with Kubernetes, GitOps such as ArgoCD, infrastructure-as-code, and cloud platforms. Comfortable working across languages such as Go, TypeScript, Python, or Rust. Experience in the following is a plus: GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers Graph databases Event-driven systems and messaging platforms such as NATS or Kafka Observability platforms such as Prometheus and Grafana Building evaluation frameworks or improving the quality and reliability of LLM-powered systems About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Senior Software Engineer, Observability

On-sitefull timeSeniorSan Francisco, United States
Apply Now

About the Role Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. The AI Infrastructure team at Together AI is at the forefront of building and scaling the foundational systems that power our generative AI platform. The storage and observability team is crucial for designing, implementing, and maintaining robust distributed storage solutions, ensuring seamless data access and management. They are also responsible for developing comprehensive observability platforms, providing critical insights into system performance and GPU utilization, and proactively identifying and resolving issues. Responsibilities Design and implement a scalable observability platform (metrics, logs, traces) using tools like Prometheus, Grafana, ClickHouse, ClickStack, and OpenTelemetry, including telemetry data pipelines and log aggregation workflows. Develop automated monitoring, alerting, and anomaly detection systems, including SLIs/SLOs, runbooks, and predictive analytics for critical services. Build and deploy custom observability tools and infrastructure-as-code using Go, Python, Terraform, Ansible, and Helm. Collaborate with engineering teams to enhance distributed tracing and application monitoring, and lead incident response with post-mortem analysis. Define observability best practices. Requirements Expertise in observability platforms (Prometheus, Grafana, ClickStack, OpenTelemetry) and cloud-native monitoring services (AWS, GCP, Azure). Strong programming skills in Go, Python, or similar languages, with proficiency in infrastructure-as-code tools (Terraform, Ansible, Helm). Experience designing, operating, and scaling large-scale distributed systems and pipelines for high-volume data ingestion and real-time querying. Deep understanding of containerization (Docker) and orchestration (Kubernetes). Knowledge of microservices architecture, service mesh technologies, CI/CD pipelines, and GitOps workflows. Expertise in managing databases (PostgreSQL, MongoDB, Redis) and time-series databases with high-cardinality data. Preferred Experience monitoring AI/ML infrastructure, GPU clusters, and custom metrics for model performance and training pipelines. Background in high-frequency, low-latency systems monitoring, chaos engineering, and reliability testing. Contributions to open-source observability projects. Familiarity with security monitoring and compliance frameworks. About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $200,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Senior Software Engineer - Together Cloud Platform

On-sitefull timeSeniorSan Francisco, United States
Apply Now

About the Role Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. As a Senior Backend Engineer, you will play a key role in building the next generation AI cloud platform – a highly available, global, blazing-fast cloud infrastructure that virtualizes cutting-edge ML hardware (GB200s/GB300s, BlueField DPUs) and enables state-of-the-art ML practitioners with self-serve AI cloud services, such as on-demand + managed Kubernetes and Slurm clusters. This platform serves both our internal StaaS products (inference, fine-tuning) and our external cloud customers, spanning dozens of data centers across the world. Some of what you’ll work on: Work on a distributed GPU scheduling system for the on-demand clusters product, Instant Clusters. Build out a global management plane for managing our data center compute, networking, and storage. Design and build new customer-facing cloud platform services, delivering killer enterprise AI cloud features. Responsibilities Identify, design, and develop foundational backend services that power Together’s cloud platform Analyze and improve the robustness and scalability of existing distributed systems, APIs, databases, and infrastructure Partner with product teams to understand functional requirements and deliver solutions that meet business needs Write clear, well-tested, and maintainable software and IaC for both new and existing systems Conduct design and code reviews, create developer documentation, and develop testing strategies for robustness and fault tolerance Participate in an on-call rotation to address critical incidents when necessary Requirements 5+ years of demonstrated experience in building large scale, fault tolerant, distributed systems and API microservices Experience designing, analyzing and improving efficiency, scalability, and stability of various system resources Excellent communication skills – able to write clear design docs and work effectively with both technical and non-technical team members Demonstrated experience with building and operating high-performance and/or globally distributed microservice architectures across one or more cloud providers (AWS, Azure, GCP) Strong systems knowledge across compute, networking, and storage, including concurrency, memory management, performant I/O, and scale Experience developing against and managing a relational database, such as PostgreSQL Expert-level programmer in one or more of programming language (Golang preferred) Proficiency in version control practices and integrating IaC with CI/CD pipelines. Experience with Kubernetes and containers preferred Experience building and operating data infrastructure (Kinesis, Airflow, Kafka, etc) a plus Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month. Compensation We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $160,000 - $230,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our privacy policy at https://www.together.ai/privacy

View more...
Software EngineeringVia Greenhouse
Verified6 days ago

Page 2 of 3