CoreWeave
Production Engineer on Herd
Not specifiedunknownUSD 139,000–204,000
Skills
Architect and build large-scale distributed systemsDesign data infrastructure for AI reasoningBuild agent orchestration and lifecycle componentsIntegrate AI SRE Platform with internal systemsLead architectural design discussionsPartner across Production Engineering, Data Engineering, ML Infrastructure, and PlatformDevelop services that interpret telemetryCodify operational best practices into services, APIs, and Kubernetes-native componentsParticipate in an on-call rotationPython (or similar), production microservices and data pipelinesKubernetes, container orchestration, cloud-native architecturesData systems (streaming, indexing, caching, ETL)Design for scalability, fault tolerance, and performanceBuilding RAG systems or embedding-based searchVector databases and text/signal retrieval systemsDeveloping or deploying agentic AI systemsStrong distributed-systems backgroundDesign data schemas and APIs for knowledge representation and operational reasoningChatOps frameworks or workflow orchestration (Temporal, Argo, Airflow)Background in MLOps, AI infrastructure, platform reliability engineeringObservability frameworks (Prometheus, Grafana, OpenTelemetry)
About the role
Production Engineering ensures CoreWeave’s cloud runs with world-class reliability, performance, and operational excellence. Herd is our newest innovation: an agentic AI platform that serves as CoreWeave’s intelligent SRE assistant - combining AI reasoning, data infrastructure, and observability into an autonomous operational intelligence layer for internal use. As a Production Engineer on Herd, you’ll define and build the systems that power a scalable agentic ecosystem. You’ll design distributed services and data pipelines that process, embed, and retrieve operational knowledge at scale, enabling LLM-powered agents to work alongside human engineers in production. This is a hands-on role at the intersection of AI operations, distributed systems, and data infrastructure. What You’ll Do- Architect and build large-scale distributed systems that power AI SRE Platforms. - Design data infrastructure for AI reasoning (embedding generation, context retrieval, vector stores) optimized for real-
Applying with JobFu tracks this in your pipeline automatically.
Free AI résumé tailoring, interview prep, and matches across 40+ countries.