Senior Software Engineer, Chaos Engineering
Datadog
Paris · فرنسا
تاريخ النشر : 01/10/2026
كل فرصة مرتبطة بمصدرها الأصلي. لا نضمن الحصول على وظيفة.
عن هذه الفرصة
قراءة النص الوارد من المصدر
النص مقدّم من المصدر. تحقّق من الشروط الكاملة في الإعلان الأصلي.
<p>Datadog’s Chaos Engineering team builds systems that surface reliability weaknesses before they become outages. As a Senior Software Engineer, you will initially focus on zonal resilience, building automation that helps services safely evacuate and recover from zonal failures, while also contributing to fault injection, incident replay, gameday orchestration, and reliability tooling.
You will work across engineering teams to design systems that safely exercise production failure modes and turn findings into verified remediation. You will also help advance the use of AI and automation to identify, test, and close resilience gaps as Datadog’s software and infrastructure evolve.</p> <p>At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table.
We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them.</p> <p><strong>What You’ll Do:</strong></p> <ul> <li>Build zonal-resilience automation that coordinates safe workload evacuations, switchovers, and recovery in partnership with the teams that own affected services.</li> <li>Design and build fault-injection systems for production environments, including infrastructure- and application-level testing, incident replay, and controlled resilience experiments.</li> <li>Develop safeguards such as blast-radius controls, kill switches, validation mechanisms, and rollback paths that keep production experiments contained and reversible.</li> <li>Build agents and automation that help propose failure scenarios, triage experiment results, and connect reliability findings to tracked remediation and verification.</li> <li>Lead gamedays from hypothesis and scenario design through execution, documented findings, remediation tracking, and validation of completed fixes.</li> <li>Design and implement reliable distributed systems, including gRPC services, Kubernetes controllers, and shared platform components, while contributing to technical design and mentoring other engineers.</li> </ul> <p><strong>Who You Are:</strong></p> <ul> <li>You have strong distributed systems fundamentals and can reason about consistency, failure modes, backpressure, idempotency, quorum, retries, and failure recovery.</li> <li>You understand Kubernetes workload lifecycles, including how pods, controllers, scheduling, draining, and eviction interact with resilient system design.</li> <li>You have experience designing, building,