쿠팡 · Tech Infrastructure

Sr. Director, Back-End Engineering

#쿠팡 채용

이 공고, 이렇게 물어볼 겁니다

Q1
당신이 이전에 일한 곳에서 구축한 SRE 조직의 성과 지표는 무엇이었나요? 구체적으로 숫자로 말할 수 있나요?
🎯 이전 경험이 실제로 숫자로 입증된 성과를 가지고 있는지 확인
Q2
당신이 이전에 경험한 가장 큰 장애는 무엇이었나요? 어떻게 해결했나요? 그 때의 경험으로부터 어떤 교훈을 얻었나요?
🎯 실제 장애 상황에서 어떻게 대응하고 해결하는지 확인
Q3
당신은 어떻게 reliability governance를 구축하고 운영했나요? 구체적인 사례를 하나 꼽아서 말씀해주세요.
🎯 reliability governance를 실제로 어떻게 구축하고 운영하는지 확인
Q4
당신이 이전에 일한 곳에서 구축한 autonomous operations 시스템은 어떤 식으로 작동했나요? 구체적인 사례를 하나 꼽아서 말씀해주세요.
🎯 autonomous operations 시스템을 실제로 어떻게 구축하고 운영하는지 확인
Q5
당신은 어떻게 incident management를 개선했나요? 구체적인 사례를 하나 꼽아서 말씀해주세요. 그리고 그 때의 경험으로부터 어떤 교훈을 얻었나요?
🎯 incident management를 실제로 어떻게 개선하는지 확인
질문만 읽으면 컨닝이에요. 소리 내어 답해보세요 — 어디서 틀어지는지 짚어드립니다.

공고 내용

Sr. Director, Site Reliability Engineering

Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define and lead company-wide reliability, resilience, scalability, and operational excellence. This leader will transform reliability from a collection of team-specific practices into platform mechanisms that services inherit by tier, while advancing incident response toward an intelligent, AI-assisted, and increasingly autonomous operating model. We are looking for a visionary, industry-recognized technology leader who has previously conceived, built, and scaled a comparable SRE, production engineering, resilience, or autonomous-operations organization at a leading global technology company. The successful candidate must combine deep technical credibility with the organizational leadership required to align executives, influence architecture across the company, and build a world-class leadership bench.

Key Responsibilities

· Set a bold, multi-year vision for company-wide reliability, resilience, and autonomous operations, and translate that vision into an executable roadmap with measurable business outcomes.

· Define and own the SRE strategy, operating model, engineering standards, and reliability governance across Coupang.

· Build platform mechanisms that allow services to inherit reliability requirements based on service tier rather than recreate them independently.

· Lead initiatives that materially improve availability, resilience, scalability, performance, and operational readiness.

· Partner with engineering, product, infrastructure, security, finance, and business leaders to align reliability investments with customer and business priorities.

· Own executive reliability metrics, including availability, detection and recovery performance, change risk, incident recurrence, capacity readiness, and operational toil.

· Build and scale a world-class SRE organization capable of influencing engineering practices across the company.

Reliability Strategy, SLOs Engineering Governance

· Establish and evolve service-tier definitions, SLOs, SLAs, error budgets, reliability scorecards, and objective certification mechanisms such as RBD/RBO.

· Create clear reliability requirements for Tier 0, Tier 1, and Tier 2 services, including redundancy, load testing, disaster recovery, observability, and incident response.

· Ensure reliability governance is embedded in architecture, development, release, and production operations rather than applied as a final review.

· Drive systematic reduction of recurring incidents, reliability risks, operational debt, and unsafe change patterns.

· Influence company-wide architecture for graceful degradation, fault isolation, load shedding, circuit breaking, and failure containment.

Incident Management Autonomous Operations

· Transform incident management into a fast, disciplined, data-driven, and increasingly autonomous operating model.

· Enable AI-assisted detection, event correlation, triage, escalation, root-cause drafting, remediation recommendations, and selected guardrailed auto-remediation.

· Improve incident command, on-call quality, escalation mechanisms, communication, post-incident learning, and corrective-action completion.

· Reduce noisy alerts, manual on-call work, repeated diagnosis, and time spent coordinating across fragmented systems.

· Use incident and telemetry data to continuously improve platform standards, testing, capacity models, and engineering roadmaps.

Disaster Recovery, Resilience Capacity

· Own the strategy and execution model for disaster recovery, regional resilience, availability-zone loss, capacity-constrained recovery, and critical business continuity.

· Build reusable DR and failover mechanisms that services inherit from the platform rather than implement as bespoke projects.

· Establish objective RPO/RTO targets, automated readiness gates, regular game days, fault injection, and evidence-based recovery certification.

· Drive proactive and intelligent capacity management using forecasting, reservations, workload prioritization, and automated response to demand and failure scenarios.

· Partner with compute, traffic, networking, storage, and application leaders to enable safe zone evacuation, regional failover, and surge readiness.

Observability, Testing Reliability Intelligence

· Partner with Observability and TestOps leaders to integrate logs, metrics, traces, continuous profiling, testing, and incident intelligence into one reliability feedback loop.

· Ensure every critical service has actionable telemetry, meaningful SLOs, release-quality signals, and production-readiness evidence.

· Use production incidents and operational patterns to drive targeted integration, load, resilience, and regression testing.

· Establish executive reliability dashboards that provide trusted views of service health, risk, capacity, and operational effectiveness.

Tal

원문에서 전체 공고 보기 →

이 회사 다른 포지션 · 비슷한 공고