쿠팡 · Coupang Intelligent Cloud (CIC)

Sr. Staff Observability Engineer (GPU Cloud & Telemetry Platform)

#쿠팡 채용

이 공고, 이렇게 물어볼 겁니다

Q1
GPUaaS 환경에서 수백만 개의 GPU ID와 멀티테넌트 레이블을 처리하는 고카디널리티(High-Cardinality) 메트릭을 수집하고 Mimir에 저장하기 위해, 기존 Prometheus 스케일링 아키텍처 대비 어떤 아키텍처적 트레이드오프를 감수하고 어떤 방식으로 데이터 압축 및 인덱싱 전략을 설계했는지 구체적인 사례를 들어 설명해 주십시오.
🎯 고카디널리티 데이터 처리의 기술적 난이도와 시스템 설계 역량을 확인한다. (Architecture & Cardinality Management)
Q2
GPU 워크로드의 버스트(Burst) 특성을 고려하여, 실시간으로 발생하는 ML 훈련 및 추론 스파이크 상황에서 데이터 손실 없이 낮은 레이턴시로 메트릭을 수집하기 위해 Vector 파이프라인에서 어떤 필터링 및 샘플링 전략을 적용했으며, 이 과정에서 발생한 오탐(False Positive) 또는 누락(Missing Data) 사례와 이를 어떻게 해결했는지 설명해 주십시오.
🎯 실시간 데이터 처리의 성능 최적화와 실제 운영상의 데이터 신뢰성 확보 능력을 확인한다. (Pipeline Engineering & Burst Handling)
Q3
GPU 인프라의 SLO를 정의하고 달성하기 위해, 단순한 시스템 상태 모니터링을 넘어 GPU 활용 효율성(Utilization Efficiency)이나 열/전력 이상(Thermal/Power Anomalies)과 같은 도메인 특화된 SLI/SLO를 어떻게 정의하고, 이 지표들이 실제 비즈니스 목표(예: 비용 효율성 또는 서비스 품질)와 어떻게 연계되어 의사결정으로 이어졌는지 구체적인 수치와 협업 이력을 제시해 주십시오.
🎯 도메인 지식(GPU/SRE)을 비즈니스 목표 및 측정 지표로 전환하는 전략적 사고와 영향력을 확인한다. (SLO Definition & Business Alignment)
Q4
Grafana Alloy, Mimir, Loki 스택을 설계할 때, 비용 효율성과 수평적 확장성(Horizontal Scalability)을 동시에 확보하기 위해 Metric/Log 파이프라인의 데이터 보존 정책(Retention Strategy)과 스토리지 컴팩션(Compaction) 튜닝은 어떻게 진행했는지, 그리고 이 과정에서 전체 인프라 비용을 몇 % 절감했는지 구체적인 수치와 함께 설명해 주십시오.
🎯 기술 스택에 대한 깊은 이해와 운영 비용(FinOps) 최적화 능력을 확인한다. (Cost Optimization & Operational Tuning)
Q5
최근에 팀 내에서 도입한 자동화(Automation) 및 SRE 원칙을 GPU Observability 플랫폼에 적용하여, 수동 대응 시간을 얼마나 단축했으며, 이 자동화 시스템이 엔지니어링 팀의 배포 속도(Deployment Velocity)에 구체적으로 어떤 긍정적인 영향을 미쳤는지 프로젝트 결과를 중심으로 설명해 주십시오.
🎯 기술 구현을 넘어 실제 조직 문화 및 프로세스 개선에 기여하는 리더십과 영향력을 확인한다. (SRE Adoption & Impact)
질문만 읽으면 컨닝이에요. 소리 내어 답해보세요 — 어디서 틀어지는지 짚어드립니다.

공고 내용

About Coupang

We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, we’re collectively disrupting the multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce.

We are proud to have the best of both worlds — a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurial surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day.

Our mission to build the future of commerce is real. We push the boundaries of what’s possible to solve problems and break traditional tradeoffs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world.

Role Overview

We are seeking a Sr. Staff Observability Engineer to lead the design and evolution of our observability platform for a GPU-as-a-Service (GPUaaS) infrastructure. This role will own the end-to-end telemetry strategy—from high-throughput metric ingestion to log pipelines and real-time visualization—powering deep insights into GPU clusters, datacenter systems, and distributed workloads.

You will architect and operate planet-scale telemetry pipelines leveraging Grafana Alloy, Mimir, Loki, and Vector , ensuring high-fidelity observability across GPU workloads, Kubernetes clusters, and datacenter infrastructure.

Key Responsibilities

Architectural Leadership Strategy

· End-to-End Observability Platform Ownership : Design and scale telemetry pipelines using:

· Grafana Alloy for metrics collection (Prometheus-compatible pipelines)

· Datadog Vector for high-throughput log ingestion and transformation

· Grafana Mimir for scalable time-series storage

· Grafana Loki for log aggregation and querying

· Strategic Roadmap : Define the multi-year vision for GPU infrastructure observability, transitioning from reactive monitoring to SLO-driven, predictive, and automated observability .

· High-Cardinality Telemetry Design : Optimize pipelines for GPU workloads characterized by:

· High-cardinality labels (GPU IDs, tenants, workloads)

· Burst-heavy workloads (ML training, inference spikes)

· Multi-tenant isolation requirements

Telemetry Pipeline Engineering

· Architect low-latency, high-throughput pipelines capable of ingesting:

· GPU metrics (utilization, memory, thermals, MIG partitions)

· Kubernetes and container telemetry

· Distributed system logs and traces

· Build and optimize metric pipelines (Alloy → Mimir) ensuring:

· Efficient remote_write tuning

· Cost-effective retention strategies

· Horizontal scalability and compaction tuning

· Design log pipelines (Vector → Loki) with:

· Structured logging and enrichment

· Intelligent filtering/sampling

· Stream partitioning for high-ingest environments

GPU Infrastructure Observability

· Establish deep observability into:

· GPU hardware (NVIDIA DCGM, MIG, NVLink, PCIe)

· Kubernetes GPU operators and scheduling behavior

· Network fabric (RDMA, InfiniBand, TCP performance)

· Define GPU-specific SLIs/SLOs such as:

· GPU utilization efficiency

· Job scheduling latency

· Cluster fragmentation

· Thermal and power anomalies

Visualization User Experience

· Build rich Grafana dashboards for:

· Real-time GPU fleet health

· Tenant-level usage and billing insights

· Capacity planning and forecasting

· Standardize dashboard frameworks and reusable panels across teams

· Enable self-service observability for platform and ML engineering teams

SRE, Automation Reliability

· Drive adoption of SRE principles :

· SLIs, SLOs, error budgets tailored to GPU workloads

· Integrate observability into CI/CD and IaC pipelines (Terraform/Kubernetes) :

· Automated canary analysis

· Observability-driven rollbacks

· Build automation (Go/Python) for:

· Pipeline health monitoring

· Dynamic routing and scaling of telemetry workloads

Incident Forensics Debugging

· Develop tooling and practices for cross-layer correlation :

· GPU → Node → Kubernetes → Application → Network

· Lead deep RCA efforts for:

· GPU contention issues

· Performance degradation in ML workloads

· Telemetry pipeline backpressure/failures

· Enable “needle-in-a-haystack” debugging using unified logs + metrics

Technical Leadership Collaboration

· Mentor engineers and lead design reviews for observability systems

· Act as a force multiplier across SRE, Infra, and ML platform teams

· Promote Observability-by-Design in all new GPU cluster deployments

Open Source Ecosystem Strategy

· Drive adoption

원문에서 전체 공고 보기 →

이 회사 다른 포지션 · 비슷한 공고