센드버드 · Engineering

Machine Learning Engineer

#센드버드 채용

이 공고, 이렇게 물어볼 겁니다

Q1
고객 컨텍스트를 기억하고 여러 채널을 연결하는 에이전트 시스템을 설계할 때, 메모리(Memory)와 검색(Retrieval) 시스템의 아키텍처를 어떻게 결정했으며, 그 결정이 실제 에이전트의 정확도와 응답 속도에 어떤 영향을 미쳤는지 구체적인 사례를 들어 설명해 주세요. (꼬리질문: RAG와 Memory를 분리했을 때의 장단점과 트레이드오프는 무엇입니까?)
🎯 복잡한 에이전트의 핵심인 컨텍스트 관리 및 검색 시스템에 대한 깊은 이해와 설계 능력을 확인한다.
Q2
실제 운영 환경에서 LLM 기반 에이전트의 신뢰성과 품질을 측정하기 위해 어떤 평가 지표(Metrics)와 평가 시스템(Evaluation System)을 구축했는지 설명해 주십시오. 특히, 단순히 정확도를 넘어 고객 경험 품질(CX Quality)과 안전성(Safety)을 어떻게 정량화했는지 구체적인 데이터와 함께 제시해 주세요. (꼬리질문: Production Regression을 감지하기 위한 자동화된 평가 파이프라인은 어떻게 구성했습니까?)
🎯 AI 모델의 성능을 측정하고, 이를 실제 비즈니스 목표(고객 경험)와 연결하여 신뢰성 있는 결과를 도출하는 능력을 확인한다.
Q3
대규모 메시징 및 음성 API를 기반으로 하는 Sendbird 환경에서, LLM을 실제 고객 여정(Customer Journey)에 통합하여 워크플로우 자동화 기능을 구현할 때 발생했던 가장 큰 엔지니어링 난제와 이를 해결하기 위해 적용했던 모델 적응(Fine-tuning/Adaptation) 전략은 무엇이었습니까? (꼬리질문: 특정 도구 사용(Tool Use)의 실패율을 줄이기 위해 어떤 방식으로 모델의 추론 과정을 제어했는지 설명해 주세요.)
🎯 AI 기술을 실제 엔터프라이즈 시스템(워크플로우 자동화, API 통합)에 적용하는 실전 경험과 모델 적응 능력을 확인한다.
Q4
수십억 건의 메시지를 처리하는 대규모 시스템에서, 추론(Inference) 시스템의 지연 시간(Latency), 처리량(Throughput), 그리고 비용을 최적화하기 위해 어떤 기술적 선택(모델 선택, 서빙 아키텍처, 배치 처리 등)을 했는지 구체적인 수치와 함께 설명해 주십시오. (꼬리질문: 실시간성과 비용 효율성 사이에서 엔지니어링 팀과 제품팀 간에 발생했던 주요 의견 충돌과 이를 어떻게 조율했습니까?)
🎯 대규모 프로덕션 환경에서의 ML 시스템 최적화 및 MLOps 엔지니어링 역량을 확인한다.
Q5
제품 관리자(PM)로부터 모호한 AI 기능 요구사항을 받아, 이를 실제 고객에게 제공할 수 있는 신뢰할 수 있는 에이전트 기능으로 전환하는 과정에서, 기술적 제약 조건과 제품 목표 사이의 우선순위를 어떻게 설정하고 협업했는지 구체적인 프로세스를 설명해 주세요. (꼬리질문: 초기 프로토타입 단계에서 기대치와 실제 구현 결과가 크게 달랐을 때, 이를 어떻게 관리하고 다음 단계의 제품 방향을 조정했습니까?)
🎯 연구자(Research)의 아이디어를 제품(Product)으로 전환하는 크로스펑셔널(Cross-functional) 협업 능력과 제품 중심 사고를 확인한다.
질문만 읽으면 컨닝이에요. 소리 내어 답해보세요 — 어디서 틀어지는지 짚어드립니다.

공고 내용

Sendbird is building AI agents for customer experience. Our platform already powers billions of conversations every month across chat, voice, video, and messaging APIs. We are now using that foundation to build agents that understand customer context, reason over business data, and take reliable action in production.

We are looking for a Machine Learning Engineer to research, build, and productionize new capabilities for those agents. This role sits at the intersection of agent product development, applied AI research, and production engineering. You will work on systems that enterprise customers depend on every day, not demos or isolated prototypes.

About Sendbird and delight.ai

Sendbird has spent more than a decade building communication infrastructure for in-app chat, voice, video, and messaging APIs. More than 4,000 brands use our platform, including DoorDash, Match Group, Noom, Yahoo Sports, and Rakuten. Our systems support more than 7 billion messages every month.

In 2024, we made a strategic shift toward AI-first customer experience. In 2025, we launched our enterprise AI agent product, delight.ai. Delight.ai helps businesses deliver customer support and engagement that is faster, more contextual, and more personal. Unlike simple FAQ bots, our agents are built to remember customer context, use tools, retrieve relevant knowledge, connect across channels, and handle real customer workflows with accuracy and control.

The Role

As a Machine Learning Engineer, you will design, build, evaluate, and ship new capabilities for our AI agents. You will work across agent architecture, retrieval, memory, planning, tool use, workflow automation, voice, evaluation, data pipelines, model adaptation, inference, and production integration.

This is a hands-on engineering role for someone who can turn AI research and product ideas into reliable customer-facing features. Some problems will require training, fine-tuning, or adapting models. Others will require better retrieval, better context handling, better tools, stronger evaluation, or a more thoughtful product design. The right person knows how to choose the right lever and ship the result.

What you will work on

· Research, prototype, evaluate, and productionize new agent capabilities, including memory, planning, tool use, workflow execution, reasoning, and context management.

· Improve the quality, reliability, and usefulness of customer-facing AI agents through better retrieval, prompting, evaluation, model adaptation, product behavior, and system design.

· Build the core intelligence layer for our agents, including retrieval systems, memory, tool-calling pipelines, workflow orchestration, and agentic reasoning.

· Train, fine-tune, and adapt LLMs and related models when model-level work is the right way to improve agent quality or product capability.

· Build data pipelines for model training, agent evaluation, and product improvement, including labeling workflows, dataset construction, quality checks, and feedback loops.

· Design evaluation systems that measure task completion, accuracy, latency, cost, reliability, safety, customer experience quality, and production regressions.

· Optimize inference systems for production, including model selection, serving architecture, batching, caching, latency, throughput, and cost.

· Integrate voice AI, internal tools, third-party APIs, and workflow systems so agents can take useful actions across real customer journeys.

· Partner with product managers, engineers, and customer-facing teams to turn ambiguous AI product requirements into shippable agent features.

What we are looking for

· 5+ years of professional experience in ML engineering, machine learning, data science, or backend engineering with substantial AI/ML ownership.

· Experience building AI-powered product features for real users, preferably involving agents, conversational AI, workflow automation, retrieval, or LLM systems.

· Hands-on experience training, fine-tuning, or adapting LLMs or other deep learning models for production use cases.

· Strong practical knowledge of model training workflows, including dataset preparation, experiment tracking, evaluation, model selection, and deployment.

· Experience building production LLM systems, such as RAG, tool calling, agents, model orchestration, prompt systems, or evaluation pipelines.

· Strong Python skills and experience building production-grade services.

· Working knowledge of inference and serving tradeoffs, including latency, throughput, GPU utilization, model size, batching, caching, and cost.

· Experience with ML infrastructure or MLOps tools used for data pipelines, training jobs, model deployment, monitoring, or evaluation.

· Strong debugging instincts for agent systems, including failure analysis, hallucination reduction, retrieval quality, model behavior, tool-use failures, and regressions.

· Clear communication skills, especially when explaining technical tradeoffs to product managers, engineers, and leade

원문에서 전체 공고 보기 →

이 회사 다른 포지션 · 비슷한 공고