Site Reliability Engineer
Key Responsibilities
- Own DragonX’s architecture, together with the CTO: participate in and drive design decisions — service boundaries, data flows, dependencies, tradeoffs — rather than just inheriting them. There’s no existing SA, so a meaningful part of this is building that function from where it stands today.
- Own deployment & CI/CD: build and maintain provisioning, automation, and release pipelines across Viettel, CMC, FPT, and Azure.
- Own incident response: be the first line on production issues — triage, mitigate, drive root-cause analysis, and close the loop with permanent fixes.
- Scale the system: make fast, sound calls on scaling up/down under real demand shifts, balancing cost, performance, and risk.
- Own security operations: hardening, access control, patching, and security monitoring across infrastructure and client-facing environments.
- Compensate for local-cloud gaps: where Viettel/CMC/FPT lack the managed services Azure/AWS take for granted (managed DBs, autoscaling, service mesh, etc.), build the equivalent yourself.
- Work directly with developers and POs: keep environments healthy and unblock delivery, including translating technical tradeoffs for non-technical stakeholders.
Skills and qualifications
(1) Must have:
- Bachelor’s degree in Information Technology, Computer Science, Software Engineering, or a related field.
- Hands-on production experience with at least one Vietnamese local cloud (Viettel, CMC, or FPT) and a major public cloud (Azure preferred).
- Strong DevOps fundamentals: CI/ CD pipelines, Infrastructure as Code (Terraform/ Ansible or equivalent), containerization/ orchestration (Docker/ Kubernetes), monitoring & observability (Prometheus/ Grafana, ELK, or equivalent).
- Strong SecOps fundamentals: hardening, IAM/ access control, patch management, security monitoring, basic compliance awareness.
- Demonstrated ability to operate without managed-service safety nets — has built or run infrastructure from bare metal / self-managed VMs before, not just consumed cloud APIs.
- Proven experience making architecture decisions, not just implementing someone else’s design — has owned or co-owned system design (service boundaries, data flow, scaling strategy) for a non-trivial production system before. This is the SA half of the role; screen for it as rigorously as the ops half.
- Fast, sound judgment under pressure — comfortable owning an incident end-to-end at 2am if needed.
- Strong communication: equally comfortable in a technical debugging session with developers and a status update to a non-technical PO.
(2) Nice-to-have
- Experience managing client-facing or multi-tenant/ multi-environment deployments.
- Background in finance, securities, banking, or another security-sensitive/regulated domain.
- Experience leading a cloud migration or running a hybrid (on-prem + public cloud) setup end-to-end.
- Working familiarity with JavaScript, Flutter, SQL, and/ or Python — enough to read code and collaborate with developers on root-causing app-level issues, not to build features