Chiehting Lee李玠廷

Senior SRE

Keeping distributed systems reliable — from Kubernetes networking to multi-cloud infrastructure as code.

SCROLL
12+ Years of experience Software and infrastructure engineering
5 / 60 Clusters / nodes Kubernetes running in production
99.95% Service level agreement Sustained across production services
4 Cloud providers Provisioned with Terraform & Terragrunt

About

Site reliability engineer based in Taipei, working where platform, network and operations meet.

Introduction

I am a site reliability engineer with more than twelve years in the software industry, split between back-end development and the infrastructure that keeps it running. My work centres on making distributed systems observable, predictable and cheap to operate — defining service levels that mean something, automating the paths that used to need a human, and understanding a platform deeply enough to debug it at three in the morning.

Most of my recent time goes into Kubernetes platform engineering: cluster networking with Cilium and eBPF, certificate and identity automation, and describing multi-cloud infrastructure entirely in Terraform and Terragrunt. I write up most of what I learn — the blog is where those working notes end up.

Focus areas

  • Reliability engineeringSLI / SLO / SLA definition, error budgets, incident response and postmortems
  • Kubernetes platformCluster lifecycle, CNI, ingress, certificate automation, workload scheduling
  • Infrastructure as codeMulti-cloud provisioning with Terraform and Terragrunt, configuration with Ansible
  • ObservabilityMetrics, logs and traces; alerting that maps to user-visible impact
  • Identity & accessOIDC, LDAP / FreeIPA, VPN, and certificate lifecycle management
  • Network & performance diagnosticsTCP, IPv6 and QUIC behaviour; measurement with iPerf, MTR and JMeter

Skills

Grouped by the problem they solve, rather than by how long the list can get.

Cloud

  • Amazon Web Services
  • Microsoft Azure
  • Alibaba Cloud
  • Huawei Cloud
  • Google Cloud

Container & Orchestration

  • Kubernetes
  • Docker
  • Helm
  • ArgoCD
  • Harbor

Infrastructure as Code & CI/CD

  • Terraform
  • Terragrunt
  • Ansible
  • GitLab CI
  • GitHub Actions

Networking

  • Cilium / eBPF
  • MetalLB
  • CoreDNS
  • ingress-nginx
  • Nginx
  • iptables

Observability

  • Prometheus
  • Grafana
  • Loki
  • OpenTelemetry
  • Elasticsearch

Identity & Security

  • OIDC / Keycloak
  • FreeIPA
  • LDAP
  • cert-manager
  • Vault
  • OpenVPN

Data & Messaging

  • MySQL
  • MariaDB
  • PostgreSQL
  • Redis
  • MQTT / EMQX

Languages & Systems

  • Go
  • Python
  • Shell Script
  • JavaScript / Node.js
  • Linux

Writing

Working notes from the platform — over a hundred posts, still going.

Latest posts

ALL POSTS →

Resume

Employment history.

  1. 立頂數位 · IoT industry Sep. 2023 — present

    Senior Software Engineer

    • Operate the production Kubernetes estate — 5 clusters across roughly 60 nodes, sustaining a 99.95% service level agreement.
    • Describe infrastructure across 3 cloud providers entirely in Terraform and Terragrunt, keeping environments reproducible and reviewable.
    • Run cluster networking on Cilium / eBPF, with MetalLB for load balancing and cert-manager for certificate lifecycle.
    • Build and maintain the observability stack — metrics, logs and alerting tied to user-visible service levels.
    • Support IoT messaging workloads (MQTT / EMQX) and the streaming pipelines around them.
  2. Blockchain Company · confidential Apr. 2020 — Aug. 2023

    DevOps Engineer

    • Built the infrastructure foundation: network segmentation, VPN, source code management, and internal applications such as GitLab and Harbor.
    • Administered more than 5 AWS accounts, 12 Kubernetes clusters across 42 nodes, and 13 MySQL plus 10 Redis instances.
    • Owned monitoring, observability and alerting for distributed systems.
    • Defined development standards and designed the automation pipelines behind them.
  3. Solartninc 恆星網路科技 Sep. 2019 — Mar. 2020

    DevOps Engineer

    • Built and maintained internal systems — Redmine, AWX, Nexus3.
    • Administered virtual machines, databases, Synology NAS and Google Cloud Platform.
    • Introduced containerisation and moved workloads onto Kubernetes.
    • Defined development standards and designed automation pipelines.
  4. Funpodium 奕兆 Apr. 2019 — Jul. 2019

    DevOps Engineer

    • Alibaba Cloud administration.
    • Introduced Kubernetes for managing containerised applications.
    • Improved the CI/CD pipeline.
  5. Larvata 果子云 Nov. 2017 — Mar. 2019

    DevOps Engineer

    • Designed, built and unit-tested highly scalable applications.
    • Wrote technical documentation and took part in test planning, integration and deployment.
    • Responsible for workflow automation, process optimisation and performance tuning.
    • Implemented a continuous delivery pipeline with Docker, Kubernetes, CircleCI, Bitbucket and Microsoft Azure.
    • Provided maintenance support and handled incident escalations.
  6. SYSTEX 精誠資訊 Mar. 2014 — Aug. 2017

    Senior Program Analyst

    • Delivered and maintained services for internal customers.
    • Ensured data was handled securely and in line with privacy law and handling best practice.
    • Ran day-to-day operation of the CRM, version control, content survey and knowledge management systems.
    • Migrated systems from Windows to Linux and containerised them with Docker.
  7. National Formosa University Sep. 2010 — Feb. 2013

    Postgraduate research · Teaching assistant