01$ whoami

jonathan ng

site reliability engineer / maker-of-things / toronto, ca

I build and keep distributed systems healthy — observability, internal platforms, and the automation that makes shipping calm. Off the clock it's homelab, custom keyboards, cooking, fishing, and far too many video games. See what I'm working on here, read the blog, or poke around the homelab.

[toronto] [observability] [homelab] [keyboards] [game-dev]

02$ cat ~/projects.md

  • getbud ↗2026–now

    A budgeting app, import/track/visualize spending with ease

  • waypoint ↗2026–now

    A collaborative travel companion, plan/live/remember your trips

  • A fast-paced team-based first person arena game

  • Keep browser tab bulk at a minimum, while preserving important tabs

  • A static site generator to render my Obsidian vault as an inter-connected blog

  • kube-cats ↗2025–now

    A fun Kubernetes workload visualizer using pixel cats written with Go and React

  • home-lab ↗2023–now

    A home lab server for learning and hosting passion projects

  • keyboards ↗2023–now

    Custom keyboard designs, firmware, layouts and builds

  • A simple pomodoro app to learn React

  • A script to pull freebies from Epic, supporting docker, k8s and github actions

  • A personal Discord bot developed to provide information from various game and movie APIs

  • luxify2020–2021

    A restock notification service using Facebook's messaging API and various online store APIs to notify userbase as soon as highly coveted items are in stock

03$ cat ~/work_experience.md

StackAdapt2024–nowDevOps Engineer (Observability and Internal Services)
  • Rollout of metadata labelling standardization across Kubernetes infrastructure, enabling support for dynamic owner-based service alert routing
  • Improve reliability of Observability stack, including migrating Grafana and Grafana OnCall backends from sqlite to postgres, and enabling Grafana to operate in high-availability mode
  • Rollout of Gatus and accompanying Prometheus alerts to enable endpoint monitoring and quick feedback to critical service teams
  • Migrating Grafana and Loki from EC2 to ECS, improving scalability and reliability of core observability services
  • Rollout structured metadata usage on Loki for k8s resources, improving cardinality and performance on Loki
  • Designed and conducted an observability workshop for Engineering department with over 60 attendees, covering metrics, logging, tracing, profiling, and alerting and providing working examples via locally runnable stack
  • Automate deployment of Oncall and alert rule changes along with rolling out live alert rule validation using pint
Tempo Software2022–2024SRE II
  • Architected internal tooling enabling developers to develop fully-scoped prod-like ephemeral environments for feature testing, used by 50+ developers across 9+ teams, improving release confidence and increasing release frequency 718%
  • Joined initiative as a developer to refactor auth service, enabling feature management via LaunchDarkly
  • Modularized legacy IaC to simplify provisioning of environments, incorporating GitOps by leveraging terragrunt and GitHub Actions to promote changes between environments
  • Co-led the migration of code base and artifacts from GitLab to GitHub consolidating 2 product lines under one organization
  • Directed the knowledge-sharing process for onboarding new hires, growing the team to 6 members
  • Implemented Datadog monitoring for k8s services, improving coverage by enabling support for opentelemetry to capture custom metrics
  • Improved stability and consistency of production rollouts with readiness gates and pod disruption budget, reducing release downtime with ALB ingress controllers to zero
  • Refactored release communication to more align with GitOps strategy, enabling consistent notifications to stakeholders as features get promoted into production and increasing service coverage to 100%
  • Improved security of CI workflows across the organization by leveraging OIDC roles for AWS authentication, eliminating need for access keys in CI as well as creating a standard for repo and service based AWS access
  • Led initiative to refactor services to leverage internal communications where possible, increasing visibility in product's inter-service communications
  • Directed new SSO permission strategy for standardizing AWS permissions for SREs and developers across the company, leaning into GitOps and management of IaC
TimePlay (now Stream6ix)2021Junior DevOps
  • Used Ansible to manage scaling game infrastructure on AWS for a projected 700% increase in player traffic
  • Deployed infrastructure using AWS API Gateway, Lambda and ECS to reduce resource spin up times 93% down to 30s for new line of on-demand games
  • Created CI/CD pipelines for Unity projects using Unity Cloud Build to leverage tailored support and integrations
Cryptonumerics (now Snowflake)2019DevOps Cloud Developer Intern
  • Designed OAS dataset retrieval API to support popular cloud storage solutions
  • Refactored Spectron testing suite, improving test consistency by 100% and cutting test time down 89%
VIA Rail Canada2019Innovation Engineer Intern
  • Developed telemetry solution using Kibana to monitor train health and activity, leveraging existing sensor data installed throughout the train and enabling engineers to diagnose issues remotely
  • Developed a bash script to automate set up of analytics solution across scalable fleet of train cars
  • Applied beacon technology to map customer journeys through trainstation via device pings
TD Bank2017–2018DevOps Engineer Intern
  • Generated a daily health report for stakeholders by compiling Jenkins build, and SonarQube results with Groovy
  • Coordinated migration of over 80 projects from several outdated instances of Jenkins, proposing a workflow to reduce over 90% of the planned work
  • Architected and introduced an experimental shared library system on Jenkins to streamline continuous integration pipeline for over 5 project teams
  • Administrator over Atlassian Toolstack, providing support and provisioning across all cloud platforms
Rave2017Backend Engineer Intern
  • Migrated video service from Postgres to Google Cloud Datastore, to improve consistency of transactions
  • Refactored location microservice, reducing redis queries by 50%

full detail on the full resume ↗

04$ git log --since=1.year

1,007 contributions in the last year · github.com/j6nca ↗

OctNovDecJanFebMarAprMayJunJulAugSep
LessMore

© Jonathan Ng