Member of the Autoscaling / Capacity / Efficiency Engineering team where I build platforms, tools, and data that ensure LinkedIn's services can scale efficiently and automatically.
- Designed and built a system for automatically vertically scaling LinkedIn services to improve efficiency, reliability, and cost effectiveness, realizing over $2.8M in savings (so far).
- Designed and built a general experimentation platform that managed the setup, teardown, monitoring, and evaluation of the performance and safety characteristics of different infrastructure configurations, enabling teams to automatically test and validate the performance and efficiency of alternative resource configurations.
- Built an extensible API to aggregate data from several disparate sources, allowing internal teams to visualize data, highlight areas of improvement and savings, and understand their service footprint (~25k services).
- Part of the team responsible for building and maintaining several distributed systems interacting with internal infrastructure and Kubernetes, large data stores and data processing pipelines, CLIs, and APIs to support both vertical and horizontal autoscaling across LinkedIn services. These systems interact with internal infra, Kubernetes, Airflow, etc.
- Helped design and build a system for orchestrating, training, and storing statistical models used to predict the capacity for LinedIn's thousands of microservices, used to power an autoscaling platform. This system trained tens of thousands of models a day and was built with Airflow, Python, GRPC/protobuf, SQLAlchemy, and MySQL