Skill requirements: Splunk, ELK, Dynatrace, AppDynamics, Grafana
1. Proficiency in site reliability engineering (sre) principles and practices.Β
2. Strong background in system administration, networking, and cloud computing.
3. Experience with monitoring tools such as prometheus, grafana, and elk stack.
4. Knowledge of containerization technologies like docker and kubernetes.
5. Ability to troubleshoot complex technical issues and perform root cause analysis.
6. Excellent communication skills and ability to work collaboratively in a team environment.
7. Strong project management and leadership skills to drive initiatives and deliver results efficiently.
8. Certifications in relevant areas such as aws certified devops engineer or google professional cloud devops engineer are a plus.
"1. Ability to plug-in the findings into CI/CD pipeline to stop code deployments beyond Beta1
2. Document the observations via feedback loop and enable the process on the feedback loop to improvise the agent.
Ability to see the automated health checks with the built pipeline with the customer's session details"