Kashif Khan

Site Reliability Engineer

Six years of production support for a Tier-1 healthcare application, held at 99.9% uptime and 98% SLA compliance. Based in Mumbai, India.

Details

Portrait of Kashif Khan
Name
Kashif Khan
Role
Site Reliability Engineer / Application Support Engineer / Production Support Engineer
Location
Mumbai, Maharashtra, India

Awards & Recognition

Technical Excellence Award

Recognized for outstanding contribution to application reliability and incident resolution.

Star of the Team

Awarded for consistent high performance and collaborative impact.

Skills

Languages and frameworks

  • Java
  • Spring Boot
  • Python

Databases and caching

  • Oracle DB
  • MongoDB
  • Redis (Valkey)

Messaging

  • Kafka
  • RabbitMQ

Cloud and systems

  • Linux
  • AWS: EC2
  • ECS
  • Lambda
  • CloudWatch

Observability

  • Dynatrace
  • Graylog

CI/CD

  • Jenkins
  • GitLab

Summary

Site Reliability Engineer and production support professional with 6 years of experience in L2/L3 production support, incident management and Root Cause Analysis (RCA) for large-scale, highly available enterprise applications. Builds SLI/SLO-based monitoring and observability with Dynatrace and Graylog, leads Major Incident Management (MIM) and retrospectives, and keeps stakeholders informed during critical outages. Hands-on with Java/Spring Boot systems, Linux administration, shell scripting, SQL, CI/CD pipelines and AWS. Has a record of improving reliability, reducing MTTR and holding SLA compliance in high-pressure, real-time production environments.

6 yrsproduction support
99.9%uptime
98%SLA compliance
-30%MTTR
+25%service throughput

Experience

I.T. Analyst

Tata Consultancy Services Limited

1 January 2020 to present
Mumbai, Maharashtra, India

  • Reliability ownership. Own end-to-end reliability of a Tier-1 healthcare enterprise application, resolving 10+ Tier 2/3 incidents weekly while keeping 98% SLA compliance and 99.9% uptime.
  • SLO monitoring. Design and implement SLO-based monitoring for application and infrastructure components in Dynatrace APM, finding performance bottlenecks and raising service throughput by 25%.
  • Log telemetry. Build telemetry and log-monitoring workflows in Graylog, with custom metrics that pinpoint failures and cut Mean Time to Resolution (MTTR) by 30%.
  • Root cause and retrospectives. Lead root cause investigations and run retrospectives for significant incidents, working with engineering on fixes that reduced recurring incidents by 10%.
  • L3 engineering. Join architecture and design discussions for Java/Spring Boot applications, giving L3 debugging support and code-level fixes for resiliency and capacity planning.
  • Automation. Wrote Python and shell scripts for operational workflows, saving the operations team 4 to 6 hours per week.
  • Data diagnostics. Validate and extract data with SQL on Oracle and MongoDB, turning business and functional requirements into technical root-cause analysis.
  • Releases. Coordinate change management and Git/Jenkins CI/CD releases to limit service disruption during deployments.
  • Documentation and mentoring. Write and maintain operational runbooks and workflow documentation, mentor junior team members and improve onboarding for the support team.
  • Communication. Technical point of contact and culture carrier for global, cross-functional teams, explaining complex issues clearly to stakeholders to speed up resolution.

Project: FLEXFIRE gym website

A single-page website built for a gym.

View project

Education

Bachelor of Engineering, Computer Engineering

Maharashtra Education Society Pillai HOC College of Engineering and Technology

Rasayani, Raigad, Maharashtra, India

Graduated 2019 · CGPA 6.98

Hobbies

Coding
Home server
Video games