Kashif Khan
Site Reliability Engineer
Six years of production support for a Tier-1 healthcare application, held at 99.9% uptime and 98% SLA compliance. Based in Mumbai, India.
Details
- Name
- Kashif Khan
- Role
- Site Reliability Engineer / Application Support Engineer / Production Support Engineer
- Location
- Mumbai, Maharashtra, India
- Phone
- +91 9821657840
- linkedin.com/in/kashifk08
Awards & Recognition
Technical Excellence Award
Recognized for outstanding contribution to application reliability and incident resolution.
Star of the Team
Awarded for consistent high performance and collaborative impact.
Skills
Languages and frameworks
- Java
- Spring Boot
- Python
Databases and caching
- Oracle DB
- MongoDB
- Redis (Valkey)
Messaging
- Kafka
- RabbitMQ
Cloud and systems
- Linux
- AWS: EC2
- ECS
- Lambda
- CloudWatch
Observability
- Dynatrace
- Graylog
CI/CD
- Jenkins
- GitLab
Summary
Site Reliability Engineer and production support professional with 6 years of experience in L2/L3 production support, incident management and Root Cause Analysis (RCA) for large-scale, highly available enterprise applications. Builds SLI/SLO-based monitoring and observability with Dynatrace and Graylog, leads Major Incident Management (MIM) and retrospectives, and keeps stakeholders informed during critical outages. Hands-on with Java/Spring Boot systems, Linux administration, shell scripting, SQL, CI/CD pipelines and AWS. Has a record of improving reliability, reducing MTTR and holding SLA compliance in high-pressure, real-time production environments.
Experience
I.T. Analyst
Tata Consultancy Services Limited
1 January 2020 to present
Mumbai, Maharashtra, India
- Reliability ownership. Own end-to-end reliability of a Tier-1 healthcare enterprise application, resolving 10+ Tier 2/3 incidents weekly while keeping 98% SLA compliance and 99.9% uptime.
- SLO monitoring. Design and implement SLO-based monitoring for application and infrastructure components in Dynatrace APM, finding performance bottlenecks and raising service throughput by 25%.
- Log telemetry. Build telemetry and log-monitoring workflows in Graylog, with custom metrics that pinpoint failures and cut Mean Time to Resolution (MTTR) by 30%.
- Root cause and retrospectives. Lead root cause investigations and run retrospectives for significant incidents, working with engineering on fixes that reduced recurring incidents by 10%.
- L3 engineering. Join architecture and design discussions for Java/Spring Boot applications, giving L3 debugging support and code-level fixes for resiliency and capacity planning.
- Automation. Wrote Python and shell scripts for operational workflows, saving the operations team 4 to 6 hours per week.
- Data diagnostics. Validate and extract data with SQL on Oracle and MongoDB, turning business and functional requirements into technical root-cause analysis.
- Releases. Coordinate change management and Git/Jenkins CI/CD releases to limit service disruption during deployments.
- Documentation and mentoring. Write and maintain operational runbooks and workflow documentation, mentor junior team members and improve onboarding for the support team.
- Communication. Technical point of contact and culture carrier for global, cross-functional teams, explaining complex issues clearly to stakeholders to speed up resolution.
Project: FLEXFIRE gym website
A single-page website built for a gym.
Education
Bachelor of Engineering, Computer Engineering
Maharashtra Education Society Pillai HOC College of Engineering and Technology
Rasayani, Raigad, Maharashtra, India
Graduated 2019 · CGPA 6.98