
Shivani S
Cloud Operations Engineer
Kompetenzen

Meine Dienstleistungen

Arbeitserfahrung
ITC Infotech
2 yrs 8 mos
Cloud Operations Engineer
Dec 2024 - Mar 2026 • 1 yr 3 mos
AWS Cloud Infrastructure Engineer with 1–2 years of hands-on experience managing 100+ EC2 instances across production and non-production environments, consistently maintaining 99.5% SLA uptime. Handled full EC2 lifecycle — provisioning, right-sizing, tagging, rebooting, and decommissioning. Managed EBS volumes, AMI pipelines, and EFS shared storage for backup, recovery, and application use. Built and maintained VPC infrastructure including subnets, route tables, internet/NAT gateways, security groups, and NACLs. Troubleshot routing, DNS, and inter-service connectivity issues. Managed IAM users, roles, and least-privilege policies. Integrated KMS for encryption key management across compute and storage. Administered S3 buckets with lifecycle policies and access controls for reports, logs, and backups. Managed Route 53 hosted zones and DNS records across application environments. Executed 2 patch cycles per month for Linux and Windows EC2 instances via AWS SSM — covering baseline setup, maintenance windows, deployment, post-reboot validation, and compliance reporting. Configured AWS Backup for EC2, EBS, and RDS with defined retention schedules. Supported ALB/NLB configurations, ECS services, and Auto Scaling Groups for production workloads. Used CloudWatch dashboards, alarms, and Logs Insights for monitoring and troubleshooting. Leveraged CloudTrail for audit trails and root cause analysis. Performed Linux and Windows server administration, participated in production deployments, and maintained SOPs, runbooks, and operational documentation for audit readiness.
Cloud Support Engineer
Jul 2023 - Dec 2024 • 1 yr 5 mos
Maintained high availability across AWS and Azure production environments through proactive monitoring and operational support. Monitored application and infrastructure health using AppDynamics, identifying performance degradation, application failures, node issues, and service disruptions. Utilized AWS CloudWatch for monitoring infrastructure health, resource utilization, log analysis, and alert investigation. Leveraged AWS CloudTrail for auditing account activities, tracking configuration changes, and supporting incident investigations. Monitored production systems, analyzed alerts, and ensured timely resolution within defined SLA targets. Delivered 24x7 cloud operations and infrastructure monitoring support across AWS and Azure environments while maintaining SLA compliance. Provided L1/L2 production support for cloud-hosted applications and infrastructure services. Managed service requests, incidents, and operational tickets using ServiceNow, JIRA, and Freshservice. Monitored infrastructure and application performance using AppDynamics, CloudWatch, and New Relic. Monitored Azure Kubernetes Service (AKS) pod health, Azure Backup jobs, ECS services, Auto Scaling activities, and ElastiCache cluster health. Managed Apache Airflow DAG monitoring and incident response; investigated Elastic Load Balancer target group health and traffic-routing issues. Configured AWS Lambda schedules for automated server start/stop operations; utilized Amazon S3 for reporting and operational activities. Prepared daily and weekly operational reports using Azure Monitor, Azure Backup, and Azure Portal dashboards covering backup failures, server health, CPU utilization, and memory metrics. Built and maintained operational runbooks, SOPs, and knowledge base documentation.