Senior Site Reliability and Platform Engineer
Senior Site Reliability and Platform Engineer with 10+ years operating production systems at scale. Specialized in AWS, Kubernetes, AI/LLM platforms, incident response, automation, and reliability engineering. Experienced building agentic AI systems using Bedrock, LangGraph, MCP, and cloud-native platforms while maintaining production SLOs through automation, observability, and operational excellence.
Production architectures, intelligent agents, and cloud-native labs
Agentic system built with LangGraph + FastAPI + FAISS that performs automated code reviews on PR diffs. Integrates real tool calling (ruff linter, test coverage), RAG over repo conventions, dynamic context trimming, and structured JSON logging. Deployed via Docker Compose with GitHub Actions CI/CD pipeline.
View Project →Tools and platforms I use daily in production