About me
GenAI Solutions Architect specializing in LLM inference optimization for agentic workloads. Helped 50+ teams redesign serving infrastructure, with documented 5-10x cost reductions. Built inferenceengineering.tech (open-source, 2K+ users). Conduct security assessments on MCP deployments and run hands-on inference workshops weekly. Day-to-day: inside production GPU clusters tuning vLLM, SGLang, and TRT-LLM. Background in custom silicon, distributed training, and quantization pipelines.