Inference platforms
Kubernetes-native model serving with KServe, vLLM and llm-d. Model-aware routing through Envoy AI Gateway, budgets and keys via LiteLLM, and queue-depth autoscaling that adds replicas before it adds GPUs.
~/ shubham tatvamasi
Kubernetes-native inference, GPU platforms, GitOps and quantum-safe networks, engineered so models run fast, scale calmly and never go dark.
Most people meet AI through a chat box. I live one layer down, where tokens become GPU cycles, GPU cycles become heat, and a single misrouted request can idle a rack. My work is making that layer boring in the best way: declarative, observable, self-healing, and secure enough that nobody has to think about it.
Making the whole stack, from the GPU die to the API gateway, behave like one calm, observable, self-healing system.
Kubernetes-native model serving with KServe, vLLM and llm-d. Model-aware routing through Envoy AI Gateway, budgets and keys via LiteLLM, and queue-depth autoscaling that adds replicas before it adds GPUs.
Telling allocated apart from actually computing. Cross-vendor telemetry that catches idle, starved and oversubscribed accelerators.
Controllers, CRDs and operators in Go. Multi-cluster networking, tenant isolation and HA storage that heals itself.
Declarative reconciliation across fleets of clusters. If it isn't in Git, it doesn't exist.
Post-quantum key exchange, signatures and certificate lifecycles, automated as Kubernetes-native resources.
Encrypted tunnels from devices behind carrier-grade NAT to private and public clouds. Fleets that stay reachable, observable and patched no matter where they live.
Quantum-safe security for Kubernetes. CRDs and controllers that generate, rotate and sign with ML-KEM and ML-DSA, and issue X.509 certificates, so post-quantum crypto becomes a manifest instead of a migration.
Multi-model LLM serving on Kubernetes: LiteLLM for keys and budgets, Envoy AI Gateway for model-aware routing, llm-d for KV-cache-aware endpoint picking, vLLM on GPU node pools, and KEDA scaling on queue depth.
A living, public GitOps tree: Flux CD bootstrap, cluster, infrastructure and app layering, and secret management patterns, all reconciled continuously from Git.
One-command, Ansible-driven provisioning of the Magma 5G core on private-cloud Kubernetes, the path that took the project beyond a single cloud.
A short operating manual for the systems I build, and the way I build them.
Every cluster, model and secret should be reproducible from a commit. If a human has to remember it, it will be forgotten.
Allocated isn’t utilised. Ready isn’t responsive. I chase the metric that tells the truth, not the one that looks good on a dashboard.
Replicas before nodes, caches before GPUs, routing before hardware. Efficiency is a design decision, not a cost report.
Harvest-now, decrypt-later is real. Post-quantum crypto belongs in today’s pipelines, not tomorrow’s roadmap.
The best infrastructure ideas get sharper in public. Ship it, document it, let others break it.
A tiny, very real-feeling terminal. Type help to look around, or try nvidia-smi.