- Designed a kernel subsystem for chunk-based storage in a distributed file system bridging multiple backends, scaling to exabyte levels with tenant performance isolation.
- Optimized read path latency by 50% on industry-standard benchmarks by eliminating redundant interactions between shared buffer caches in the kernel's VM subsystem.
- Resolved 60% CPU overutilization from lock contention on hot delta blocks, improving cluster stability.
- Built NVMe-oF device driver subsystems with RDMA transport, achieving 2× higher per-node throughput.
- Performance analysis of a GPU-orchestrated file system against GPUfs, cuFile, GeminiFS — showing 1.61× speedup over existing storage paths.
Experience
- Built a lock-free I/O sampling subsystem exposing per-application Lustre file system metrics from kernel to userspace.
- Developed portable eBPF programs with libbpf instrumenting MPI I/O libraries, syscalls, and the VFS with low-overhead probes.
- Roofline modeling across Nvidia A100 and AMD MI250X; compared vendor and open-source device compilers to identify performance gaps.
- Performance tuning of CXI libfabric on HPE Slingshot for GPU-NIC RDMA and GPU Peer2Peer IPC MPI transfers.
- Analyzed a Python/C++ BLAS framework across NUMA configs on Cray x86 and IBM Power9 supercomputers.
- Taught and mentored 150+ undergrad students over 3 semesters — C memory management, CPU microarchitecture, systems programming.
Education
Research & Publications
Skills
Blog
No posts yet.
Thoughts on systems, storage, and performance — coming soon.