SP
Shubh Pachchigar
Software Engineer 2
HPC Storage · Dell Technologies
Seattle, WA
Kernel Dev & HPC Storage

I'm a software engineer working on distributed file systems and HPC storage infrastructure at Dell Technologies. My work lives deep in the kernel — building chunk-based storage subsystems, tuning buffer caches, and designing NVMe-oF device drivers with RDMA transport.

Previously at NERSC / Lawrence Berkeley National Laboratory, I built eBPF-based tools to characterize HPC workloads at the system level — instrumenting MPI I/O libraries, VFS, and syscalls on Cray x86 and IBM Power9 supercomputers.

MS in Computer Science from SF State (GPA 3.9). I like problems where the hardware and software boundary blurs — storage stacks, GPU memory, RDMA, and anything where getting closer to the metal means getting faster.

Distributed File Systems Kernel Development eBPF NVMe-oF / RDMA GPU Computing HPC Storage MPI

Experience

Dell Technologies
Software Engineer 2 · File System and Data Services Group (HPC Storage)
  • Designed a kernel subsystem for chunk-based storage in a distributed file system bridging multiple backends, scaling to exabyte levels with tenant performance isolation.
  • Optimized read path latency by 50% on industry-standard benchmarks by eliminating redundant interactions between shared buffer caches in the kernel's VM subsystem.
  • Resolved 60% CPU overutilization from lock contention on hot delta blocks, improving cluster stability.
  • Built NVMe-oF device driver subsystems with RDMA transport, achieving 2× higher per-node throughput.
  • Performance analysis of a GPU-orchestrated file system against GPUfs, cuFile, GeminiFS — showing 1.61× speedup over existing storage paths.
NERSC, Lawrence Berkeley National Laboratory
Software Engineering Intern · HPC
  • Built a lock-free I/O sampling subsystem exposing per-application Lustre file system metrics from kernel to userspace.
  • Developed portable eBPF programs with libbpf instrumenting MPI I/O libraries, syscalls, and the VFS with low-overhead probes.
  • Roofline modeling across Nvidia A100 and AMD MI250X; compared vendor and open-source device compilers to identify performance gaps.
  • Performance tuning of CXI libfabric on HPE Slingshot for GPU-NIC RDMA and GPU Peer2Peer IPC MPI transfers.
  • Analyzed a Python/C++ BLAS framework across NUMA configs on Cray x86 and IBM Power9 supercomputers.
San Francisco State University
Graduate Teaching Associate
  • Taught and mentored 150+ undergrad students over 3 semesters — C memory management, CPU microarchitecture, systems programming.

Education

San Francisco State University
MS Computer Science · GPA 3.9 / 4.0
Nirma University, India
BTech Electrical Engineering · GPA 8 / 10

Research & Publications

HPC Workload Characterization Using eBPF
Pachchigar, S., Friesen, B., Cook, B.
Cray User Group 2025
Beyond the Imitation Game (BIG-bench)
Srivastava, A., A., et al. (incl. Pachchigar, S.)
TMLR 2023

Skills

Languages
CC++CUDA PythonBash x86 ASMARM ASM
Tools
eBPF / libbpfGDBPerf CMakeDockerMLIR NSysNCUGit
Domains
Distributed File SystemsKernel Dev HPC StorageNVMe-oF / RDMA GPU ComputingMPI

Blog

No posts yet.

Thoughts on systems, storage, and performance — coming soon.