Muneeb Ahmad

Hi, I'm an M.S. student in Electrical and Computer Engineering at UC Davis (2025–2027), advised by Prof. John Owens. I work on GPU architecture and performance, and on the GPU–ML stack: custom attention kernels for LLM inference, KV cache compression, and vLLM integration.

This summer I interned with Qualcomm's GPU Research team in San Diego, analyzing and improving ray-tracing performance on Adreno GPUs. I received my B.Tech. in Computer Engineering from Jamia Millia Islamia.

I am looking for roles in GPU architecture, GPU performance, and ML inference.

At a glance

Picture of Muneeb Ahmad in an Arsenal Football Club Home Jersey.
Yes, I am a Gunner

Experience

GPU Research Intern, Qualcomm

San Diego, CA

  • Analyzed ray-tracing performance on Adreno GPUs for both Vulkan and DirectX12, on both ray pipeline and ray query workloads.
    • Identified bottlenecks using hardware performance counters and traced fixes back to the shader compiler codegen.
    • Fixes resulted in a 2x speedup in targeted configurations.
    • Identified memory overallocation bugs in the compiler.
  • Proposed memory layouts that achieve 100% (2x existing) memory bandwidth utilization for ray-tracing internal data by improving coalescing behavior; the design also improves cache-friendliness.
  • Created AI agent-based tooling to aid in end-to-end on-device performance analysis, regression testing, and root-causing.

Graduate Student Researcher, Owens Group

UC Davis

  • Created optimized custom GPU kernels for compressed paged attention with fused RoPE using DSLs at various levels of abstraction Triton, Gluon and HipKittens.
  • Integrated custom attention paths into vLLM, including managing the KV cache in bespoke formats.
  • Implemented gesvda on top of rocSOLVER for 10x faster approximate singular value decomposition on AMD GPUs.
  • Developement of hardware-efficient methods to compress the KV cache 8–30x with minimal accuracy degradation (<2% on task performance).
  • Extensively evaluated LLM task performance using models like Llama and Qwen3 on benchmarks like LongBench and GSM-8K.

Researcher, BeyondDefence Lab

University of New Mexico (Remote)

  • Created a framework to find the optimal deployment strategy for a given deep learning model on a given edge device, targeting satellites and robotics.
  • Achieved a ~400% improvement in latency on the Coral Edge TPU dev board.
  • Worked with hardware accelerators such as the Coral Edge TPU and NVIDIA Jetson Nano.
  • Applied quantization to models and transformed formats for compatibility with different hardware.

Selected Projects

All projects →

Education

M.S., Electrical and Computer Engineering

University of California, Davis

  • Advisor: John Owens
  • Coursework: Modelling & Optimization in Computer Engineering, Large-scale Scientific Computing, High Performance Statistical Computing, Computer Architecture, Information Theory & Coding

B.Tech., Computer Engineering

Jamia Millia Islamia, New Delhi, India

  • CGPA: 9.05 / 10
  • Coursework: Computer Networks, Network Security, Computer Architecture, Digital System Design, Embedded Systems, Operating Systems, Parallel & Distributed Computing

Skills

Contact

Email: muneeb0ahmed3 [at] gmail [dot] com
Resume: PDF