TinyGPU
A custom SIMD vector accelerator wired into a RISC-V core — running on real FPGA silicon, not simulation.
15.61x speed up compared to CVA6
View repohi! im atharva, going through life one thread at a time.
See my workA custom SIMD vector accelerator wired into a RISC-V core — running on real FPGA silicon, not simulation.
15.61x speed up compared to CVA6
View repoHand-written CUDA GEMM kernels, iteratively optimized against cuBLAS using NSight Compute profiling.
Naive → optimized, benchmarked in GFLOPs across matrix sizes
View repoI believe hardware is the future for accessible, intelligent systems and I want to work on improving the current capabilities chips that handle AI workloads.