gpu_nn_framework
A neural-network framework built from scratch in C++17 with optional CUDA acceleration. Implements a full forward/backward pass via a custom reverse-mode autograd engine, common layers, losses and optimizers — no external ML libraries.
v2 is a correctness rewrite. The original autograd had architectural bugs that prevented any learning at all (see What was fixed). This version passes a 2,000+ assertion test suite including finite-difference gradient checks over every parameter of a full MLP.
Features
- Tensor — n-dimensional float array; CPU + CUDA storage; NumPy-style
broadcasting on
+ - * /; views (reshape),transpose,squeeze,clone/detach - Autograd — reverse-mode AD over a dynamic graph; single-pass
topological engine (correct for fan-out/diamond graphs);
NoGradGuardfor inference;manual_seed()for reproducibility - Layers —
Linear,ReLULayer,SigmoidLayer,TanhLayer,Dropout,Sequential - Losses — MSE, cross-entropy (fused stable softmax+NLL), binary cross-entropy, NLL — all fully differentiable
- Optimizers —
SGD(momentum, weight decay, Nesterov),Adam - CUDA kernels — tiled GEMM, elementwise activations, axpy/fill
- OpenMP — multithreaded CPU matmul and Linear forward/backward
Tensors are handles: copying a Tensor copies a pointer, and both handles
refer to the same data and the same autograd node (exactly like PyTorch). Use
clone() for a deep copy and detach() to drop out of the graph.
What was fixed
The v1 autograd could not train anything. The root causes:
- Graph severed at almost every op. The copy constructor deep-copied data
but silently dropped
grad_fn, and ops “kept” their inputs by deep-copying any tensor not created throughmake_tensor.cross_entropy_losstherefore stored a grad-less copy of the logits, ReLU a grad-less copy of the linear output, and so on — backprop stopped one node below the loss and parameters never received gradients. Fixed by the handle/TensorImplredesign: every consumer holds the same graph node, andGradFn::inputsowns real handles. - Use-after-free in
LinearBackward. It held a fake non-owningshared_ptrto a stack local insideSequential::forward, which was destroyed beforebackward()ran. - Double gradient propagation. The engine iterated a topological list and each backward recursively invoked its inputs’ backwards, so gradients below the root were counted multiple times (up to 2^depth on diamond graphs). The engine now does a single topological pass; each node fires exactly once with its fully accumulated gradient.
Tensor::sum(dim)had undefined behaviour —for (int i = dim + i; …)readsiin its own initializer (should bedim + 1).MatMulBackwardcomputed wrong numbers. It fed a strided transpose view into a matmul that reads raw pointers assuming contiguous memory.transpose()now materializes a contiguous copy.sigmoid()on CUDA launched the ReLU kernel.- Scalar
+,-,/silently detached the graph (only*had a backward). All four now have backwards. binary_cross_entropy_lossandnll_losshad no backward at all.reshape()on non-contiguous views silently reordered data — now guarded; views also propagate gradients (they previously dropped them).- The
CUerror macro inmatmul.cuwas missing line continuations, so the CUDA translation unit could not compile; and there was noCMakeLists.txtdespite the README referencing one. - Init wasn’t Kaiming.
tanh(randn)·boundreplaced with PyTorch’s actualnn.Lineardefault:U(-1/√fan_in, +1/√fan_in)for weight & bias.
Improvements beyond fixes: broadcasting on all elementwise ops (with correct
gradient reduction), softmax/log_softmax with backwards, Dropout, Adam,
SGD momentum/weight-decay/Nesterov, NoGradGuard, manual_seed, double
accumulators in reductions/losses, OpenMP + SIMD-vectorized Linear kernels,
fused single-loop optimizer updates (no per-step temporaries), and a
benchmark pair for head-to-head comparison with PyTorch.
Both run the identical workloads (square matmuls, Linear inference, a full MLP training step, elementwise add). Honest expectations:
- Large GEMMs: PyTorch calls MKL/OpenBLAS (CPU) or cuBLAS (GPU) — decades of hand-tuned assembly. A hand-rolled kernel will not beat those; expect PyTorch to win by several× on 512²+ matmuls. Closing that gap means register blocking + explicit AVX/FMA kernels (on the roadmap).
- Small tensors / small-batch steps: this framework has near-zero per-op overhead (no dispatcher, no dtype/device dispatch, no Python), so on tiny workloads — e.g. small-MLP training steps at batch ≤ 64 — it can match or beat PyTorch’s per-step latency, especially single-threaded. This is the realistic place to “win”, and the benchmark pair lets you measure it on your machine rather than take anyone’s word for it.
Operation coverage
| Op | CPU | CUDA | Backward |
|---|---|---|---|
+ − × ÷ (tensor, broadcast) |
✅ | — | ✅ |
+ − × ÷ (scalar) |
✅ | — | ✅ |
matmul |
✅ (OpenMP) | ✅ (tiled) | ✅ |
relu / sigmoid / tanh |
✅ | ✅ | ✅ |
softmax / log_softmax |
✅ | — | ✅ |
exp / log / pow / sqrt |
✅ | — | ✅ |
sum / mean (full or dim) |
✅ | — | ✅ |
reshape / transpose / squeeze / unsqueeze |
✅ | reshape only | ✅ |
| losses (MSE / CE / BCE / NLL) | ✅ | — | ✅ |
Linear fused fwd/bwd |
✅ (OpenMP+SIMD) | — | ✅ |
in-place add_ / mul_ / fill_ / zero_ |
✅ | ✅ | n/a |
End-to-end GPU training is not possible yet (losses/reductions are CPU-only); the CUDA kernels accelerate the ops marked above. The CUDA path compiles against CUDA 11+ but was not run in this environment — treat it as best-effort until you’ve run the test suite on a GPU box.
Roadmap
- [ ] CPU GEMM register blocking + AVX/FMA micro-kernel
- [ ] CUDA paths for sum, softmax, cross-entropy → full GPU training
- [ ] Stream-based async CUDA (currently syncs per launch)
- [ ] Batch normalization, convolutional layers
- [ ] Dataloader / mini-batching utilities
- [ ] Graph freeing after backward (
retain_graph=falsesemantics)
Migration notes (v1 → v2)
make_tensor(...)/Tensor::self_are gone — plainTensorvalues track gradients correctly now:Tensor x({3.f},{1},Device::CPU,true); (x*x).backward();t.grad(public member) →t.grad()returns aTensorhandle; check with.defined().Layer::parameters()returnsstd::vector<Tensor>handles instead of rawTensor*.SGDmoved fromlinear.htooptim.h(and gained momentum/weight decay);Adamis new.- Copying a
Tensorno longer deep-copies — useclone()for that. transpose()returns a materialized contiguous tensor, not a view.cuda/relu.cu→cuda/elementwise.cu(addsaxpy,fill).