Self-taught engineering deep dives — built in public and taught from first principles, against real hardware.
How LLM inference really works — from token to GPU kernel — bridging the ops layer down into the hardware, taught against a real 4×H100 cluster.