Writing
Currently, my technical writing is primarily hosted on 🤗 Hugging Face Hub and Medium. Below is a structured index of my research implementations, distributed systems debugging, and deep dives.
🚀 Featured Original Research
[Original Research][Post-Mortem]Hopper: The Optimizer That Learns Parallelism 2x Faster Than Adam — A novel optimizer I developed from scratch specifically to handle the unique optimization landscapes of Reinforcement Learning. Includes a breakdown of its 2x training acceleration over Adam and post-mortem insights into its development.[Original Research][Post-Mortem]Field Notes: Why Muon “Hollows Out” in RL (and What We Plan To DO Next) and Field Notes: The Dilemma of Training Reasoning with Muon[Ablation][Post-Mortem]When Memory Meets Memory: Engram on Transformers vs GDN Hybrids — Integrating DeepSeek’s Engram module into OLMo-core for a 2x2 architectural study. Includes a post-mortem on how a silent execution bug mimicked a false architectural gap.
⚖️ Optimization & Numerical Stability
[Post-Mortem][Investigation]The Three Horsemen of Numerical Divergence in Hybrid Models — Unpacking why certain hybrid models experience catastrophic KL divergence spikes at Step 0 of GRPO training. Testing the MiniMax fix and establishing the practical boundary for precision requirements.[Deep Dive]Going Beyond AdamW: A Practical Guide to the Muon Optimizer
🌐 Distributed Computing & Systems Engineering
[Deep Dive]The “Turtle Speed” Breakthrough 🐢✨, Part 3: My Map of the Distributed Nightmare[Deep Dive]The “Turtle Speed” Breakthrough 🐢✨ Part 2: The Blueprint for Distributed Chaos[Deep Dive]The “Turtle Speed” Breakthrough 🐢✨: Decoding Distributed Optimizers from FSDP to Muon’s Secret Sauce[Glossary]No-BS Glossary: Distributed Training 🚀 + The WHY Behind The How[Implementation]A Single Function, A World of Engineering: Deep Dive into Memory-Efficient Token Probability Calculation in Hugging Face TRL
🤖 Reinforcement Learning for Reasoning Tasks
[Implementation]Generate Once, Train Many: Implementing GRPO’s Efficient Sampling from Scratch[Deep Dive]Beyondgenerate(): A Deep Dive into Stateful, Multi-Turn LLM Rollouts for Tool Use[Deep Dive]The Evolution of Policy Optimization: Understanding GRPO, DAPO, and Dr. GRPO’s Theoretical Foundations[Deep Dive]Bridging Theory and Practice: Understanding GRPO Implementation Details in Hugging Face’s TRL Library
🛡️ AI Ethics
📰 Research Feature Writing (DeepLearning.AI)
I regularly cover cutting-edge machine learning research breakthroughs for The Batch:
[Analysis]10 Million Tokens of Input Context[Analysis]Better Spatial Perception for Robots