The Building Blocks of an Agent Memory SystemThe Building Blocks of an Agent Memory System Most agent "memory" systems retrieve too much. They paste the last N turns into the context window, or they dump the top-K results from a vector search, and they hope the model finds the signal. The model...Apr 26, 2026·14 min read
How I Got a 60x Speedup with SoA MegakernelsLessons learned from building GPU-accelerated RL environments with NVIDIA Warp When I started building vectorized physics simulations for reinforcement learning, I made all the classic mistakes. Multiple kernel launches. Object-oriented data layouts...Jan 3, 2026·5 min read
When Less Is More: What 36 Lidar Experiments Taught Me About RL Agent PerceptionDuring my recent work on reinforcement learning agents navigating unknown environments, I observed something puzzling. My agents would often get stuck in circular scanning patterns—spinning endlessly as if searching for something. At first, I attribu...Nov 2, 2025·4 min read
Reality Navigation: Baseline Resultshttps://www.youtube.com/watch?v=U4ckkoolNH0 Quick update on our Reality simulator navigation experiments. We've moved from Y-axis rewards to proper target-based navigation with compass guidance. What Changed We switched from rewarding Y-axis movement...Oct 5, 2025·1 min read
Why Reward Shaping Sucks and What You Can Do About ItIn my previous article on reward shaping, I walked through four hard-learned lessons about balancing collision penalties in a navigation task. I eventually found the "Goldilocks solution" - a -0.1 penalty that let my agent learn to navigate obstacles...Sep 18, 2025·6 min read
The Art of Reward Shaping: Four Hard-Learned Lessons from Navigation AII was building a navigation agent with a 128-beam LIDAR scanner and FPS-style movement controls - think WASD movement with strafing and rotation. The setup seemed straightforward: navigate through a level filled with obstacles, with progress measured...Sep 15, 2025·4 min read
You Should Care About Action EntropyYou Should Care About Action Entropy I had a working RL setup. My agent successfully navigated a simple obstacle course, the reward curve was beautiful - a smooth monotonic rise to 1.0. Life was good. Then I switched to a more complex environment wit...Sep 10, 2025·3 min read