Empirical SRAM Gating and Semantic Ghosts
A V2 architecture for KV cache optimization that replaces crude token dropping with empirical compression, squashing low-attention tokens into dense semantic ghosts.
Practical AI systems notes from building and testing real-world AI products.
These are not formal academic papers. They are technical memos, experiments, and observations from hands-on work with AI systems.
A V2 architecture for KV cache optimization that replaces crude token dropping with empirical compression, squashing low-attention tokens into dense semantic ghosts.
An architectural KV-cache experiment that retained 16.5% of the baseline cache footprint and increased measured throughput by 27% in a high-concurrency Qwen 0.5B stress test, while revealing limitations of blunt token eviction.
Just a project I was tinkering with because I think physics is cool. A custom activation approximating the Brachistochrone curve for LLM training.
M.Sc. thesis work on reducing object detection annotation effort through uncertainty-based active learning.
Undergraduate thesis work on GPS-free position estimation using IMU data, quaternion-based orientation estimation, and deterministic drift correction.