I'm a developer exploring how large language models can be useful outside the cloud β on laptops, phones, and other resource-constrained devices.
Right now I'm focused on:
- π§ LLMs β inference, prompting, evaluation, and practical use cases
- π± Edge / on-device AI β quantization, GGUF, CPU-friendly runtimes
- β‘ Systems thinking β how to ship models that stay fast and small
- π Growing in public β learning by building, writing, and shipping
I like turning research ideas into something that actually runs on real hardware.
- Quantizing & optimizing open-source models for CPU / low-memory machines
- Local inference stacks (
llama.cpp, Hugging Face, PyTorch) - Small experiments that connect ML systems with everyday products
β¨ Learning in public Β· building toward edge-ready AI β¨


