AshuGPT
Decoder-only LLM, pretrained from scratch
A 124M-parameter transformer trained end to end in raw PyTorch — RoPE, RMSNorm, SwiGLU, KV-cache attention and a byte-level BPE tokenizer, all hand-written across 6,000 lines. No transformers.AutoModel: every layer implemented and measured.
- 23.53 validation perplexity on 2.46B tokens of FineWeb-Edu — 20,000 steps, 27 h on a single RTX 2080 Ti at 25K tok/s.
- Cut peak memory 50.9% with gradient checkpointing, accumulation and fused attention, all gradient-identical; FSDP with CPU offload shards a 1.23B config to 2.04 GB/GPU.
- Sequence packing reclaimed the 89% of each fine-tuning step lost to padding, for 4.40x supervised throughput; DPO on HH-RLHF moved the win rate from 50% to 57.1%.
- Served from Hugging Face Spaces behind a streaming, KV-cached FastAPI endpoint at 80 tok/s, under 441 tests on GitHub Actions.
- PyTorch
- CUDA
- DDP / FSDP
- FastAPI
- HF Spaces