MiniMind

Github
2025-10-20

Description

MiniMind is an ultra-lightweight, open-source large language model project that boasts a version as small as 1/7,000 the size of GPT‑3 — making it feasible to train on a standard single consumer GPU. It includes not just the model architecture (built from scratch in native PyTorch, without relying on high-level libraries) but also full pipelines for mixed-expert modules (MoE), data cleaning, pre-training, supervised fine-tuning (SFT), LoRA tuning, direct preference optimization (DPO), and model distillation. With a vision toward both researchers and beginners, MiniMind also offers a vision-multimodal variant (MiniMind-V) and emphasizes hands-on learning and democratizing access to large-model creation.

User Reviews

No reviews yet for this resource.