verl for a single consumer GPU. PPO, GRPO and on-policy distillation on NVIDIA GPUs.
-
Updated
Sep 15, 2026 - Python
verl for a single consumer GPU. PPO, GRPO and on-policy distillation on NVIDIA GPUs.
Non-intimidating guide to create a KVM GPU Passthrough via libvirt/virt-manager on systems with only one GPU.
Train Dense Passage Retriever (DPR) with a single GPU
A 124.6M LLM trained from scratch on 13B tokens on a single RTX 4090 — tokenizer, pretraining, post-training, evaluation, and reproducible inference.
Seamless NVIDIA GPU hot handoff between Proxmox host and VM — bind/unbind nvidia ⇆ vfio-pci safely, no reboots.
A no-code browser-based tool that enables domain experts to fine-tune AI language models using their own knowledge with nothing more than a CSV file.
🚀 Achieve rapid training of NanoGPT (GPT-2 124M) on a single RTX 4090, targeting a validation loss below 3.28 with FineWeb-Edu data.
🔌 单卡 GPU LLM 推理网关 · 模型即插件 · 三态 GPU · 9 云端预设 · macOS Dashboard
Reproducibly turn a compatible Swift-Qwen3.8-27B W4A16 checkpoint into a syv-optimised single-GPU fast variant: target-calibrated draft vocabulary, GPTQ int4 lm_head/MTP, structural verification, RTX 3090 benchmarks.
Cog Single GPU Quantized Implementation of Step-Video-T2V
GPT-2-class language models trained from scratch in PyTorch on one RTX 3090, with 10B-token data curation and full GPT-2 comparisons.
Autonomous research stack for continuously improving LLM training through automated experimentation. Single-GPU research labs. Karpathy-inspired.
A lightweight, end-to-end implementation of Stable Diffusion built from first principles on a single T4 GPU. Features a custom 192-channel U-Net, VAE, and a CLIP encoder, optimized for consumer hardware and trained on approx. 168k images.
Evidence-first collection of from-scratch language models trained on a single NVIDIA L20
Evidence-first, from-scratch 1.1B English LM pretraining on 20B tokens using one NVIDIA L20
🧠 Minimal, hackable Group Relative Policy Optimization (GRPO) for LLM alignment — the algorithm behind DeepSeek-R1. Train reasoning models on a single GPU.
From-scratch 135M Transformer pretraining on 10B FineWeb-Edu tokens using a single NVIDIA L20, with public checkpoint and lm-eval comparisons.
A reproducible, memory-efficient pipeline for training a 1.2B-parameter bilingual Chinese-English GPT language model from scratch on a single RTX 4090. 一个可复现且显存高效的训练管线,用于在单张 RTX 4090 上从零训练 12 亿参数的中英双语 GPT 语言模型。
GitHub template for reproducible single-GPU ML experiments - config-driven training, multi-seed evals, VRAM budgeting.
To associate your repository with the single-gpu topic, visit your repo's landing page and select "manage topics."