AI 日报hiw3c.com

Show HN: TurboGPT: train 22KiB transformer in 13s

Hacker News Top github.com 网页快照
正文为英文,可一键机器翻译(仅首次需要等待)

Notifications You must be signed in to change notification settings

Latest commit

History

Folders and files

Repository files navigation

turboGPT

Tiny byte-level GPT training in CUDA C++. MIT.

Build

nix-build -o build/nix-result

Windows, Visual Studio 2022 C++ tools, and CUDA 13.4:

.\build.ps1 - CudaArch 86

CudaArch is the GPU compute capability from NVIDIA's CUDA GPU list .

Run

.\build\ turbogpt.exe -- data hn1g.txt -- log - to runs / ctx4

The run stores its checkpoint at runs/ctx4/ctx4.pt , containing model, optimizer, scheduler, and trainer state. Use --load CHECKPOINT.pt to resume it.

runs/ctx4/report.json is derived from the log directory. Logs are TensorBoard-compatible: one report per batch, capped at 8Mi reports, and flushed with periodic or final checkpoints.

Result

hn1g after 1.5G training tokens: 2.5295 BPB .

Tests

python tests\verify.py

About

Train a tiny GPT in under a minute (CUDA only)

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages