🧬 micro-GPT
← Portfolio
Loss — against the two numbers that bracket it
Training loss (nats/character)
Bigram entropy — counting alone
ln V — knowing nothing at all
What it writes, right now
Attention — who each character is looking at
0
1
Data
Corpus
This site, both languages
This site, English only
This site, Romanian only
Model
Width
32
Layers
2
Heads
2
Context
32
Peak learning rate
0.0100
Train
Reset
Sampling
Prompt
Temperature
0.70
Top-k
10
Start again from the prompt
Numbers