Benchmark results
Sphragis
1 model
Select a row with diagnostics to compare accuracy by author
| Sentence tasks | Source-held-out tasks | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| OLMo-1B authorial language models | 1.18B parameters per author | Lowest per-token perplexity | 62.36 | 86.84 | 89.53 | 92.44 | 57.80 | 69.88 | 67.59 | 56.72 | Per-author base model and epoch selected on validation attribution; the source-held-out track has its own models, trained on its own split |
Benchmark results
Sphragis-Metre
1 model
Select a row with diagnostics to compare accuracy by author
| Verse tasks | |||||||
|---|---|---|---|---|---|---|---|
| OLMo-1B authorial language models | 1.18B parameters per author | Lowest per-token perplexity | 56.81 | 76.15 | 80.99 | 72.88 | Seventeen models trained on verse_1; per-author base model and epoch selected on validation attribution |